検索を外部の検索エンジンに外出しする
本体の標準機能ではありません
プリザンター本体の全文検索は RDBMS の機能だけで動いています。このページは外部の検索エンジンに移す場合の設計メモです。本体を改修せずに外部の検索エンジンを使う例は Fess で全文検索にあります。
前提にした現行実装
詳しくは検索機能の内部実装にまとめています。改修に関係する点だけ挙げます。
- 各レコードの内容をつないだ文字列を
Items.FullTextに入れ、検索はこの列を見る。書き込みはIndexes.CreateFullText()(Indexes.cs#L131-L145)で、FullTextとSearchIndexCreatedTimeだけを更新する。 - バックグラウンドの
Indexes.RebuildSearchIndexes()が、SearchIndexCreatedTimeが古いレコードを作り直す(Indexes.cs#L1048)。 - 横断検索は
Indexes.Get()(Indexes.cs#L748-L760)がISqlCommandTextの RDBMS ごとの実装で SQL を作る(SQL Server はCONTAINS、PostgreSQL はpg_trgmの%>、MySQL はMATCH ... AGAINST)。権限は同じ SQL の中でDef.Sql.CanReadを条件に入れて絞る。 - 一覧の検索はサイトの検索方式(
FullText・PartialMatch・MatchInFrontOfTitle・BroadMatchOfTitle)に従い、View.SetSearchWhere()が別に条件を作る。Indexes.Get()を通らない。 - 検索語は
SearchIndexes()→Words()→FullTextClause()で分割・かな変換される。Libraries/Search/WordBreaker.csという文字種で分割するクラスもあるが、1.5.8.1 ではどこからも参照されていない。 - 添付ファイルの中身の検索(
SearchDocuments)はBinaries.Binを RDBMS の全文検索で直接見るもので、実用になるのは SQL Server だけ。
候補のエンジン
エンジンの特徴は各プロジェクトの公開情報に基づきます(調査時点)。
| エンジン | .NET クライアント | ライセンス | セルフホスト | 分散構成 | 日本語解析 |
|---|---|---|---|---|---|
| Elasticsearch 8 | Elastic.Clients.Elasticsearch(公式。旧 NEST は 7.x 用で EOL) | SSPL / Elastic License 2.0(要確認) | 可 | 可 | kuromoji(analysis-kuromoji)・ICU |
| OpenSearch 2 | OpenSearch.Client(公式。NEST のフォーク) | Apache 2.0 | 可 | 可(AWS は Amazon OpenSearch Service) | kuromoji |
| Apache Solr 9 | SolrNet(コミュニティ、1.2.1。9.x 対応は一部ベータ) | Apache 2.0 | 可(Java 21 以上) | SolrCloud | ICU・カスタム設定。kuromoji は既定で同梱されない |
| Meilisearch | Meilisearch(公式) | MIT | 可(Windows は Server 2022 以降が公式対応) | 不可(単一ノード) | Lindera(形態素解析、内蔵) |
| Typesense | Typesense(コミュニティ、8.1.0) | GPL-3.0(要確認) | 可(Windows は Docker 経由のみ) | 全量複製の冗長構成 | 形態素解析なし(事前に分割が要る) |
| Manticore Search | manticoresearch-net(公式)。MySQL プロトコル互換なので MySQL 用クライアントでも可 | GPL-2.0(要確認) | 可(Windows インストーラあり) | Galera | ngram(bigram) |
| Azure AI Search | Azure.Search.Documents(公式、11.7.0) | Azure のサービス | 不可 | マネージド | ja.microsoft |
| Algolia | Algolia.Search(公式、7.38.x) | SaaS | 不可 | マネージド | 形態素解析なし(事前に分割が要る) |
いずれも net10.0(1.5.8.1 のターゲット)から使えます。選び方の目安は次のとおりです。
| 重視すること | 候補 |
|---|---|
| 日本語の検索精度 | Elasticsearch / OpenSearch(kuromoji) |
| 導入の手軽さ | Meilisearch |
| ライセンスの制約がないこと | OpenSearch / Apache Solr |
| 少ない資源で動かす | Manticore Search |
| Azure 前提 | Azure AI Search |
| AWS 前提 | Amazon OpenSearch Service |
| インフラを持たない | Algolia(データをクラウドに送ること、量が増えると費用がかさむことに注意) |
改修の構成
RDB の Items.FullText への書き込みはそのまま残し、外部エンジンにも同じ内容を送ります。検索は外部エンジンから一致した ReferenceId の一覧を受け取り、権限の絞り込みは DB で行います。
図を読み込み中…
Parameters.Search.ExternalSearch.Enabled のようなフラグで切り替え、無効なら現行の SQL 検索をそのまま使う形にすると、外部エンジンを使わない環境に影響しません。
改修するファイル
| ファイル | 内容 |
|---|---|
Implem.ParameterAccessor/Parts/Search.cs・App_Data/Parameters/Search.json | 接続情報の追加(下記) |
新規 Libraries/Search/IFullTextSearchEngine.cs | エンジンを差し替えるためのインターフェース |
新規 Libraries/Search/ElasticsearchEngine.cs など | エンジンごとの実装 |
Libraries/Search/Indexes.cs の CreateFullText() | DB 更新の後に外部エンジンへ送る |
Libraries/Search/Indexes.cs の Get() | 外部エンジンから ID を取って DB で絞る |
Libraries/Search/Indexes.cs の RebuildSearchIndexes() | 外部エンジンへの一括再送 |
レコードの削除処理(IssueModel・ResultModel・WikiModel の Delete() の呼び出し元) | 外部エンジンからも削除 |
一覧の検索(View.SetSearchWhere())も外部エンジンに向けるなら、そちらも別に改修が要ります。
パラメータ
1.5.8.1 の Search クラスの項目は SearchDocuments・CreateIndexes・PageSize・DisableCrossSearch・DisableCrossSearchSites・FullTextIncludeBreadcrumb・FullTextIncludeSiteId・FullTextIncludeSiteTitle・FullTextNumberOfMails・FullTextMaxNumberOfMails です(Search.cs)。ここに外部エンジンの設定を足します。
public class Search
{
// 既存の項目はそのまま
public ExternalSearch ExternalSearch;
}
public class ExternalSearch
{
public bool Enabled; // 外部エンジンを使うか
public string Engine; // "Elasticsearch" | "OpenSearch" | "Solr" | "Meilisearch" など
public string Url; // 例: "http://localhost:9200"
public string IndexName; // 例: "pleasanter"
public string ApiKey; // 省略可
public string CertificateFingerprint; // 省略可
}インターフェース
public interface IFullTextSearchEngine
{
void Index(long referenceId, long siteId, int tenantId,
string referenceType, string fullText);
void Delete(long referenceId);
IEnumerable<long> Search(string searchText,
IEnumerable<long> siteIdList,
int tenantId,
int offset,
int pageSize);
}Elasticsearch の実装例(Elastic.Clients.Elasticsearch)
using Elastic.Clients.Elasticsearch;
public class ElasticsearchEngine : IFullTextSearchEngine
{
private readonly ElasticsearchClient _client;
private readonly string _indexName;
public ElasticsearchEngine(string url, string indexName, string apiKey = null)
{
var settings = new ElasticsearchClientSettings(new Uri(url))
.DefaultIndex(indexName);
if (!string.IsNullOrEmpty(apiKey))
settings = settings.Authentication(new ApiKey(apiKey));
_client = new ElasticsearchClient(settings);
_indexName = indexName;
}
public void Index(long referenceId, long siteId, int tenantId,
string referenceType, string fullText)
{
_client.Index(new PleasanterDocument
{
ReferenceId = referenceId,
SiteId = siteId,
TenantId = tenantId,
ReferenceType = referenceType,
FullText = fullText
}, i => i.Index(_indexName).Id(referenceId));
}
public void Delete(long referenceId)
{
_client.Delete<PleasanterDocument>(
referenceId, d => d.Index(_indexName));
}
public IEnumerable<long> Search(string searchText,
IEnumerable<long> siteIdList,
int tenantId,
int offset,
int pageSize)
{
var response = _client.Search<PleasanterDocument>(s => s
.Index(_indexName)
.From(offset)
.Size(pageSize)
.Query(q => q
.Bool(b => b
.Must(m => m
.Match(mm => mm
.Field(f => f.FullText)
.Query(searchText)))
.Filter(
f => f.Term(t => t.TenantId, tenantId),
siteIdList?.Any() == true
? f => f.Terms(t => t
.Field(ff => ff.SiteId)
.Terms(new TermsQueryField(
siteIdList.Select(id => FieldValue.Long(id))
.ToArray())))
: null))));
return response.IsValidResponse
? response.Documents.Select(d => d.ReferenceId)
: Enumerable.Empty<long>();
}
}
internal class PleasanterDocument
{
public long ReferenceId { get; set; }
public long SiteId { get; set; }
public int TenantId { get; set; }
public string ReferenceType { get; set; }
public string FullText { get; set; }
}OpenSearch は同じ構造で OpenSearch.Client を使います(API が NEST とほぼ同じ)。Solr は SolrNet の ISolrOperations<T> で Add・Delete・Query(tenantId・siteId は FilterQueries)を、Manticore は manticoresearch-net の IndexApi.Replace・SearchApi.Search を使う形になります。
CreateFullText() と Get() の変更
private static void CreateFullText(Context context, long id, string fullText)
{
if (fullText != null)
{
// 既存の UPDATE Items はそのまま
if (Parameters.Search.ExternalSearch?.Enabled == true)
{
var engine = SearchEngineFactory.Get();
// siteId・referenceType は Items から別に取る必要がある
engine?.Index(referenceId: id, /* ... */);
}
}
}
public static DataSet Get(Context context, string searchText, /* 既存の引数 */)
{
if (Parameters.Search.ExternalSearch?.Enabled == true)
{
var matchedIds = SearchEngineFactory.Get()?.Search(
searchText: searchText,
siteIdList: siteIdList,
tenantId: context.TenantId,
offset: offset,
pageSize: pageSize);
if (matchedIds?.Any() != true) return null;
// matchedIds を ReferenceId の条件にして CanRead 付きで取得する
return GetByIds(context, matchedIds, dataTableName);
}
// 既存の SQL 全文検索
}CreateFullText() が受け取るのは id と fullText だけなので、外部エンジンに送る siteId・referenceType は Items から取り直すか、引数を増やします。
日本語の解析
外部エンジンに移すと、分かち書きはエンジン側のアナライザーが行います。現行の FullTextClause() が検索語にカタカナ・ひらがなの両形を足している処理は、kuromoji なら readingform フィルターなどで代わりに吸収できます。
| エンジン | 方式 | 現行の検索語の加工 |
|---|---|---|
| Elasticsearch / OpenSearch | kuromoji(形態素解析) | 不要(エンジン側で分割) |
| Apache Solr | ICU など(カスタム設定) | 不要 |
| Meilisearch | Lindera(形態素解析) | 不要 |
| Manticore Search | ngram(bigram) | 不要だが精度は形態素解析に劣る |
| Azure AI Search | ja.microsoft | 不要 |
| Typesense / Algolia | 形態素解析なし | 登録前にアプリ側で分割が要る |
Elasticsearch のインデックス設定の例です。
{
"settings": {
"analysis": {
"analyzer": {
"pleasanter_ja": {
"type": "custom",
"tokenizer": "kuromoji_tokenizer",
"filter": ["kuromoji_baseform", "kuromoji_part_of_speech", "cjk_width", "ja_stop", "lowercase"]
}
}
}
},
"mappings": {
"properties": {
"fullText": { "type": "text", "analyzer": "pleasanter_ja" },
"referenceId": { "type": "long" },
"siteId": { "type": "long" },
"tenantId": { "type": "integer" },
"referenceType": { "type": "keyword" }
}
}
}Manticore Search で日本語を bigram で扱うテーブルの例です。
CREATE TABLE pleasanter (
reference_id BIGINT,
site_id BIGINT,
tenant_id INTEGER,
reference_type STRING,
full_text FIELD
)
charset_table='japanese'
ngram_len='2'
ngram_chars='japanese';権限とページング
外部エンジンは権限を知らないので、権限の判定は必ず DB 側(CanRead)で行います。
図を読み込み中…
外部エンジンで offset・pageSize を当ててから DB で権限を絞る構成になるので、外部エンジンが返した件数と、権限で絞った後の件数が一致しないことがあります。現行の DB 側のページングと整合させる方法は、設計時に決めておく必要があります。
マルチテナント
インデックスには tenantId を持たせ、検索には常に tenantId の条件を付けます。テナントごとにインデックスを分ける(pleasanter_tenant1 など)か、1 つのインデックスを tenantId で絞るかは、テナント数とデータ量で決めます。