Skip to content

検索を外部の検索エンジンに外出しする ​

第1版作成 最終更新 (日本時間)
確認バージョン1.5.1.01.5.8.1

本体の標準機能ではありません

プリザンター本体の全文検索は RDBMS の機能だけで動いています。このページは外部の検索エンジンに移す場合の設計メモです。本体を改修せずに外部の検索エンジンを使う例は Fess で全文検索にあります。

前提にした現行実装 ​

詳しくは検索機能の内部実装にまとめています。改修に関係する点だけ挙げます。

  • 各レコードの内容をつないだ文字列を Items.FullText に入れ、検索はこの列を見る。書き込みは Indexes.CreateFullText()(Indexes.cs#L131-L145)で、FullText と SearchIndexCreatedTime だけを更新する。
  • バックグラウンドの Indexes.RebuildSearchIndexes() が、SearchIndexCreatedTime が古いレコードを作り直す(Indexes.cs#L1048)。
  • 横断検索は Indexes.Get()(Indexes.cs#L748-L760)が ISqlCommandText の RDBMS ごとの実装で SQL を作る(SQL Server は CONTAINS、PostgreSQL は pg_trgm の %>、MySQL は MATCH ... AGAINST)。権限は同じ SQL の中で Def.Sql.CanRead を条件に入れて絞る。
  • 一覧の検索はサイトの検索方式(FullText・PartialMatch・MatchInFrontOfTitle・BroadMatchOfTitle)に従い、View.SetSearchWhere() が別に条件を作る。Indexes.Get() を通らない。
  • 検索語は SearchIndexes() → Words() → FullTextClause() で分割・かな変換される。Libraries/Search/WordBreaker.cs という文字種で分割するクラスもあるが、1.5.8.1 ではどこからも参照されていない。
  • 添付ファイルの中身の検索(SearchDocuments)は Binaries.Bin を RDBMS の全文検索で直接見るもので、実用になるのは SQL Server だけ。

候補のエンジン ​

エンジンの特徴は各プロジェクトの公開情報に基づきます(調査時点)。

エンジン.NET クライアントライセンスセルフホスト分散構成日本語解析
Elasticsearch 8Elastic.Clients.Elasticsearch(公式。旧 NEST は 7.x 用で EOL)SSPL / Elastic License 2.0(要確認)可可kuromoji(analysis-kuromoji)・ICU
OpenSearch 2OpenSearch.Client(公式。NEST のフォーク)Apache 2.0可可(AWS は Amazon OpenSearch Service)kuromoji
Apache Solr 9SolrNet(コミュニティ、1.2.1。9.x 対応は一部ベータ)Apache 2.0可(Java 21 以上)SolrCloudICU・カスタム設定。kuromoji は既定で同梱されない
MeilisearchMeilisearch(公式)MIT可(Windows は Server 2022 以降が公式対応)不可(単一ノード)Lindera(形態素解析、内蔵)
TypesenseTypesense(コミュニティ、8.1.0)GPL-3.0(要確認)可(Windows は Docker 経由のみ)全量複製の冗長構成形態素解析なし(事前に分割が要る)
Manticore Searchmanticoresearch-net(公式)。MySQL プロトコル互換なので MySQL 用クライアントでも可GPL-2.0(要確認)可(Windows インストーラあり)Galerangram(bigram)
Azure AI SearchAzure.Search.Documents(公式、11.7.0)Azure のサービス不可マネージドja.microsoft
AlgoliaAlgolia.Search(公式、7.38.x)SaaS不可マネージド形態素解析なし(事前に分割が要る)

いずれも net10.0(1.5.8.1 のターゲット)から使えます。選び方の目安は次のとおりです。

重視すること候補
日本語の検索精度Elasticsearch / OpenSearch(kuromoji)
導入の手軽さMeilisearch
ライセンスの制約がないことOpenSearch / Apache Solr
少ない資源で動かすManticore Search
Azure 前提Azure AI Search
AWS 前提Amazon OpenSearch Service
インフラを持たないAlgolia(データをクラウドに送ること、量が増えると費用がかさむことに注意)

改修の構成 ​

RDB の Items.FullText への書き込みはそのまま残し、外部エンジンにも同じ内容を送ります。検索は外部エンジンから一致した ReferenceId の一覧を受け取り、権限の絞り込みは DB で行います。

図を読み込み中…

Parameters.Search.ExternalSearch.Enabled のようなフラグで切り替え、無効なら現行の SQL 検索をそのまま使う形にすると、外部エンジンを使わない環境に影響しません。

改修するファイル ​

ファイル内容
Implem.ParameterAccessor/Parts/Search.cs・App_Data/Parameters/Search.json接続情報の追加(下記)
新規 Libraries/Search/IFullTextSearchEngine.csエンジンを差し替えるためのインターフェース
新規 Libraries/Search/ElasticsearchEngine.cs などエンジンごとの実装
Libraries/Search/Indexes.cs の CreateFullText()DB 更新の後に外部エンジンへ送る
Libraries/Search/Indexes.cs の Get()外部エンジンから ID を取って DB で絞る
Libraries/Search/Indexes.cs の RebuildSearchIndexes()外部エンジンへの一括再送
レコードの削除処理(IssueModel・ResultModel・WikiModel の Delete() の呼び出し元)外部エンジンからも削除

一覧の検索(View.SetSearchWhere())も外部エンジンに向けるなら、そちらも別に改修が要ります。

パラメータ ​

1.5.8.1 の Search クラスの項目は SearchDocuments・CreateIndexes・PageSize・DisableCrossSearch・DisableCrossSearchSites・FullTextIncludeBreadcrumb・FullTextIncludeSiteId・FullTextIncludeSiteTitle・FullTextNumberOfMails・FullTextMaxNumberOfMails です(Search.cs)。ここに外部エンジンの設定を足します。

csharp
public class Search
{
    // 既存の項目はそのまま
    public ExternalSearch ExternalSearch;
}

public class ExternalSearch
{
    public bool Enabled;                  // 外部エンジンを使うか
    public string Engine;                 // "Elasticsearch" | "OpenSearch" | "Solr" | "Meilisearch" など
    public string Url;                    // 例: "http://localhost:9200"
    public string IndexName;              // 例: "pleasanter"
    public string ApiKey;                 // 省略可
    public string CertificateFingerprint; // 省略可
}

インターフェース ​

csharp
public interface IFullTextSearchEngine
{
    void Index(long referenceId, long siteId, int tenantId,
               string referenceType, string fullText);
    void Delete(long referenceId);
    IEnumerable<long> Search(string searchText,
                             IEnumerable<long> siteIdList,
                             int tenantId,
                             int offset,
                             int pageSize);
}
Elasticsearch の実装例(Elastic.Clients.Elasticsearch)
csharp
using Elastic.Clients.Elasticsearch;

public class ElasticsearchEngine : IFullTextSearchEngine
{
    private readonly ElasticsearchClient _client;
    private readonly string _indexName;

    public ElasticsearchEngine(string url, string indexName, string apiKey = null)
    {
        var settings = new ElasticsearchClientSettings(new Uri(url))
            .DefaultIndex(indexName);
        if (!string.IsNullOrEmpty(apiKey))
            settings = settings.Authentication(new ApiKey(apiKey));
        _client = new ElasticsearchClient(settings);
        _indexName = indexName;
    }

    public void Index(long referenceId, long siteId, int tenantId,
                      string referenceType, string fullText)
    {
        _client.Index(new PleasanterDocument
        {
            ReferenceId = referenceId,
            SiteId = siteId,
            TenantId = tenantId,
            ReferenceType = referenceType,
            FullText = fullText
        }, i => i.Index(_indexName).Id(referenceId));
    }

    public void Delete(long referenceId)
    {
        _client.Delete<PleasanterDocument>(
            referenceId, d => d.Index(_indexName));
    }

    public IEnumerable<long> Search(string searchText,
                                    IEnumerable<long> siteIdList,
                                    int tenantId,
                                    int offset,
                                    int pageSize)
    {
        var response = _client.Search<PleasanterDocument>(s => s
            .Index(_indexName)
            .From(offset)
            .Size(pageSize)
            .Query(q => q
                .Bool(b => b
                    .Must(m => m
                        .Match(mm => mm
                            .Field(f => f.FullText)
                            .Query(searchText)))
                    .Filter(
                        f => f.Term(t => t.TenantId, tenantId),
                        siteIdList?.Any() == true
                            ? f => f.Terms(t => t
                                .Field(ff => ff.SiteId)
                                .Terms(new TermsQueryField(
                                    siteIdList.Select(id => FieldValue.Long(id))
                                    .ToArray())))
                            : null))));
        return response.IsValidResponse
            ? response.Documents.Select(d => d.ReferenceId)
            : Enumerable.Empty<long>();
    }
}

internal class PleasanterDocument
{
    public long ReferenceId { get; set; }
    public long SiteId { get; set; }
    public int TenantId { get; set; }
    public string ReferenceType { get; set; }
    public string FullText { get; set; }
}

OpenSearch は同じ構造で OpenSearch.Client を使います(API が NEST とほぼ同じ)。Solr は SolrNet の ISolrOperations<T> で Add・Delete・Query(tenantId・siteId は FilterQueries)を、Manticore は manticoresearch-net の IndexApi.Replace・SearchApi.Search を使う形になります。

CreateFullText() と Get() の変更 ​

csharp
private static void CreateFullText(Context context, long id, string fullText)
{
    if (fullText != null)
    {
        // 既存の UPDATE Items はそのまま
        if (Parameters.Search.ExternalSearch?.Enabled == true)
        {
            var engine = SearchEngineFactory.Get();
            // siteId・referenceType は Items から別に取る必要がある
            engine?.Index(referenceId: id, /* ... */);
        }
    }
}

public static DataSet Get(Context context, string searchText, /* 既存の引数 */)
{
    if (Parameters.Search.ExternalSearch?.Enabled == true)
    {
        var matchedIds = SearchEngineFactory.Get()?.Search(
            searchText: searchText,
            siteIdList: siteIdList,
            tenantId: context.TenantId,
            offset: offset,
            pageSize: pageSize);
        if (matchedIds?.Any() != true) return null;
        // matchedIds を ReferenceId の条件にして CanRead 付きで取得する
        return GetByIds(context, matchedIds, dataTableName);
    }
    // 既存の SQL 全文検索
}

CreateFullText() が受け取るのは id と fullText だけなので、外部エンジンに送る siteId・referenceType は Items から取り直すか、引数を増やします。

日本語の解析 ​

外部エンジンに移すと、分かち書きはエンジン側のアナライザーが行います。現行の FullTextClause() が検索語にカタカナ・ひらがなの両形を足している処理は、kuromoji なら readingform フィルターなどで代わりに吸収できます。

エンジン方式現行の検索語の加工
Elasticsearch / OpenSearchkuromoji(形態素解析)不要(エンジン側で分割)
Apache SolrICU など(カスタム設定)不要
MeilisearchLindera(形態素解析)不要
Manticore Searchngram(bigram)不要だが精度は形態素解析に劣る
Azure AI Searchja.microsoft不要
Typesense / Algolia形態素解析なし登録前にアプリ側で分割が要る

Elasticsearch のインデックス設定の例です。

json
{
    "settings": {
        "analysis": {
            "analyzer": {
                "pleasanter_ja": {
                    "type": "custom",
                    "tokenizer": "kuromoji_tokenizer",
                    "filter": ["kuromoji_baseform", "kuromoji_part_of_speech", "cjk_width", "ja_stop", "lowercase"]
                }
            }
        }
    },
    "mappings": {
        "properties": {
            "fullText": { "type": "text", "analyzer": "pleasanter_ja" },
            "referenceId": { "type": "long" },
            "siteId": { "type": "long" },
            "tenantId": { "type": "integer" },
            "referenceType": { "type": "keyword" }
        }
    }
}

Manticore Search で日本語を bigram で扱うテーブルの例です。

sql
CREATE TABLE pleasanter (
    reference_id BIGINT,
    site_id BIGINT,
    tenant_id INTEGER,
    reference_type STRING,
    full_text FIELD
)
charset_table='japanese'
ngram_len='2'
ngram_chars='japanese';

権限とページング ​

外部エンジンは権限を知らないので、権限の判定は必ず DB 側(CanRead)で行います。

図を読み込み中…

外部エンジンで offset・pageSize を当ててから DB で権限を絞る構成になるので、外部エンジンが返した件数と、権限で絞った後の件数が一致しないことがあります。現行の DB 側のページングと整合させる方法は、設計時に決めておく必要があります。

マルチテナント ​

インデックスには tenantId を持たせ、検索には常に tenantId の条件を付けます。テナントごとにインデックスを分ける(pleasanter_tenant1 など)か、1 つのインデックスを tenantId で絞るかは、テナント数とデータ量で決めます。

関連ページ ​

変更履歴

第1版レートリミッターの解説と、フォーム投稿の制限・ウイルススキャン・拡張子制限・外部検索エンジン・管理画面設定の改修・設計メモを追加