使用专用的搜索引擎——基于倒排索引(inverted index)的 Elasticsearch/OpenSearch——通过 CDC 从你的源 DB 供给数据,并对 index 分片以扩展规模。你的主数据库(Postgres/MySQL)无法在数亿行上做按相关性排序的文本搜索;LIKE '%term%' 会扫描每一行。搜索引擎把这个问题反转过来。
Source DB ─▶ CDC (Debezium) ─▶ Kafka ─▶ Indexer ─▶ Elasticsearch/OpenSearch
(Postgres) (change stream) ├─ Shard 0 (inverted index) + replicas
├─ Shard 1
Query ─▶ Coordinator ─▶ scatter to shards ─▶ gather ─┴─ Shard N ── ranked by BM25
