Replies: 1 comment
|
数据点 + 两条实现约束,来自一个独立实现(trigram 路线,不是扩展加载路线) I maintain a separate fork implementing CJK search on the same seam, and I reproduced the core of this proposal before reading it. Two things that may help scope the change: 1. Independent confirmation of the premise. Measured on
Your 2. Constraint worth putting in the proposal: the tokenizer choice alone does not fix short queries.
3. A no-extension alternative worth having in the same discussion. My fork avoids Source with the measurements and the test suite: https://github.com/QIANLING-0831/dsh-memory-plus (the Also relevant: @xzy-jason's jieba-based approach in #1456 reports 94% precision vs 45% for trigram+LIKE on 1,000 Chinese entries. If the team is weighing a tokenizer seam, that thread's data plus yours covers both the extension route and the pure-JS route.
|
Uh oh!
There was an error while loading. Please reload this page.
Feature proposal: configurable FTS5 tokenizer for session search (Chinese support)
中文摘要:
session-query-sqlite的全文索引目前硬编码tokenize = 'unicode61'(SQLite 默认分词器)。unicode61不做中文分词,一段连续中文(含中英混排)会被当成一个超长 token,导致中文 session 搜索几乎无法命中。建议给该包增加tokenizer/tokenizerExtension两个配置项,允许加载第三方分词扩展(如 wangfenjin/simple,SQLite FTS5 中文+拼音分词,MIT/GPLv3 双许可),并顺手把索引 schema 版本升到 9 以在切换分词器时自动重建索引。Problem
With the default
unicode61tokenizer, a document like查询OpenRouter免费视觉模型becomes one index token (查询openrouter免费视觉模型). Querying视觉模型orOpenRoutermatches nothing:视觉模型免费OpenRouter(inside CJK run)Proposed change
@deepseek-ai/dsh-session-query-sqlitegains two optional config keys (default behavior unchanged):Implementation notes:
tokenizeris restricted to[A-Za-z0-9_]+(validated in the schemastery schema, matching howopenAtis validated) so configuration can never interpolate SQL.tokenizerExtensionis set, the engine opensnode:sqlitewithallowExtension, callsenableLoadExtension+loadExtension, then creates the FTS5 tables with the configured tokenizer; otherwise everything stays exactly as today (unicode61, no extension loading).SESSION_QUERY_SQLITE_SCHEMA_VERSIONbumps 8 → 9 so existing persistent derived indexes rebuild with the configured tokenizer.tokenizerwithouttokenizerExtensionis rejected at config resolution.Validation
simpleextension and searches Chinese phrases (skipped in CI when the extension is absent).pnpm run typecheck, oxlint,verify-config-catalogall pass;docs/config-catalog.mdregenerated.Ready-to-merge branch
The full change is committed on
feat/session-query-tokenizerin my fork: https://github.com/corntrace/deepseek-harness/tree/feat/session-query-tokenizerI understand external PRs are not accepted at the moment (per CONTRIBUTING.md) — posting this as an Idea per the contributing guidance. If the team is interested, the branch is one click away from a PR; happy to adjust scope, add i18n docs, or split the change.
All reactions