v0.5.5 — DeepSeek default + LLM runner cleanup
TL;DR
Significant LLM pipeline upgrade. Default model swapped tencent/hy3-preview → deepseek/deepseek-v4-flash after side-by-side benchmark showed 8× faster, 6× cheaper, qualitatively better extraction. Plus removed two latent bugs in the LLM runner that were causing silent failures.
This is a drop-in upgrade from 0.5.4 — no DB schema change, no env var change, same OPENROUTER_API_KEY.
Upgrade from 0.5.4
npm i -g github:VKirill/TencentDB-Memory-Claude-Code#v0.5.5
# Restart Claude Code (so MCP picks up new binary)Users with explicit model field in their config.json are unaffected — config wins over template. Fresh installs and template re-generation will use DeepSeek by default.
What changed
The headline: model swap
Side-by-side benchmark on identical L1 extraction prompt:
| tencent/hy3-preview (old) | deepseek/deepseek-v4-flash (new) | |
|---|---|---|
| Latency | 61 seconds | 7 seconds ⚡ |
| Reasoning tokens | 10,955 | 818 |
| Cost per call | $0.003 | $0.0005 💰 |
| Facts extracted | 2 of 3 types | 3 of 3 types ✅ |
| Context window | 32K | 1M |
Native response_format |
❌ | ✅ |
Native structured_outputs |
❌ | ✅ |
Hy3-preview is a reasoning model with always-on chain-of-thought. On real L1-extraction batches (5–15K input tokens), it burned 10K+ output tokens on hidden reasoning before producing the actual JSON — frequently hitting finish_reason=length with empty content. DeepSeek v4-flash is fast inference with optional reasoning (opt-in via reasoning_effort parameter), which we don't enable for extraction.
Latent bugs fixed
-
maxTokens=4096cap removed (db9fe85)
The arbitrary 4096 ceiling was inherited from a gpt-3.5 era. For reasoning models it caused silent failures — model would burn the output budget on hidden reasoning before producing visible content. WithmaxTokensundefined, AI SDK omits the parameter and each model uses its own (typically larger) provider-side default. Per-caller and per-config overrides still work. -
Dead
compatibility: "compatible"option removed (175edc9)
The field was dropped fromOpenAIProviderSettingsin@ai-sdk/openai3.x. Silently ignored at runtime but polluted TypeScript diagnostics. In SDK 3.x the chat endpoint is selected viaprovider.chat(model)(which the code already does); no extra flag is needed.
Migration notes
- No DB migration — existing
vectors.dbfiles continue working. - No env change —
OPENROUTER_API_KEYis the same; the model swap is provider-side, both models served via OpenRouter. - Config precedence unchanged — your explicit
modelsetting inconfig.jsonalways wins over the template default. - Per-model overrides — you can still set
llm.model,extraction.model,persona.modelindependently if you want fine-grained control.
What was NOT changed (intentionally)
- Embedding model and dimensions unchanged (
text-embedding-3-large@ 1024-d from 0.5.3). - MCP server tooling unchanged (4 tools, McpServer migration from 0.5.4 stable).
- Scheduler 30-min cadence unchanged.