Skip to content
This repository was archived by the owner on Jun 23, 2026. It is now read-only.

v0.5.5 — DeepSeek default + LLM runner cleanup

Choose a tag to compare

@VKirill VKirill released this 17 May 19:19
· 5 commits to main since this release

TL;DR

Significant LLM pipeline upgrade. Default model swapped tencent/hy3-preview → deepseek/deepseek-v4-flash after side-by-side benchmark showed 8× faster, 6× cheaper, qualitatively better extraction. Plus removed two latent bugs in the LLM runner that were causing silent failures.

This is a drop-in upgrade from 0.5.4 — no DB schema change, no env var change, same OPENROUTER_API_KEY.

Upgrade from 0.5.4

npm i -g github:VKirill/TencentDB-Memory-Claude-Code#v0.5.5
# Restart Claude Code (so MCP picks up new binary)

Users with explicit model field in their config.json are unaffected — config wins over template. Fresh installs and template re-generation will use DeepSeek by default.

What changed

The headline: model swap

Side-by-side benchmark on identical L1 extraction prompt:

tencent/hy3-preview (old) deepseek/deepseek-v4-flash (new)
Latency 61 seconds 7 seconds ⚡
Reasoning tokens 10,955 818
Cost per call $0.003 $0.0005 💰
Facts extracted 2 of 3 types 3 of 3 types ✅
Context window 32K 1M
Native response_format ❌ ✅
Native structured_outputs ❌ ✅

Hy3-preview is a reasoning model with always-on chain-of-thought. On real L1-extraction batches (5–15K input tokens), it burned 10K+ output tokens on hidden reasoning before producing the actual JSON — frequently hitting finish_reason=length with empty content. DeepSeek v4-flash is fast inference with optional reasoning (opt-in via reasoning_effort parameter), which we don't enable for extraction.

Latent bugs fixed

  1. maxTokens=4096 cap removed (db9fe85)
    The arbitrary 4096 ceiling was inherited from a gpt-3.5 era. For reasoning models it caused silent failures — model would burn the output budget on hidden reasoning before producing visible content. With maxTokens undefined, AI SDK omits the parameter and each model uses its own (typically larger) provider-side default. Per-caller and per-config overrides still work.

  2. Dead compatibility: "compatible" option removed (175edc9)
    The field was dropped from OpenAIProviderSettings in @ai-sdk/openai 3.x. Silently ignored at runtime but polluted TypeScript diagnostics. In SDK 3.x the chat endpoint is selected via provider.chat(model) (which the code already does); no extra flag is needed.

Migration notes

  • No DB migration — existing vectors.db files continue working.
  • No env change — OPENROUTER_API_KEY is the same; the model swap is provider-side, both models served via OpenRouter.
  • Config precedence unchanged — your explicit model setting in config.json always wins over the template default.
  • Per-model overrides — you can still set llm.model, extraction.model, persona.model independently if you want fine-grained control.

What was NOT changed (intentionally)

  • Embedding model and dimensions unchanged (text-embedding-3-large @ 1024-d from 0.5.3).
  • MCP server tooling unchanged (4 tools, McpServer migration from 0.5.4 stable).
  • Scheduler 30-min cadence unchanged.