Releases: zhouronghua/code-agent
Releases · zhouronghua/code-agent
Release list
v0.3.36
feat: visible model switching + fallback recovery probe
Root cause:
The wall of "Let me do it." filler in recent sessions was NOT a prompt change
but a degenerate reasoning loop from hy3 running as the 保底 (fallback) model:
task_1789444306501_2dspe8.json ends with model=hy3 and a single step carrying
98,320 chars of reasoning ("Let me do it." repeated 1617 times) and exactly
32,768 completion tokens = its whole output budget, with an empty answer that
the loop accepted as final. Switching to the 保底 model was also silent, so
there was no way to tell which model produced an answer. Separately, a model
switch only swapped the provider: the context kept the previous model's window
budget, so a 497k-token history was handed to a 256k-window model (observed:
promptTokens 497k -> 113k with cachedTokens 0 = a forced compaction), i.e.
"context breaks after a model switch".
Solution:
- Report every runtime model switch as one [MODEL] line (fallback-timeout /
primary-recovered / routing / profile) and record it in the task log's new
modelSwitches field.
- Supervise the primary model in the background after falling back: probe it
with a cheap healthCheck (30s, backing off to 120s, unref'd timer) and switch
back as soon as it answers; tunable via model_routing.fallback_probe_interval_s
or AgentLoop.setProbeSchedule().
- AgentContext budgets now follow the switched-to model and history is compacted
when it no longer fits; the previous model's reasoning_content is dropped when
moving to a non-thinking model (content/tool_calls preserved).
- healthCheck() must not use max_tokens:1 (GPT-5.x answers 400, which made a
healthy model look unreachable); the reachable/unreachable status policy was
measured against the real gateway.
- Detect degenerate reasoning loops, drop that thinking and retry with a
corrective instruction (max 2 rounds); an empty answer is never accepted as
the final answer. Printed THINKING is truncated to 4k chars.
- Default 保底 model is now gpt-5.6-luna (non-thinking) in config/models templates.
Test:
npm run test:model-switch (63 assertions: loop detection, switch reporting,
context hygiene, fallback + recovery probe, routing resolution, health-check
status policy, loop recovery); E2E with the installed package against a mock
gateway (504 -> [MODEL] fallback -> probe -> [MODEL] switch back -> answer from
the primary model); real-gateway gpt-5.6-luna healthCheck=true.
Impact area:
code-agent CLI (agent loop, context manager, LLM providers, config templates)
Fix status:
done
v0.3.25
feat: benchmark v3 stability + skills triggers in system prompt Benchmark v3 improvements: - Multi-run averaging (BENCHMARK_RUNS=3) with median aggregation to smooth LLM non-determinism - Relaxed thresholds (50-150% tolerance vs old 10-20%) to account for natural LLM variance - BENCHMARK_DISABLE=1 env var to skip checks entirely - BENCHMARK_STRICT=1 to enable blocking mode (default: warn-only, exit 0) - New npm script: benchmark:strict Skills loading: - Include trigger keywords in skill header output so agent can self-match at runtime - buildFullContextPrompt now supports task description for auto-matching - Updated agent prompt template to reference "Triggers" field Version: 0.3.25
v0.3.18
v0.3.17
feat: async tool execution with /btw cancel support (v0.3.17) - Add AbortSignal support to all tools for cancellable execution - AgentLoop tracks current tool via AbortController, exposed via cancelCurrentTool() - /btw cancel and /btw abort commands immediately stop running tools - runTerminal: kills spawned process on abort - poll: checks abort signal before each attempt and during sleep - readFile/writeFile/editFile/listDir/searchText/searchFiles: check signal before execution - Updated help text and tips to mention /btw cancel feature
v0.3.15
v0.3.15: models.json support, config dedup
v0.3.14
v0.3.14: revert MAX_VERIFICATION_ROUNDS to 2
v0.3.13
feat: display full reasoning_content in live output, add thinking par…
v0.3.12
v0.3.12: Auto-execute plan after Plan mode
v0.3.10
v0.3.10: fix CI by using public npm registry in lockfile
v0.2.5
Auto-inject version from package.json via esbuild define - Replace hardcoded AGENT_VERSION constant with __AGENT_VERSION__ placeholder - esbuild define substitutes package.json version at build time - Eliminates manual version sync between package.json and main.ts Made-with: Cursor