Skip to content

Releases: zhouronghua/code-agent

v0.3.36

Choose a tag to compare

@github-actions github-actions released this 16 Sep 00:31
feat: visible model switching + fallback recovery probe

Root cause:
    The wall of "Let me do it." filler in recent sessions was NOT a prompt change
    but a degenerate reasoning loop from hy3 running as the 保底 (fallback) model:
    task_1789444306501_2dspe8.json ends with model=hy3 and a single step carrying
    98,320 chars of reasoning ("Let me do it." repeated 1617 times) and exactly
    32,768 completion tokens = its whole output budget, with an empty answer that
    the loop accepted as final. Switching to the 保底 model was also silent, so
    there was no way to tell which model produced an answer. Separately, a model
    switch only swapped the provider: the context kept the previous model's window
    budget, so a 497k-token history was handed to a 256k-window model (observed:
    promptTokens 497k -> 113k with cachedTokens 0 = a forced compaction), i.e.
    "context breaks after a model switch".

Solution:
    - Report every runtime model switch as one [MODEL] line (fallback-timeout /
      primary-recovered / routing / profile) and record it in the task log's new
      modelSwitches field.
    - Supervise the primary model in the background after falling back: probe it
      with a cheap healthCheck (30s, backing off to 120s, unref'd timer) and switch
      back as soon as it answers; tunable via model_routing.fallback_probe_interval_s
      or AgentLoop.setProbeSchedule().
    - AgentContext budgets now follow the switched-to model and history is compacted
      when it no longer fits; the previous model's reasoning_content is dropped when
      moving to a non-thinking model (content/tool_calls preserved).
    - healthCheck() must not use max_tokens:1 (GPT-5.x answers 400, which made a
      healthy model look unreachable); the reachable/unreachable status policy was
      measured against the real gateway.
    - Detect degenerate reasoning loops, drop that thinking and retry with a
      corrective instruction (max 2 rounds); an empty answer is never accepted as
      the final answer. Printed THINKING is truncated to 4k chars.
    - Default 保底 model is now gpt-5.6-luna (non-thinking) in config/models templates.

Test:
    npm run test:model-switch (63 assertions: loop detection, switch reporting,
    context hygiene, fallback + recovery probe, routing resolution, health-check
    status policy, loop recovery); E2E with the installed package against a mock
    gateway (504 -> [MODEL] fallback -> probe -> [MODEL] switch back -> answer from
    the primary model); real-gateway gpt-5.6-luna healthCheck=true.

Impact area:
    code-agent CLI (agent loop, context manager, LLM providers, config templates)

Fix status:
    done

v0.3.25

Choose a tag to compare

@github-actions github-actions released this 11 Aug 07:04
feat: benchmark v3 stability + skills triggers in system prompt

Benchmark v3 improvements:
- Multi-run averaging (BENCHMARK_RUNS=3) with median aggregation to smooth LLM non-determinism
- Relaxed thresholds (50-150% tolerance vs old 10-20%) to account for natural LLM variance
- BENCHMARK_DISABLE=1 env var to skip checks entirely
- BENCHMARK_STRICT=1 to enable blocking mode (default: warn-only, exit 0)
- New npm script: benchmark:strict

Skills loading:
- Include trigger keywords in skill header output so agent can self-match at runtime
- buildFullContextPrompt now supports task description for auto-matching
- Updated agent prompt template to reference "Triggers" field

Version: 0.3.25

v0.3.18

Choose a tag to compare

@github-actions github-actions released this 16 Jul 01:33
v0.3.18

v0.3.17

Choose a tag to compare

@github-actions github-actions released this 10 Jul 09:15
feat: async tool execution with /btw cancel support (v0.3.17)

- Add AbortSignal support to all tools for cancellable execution
- AgentLoop tracks current tool via AbortController, exposed via cancelCurrentTool()
- /btw cancel and /btw abort commands immediately stop running tools
- runTerminal: kills spawned process on abort
- poll: checks abort signal before each attempt and during sleep
- readFile/writeFile/editFile/listDir/searchText/searchFiles: check signal before execution
- Updated help text and tips to mention /btw cancel feature

v0.3.15

Choose a tag to compare

@github-actions github-actions released this 16 Jul 01:33
v0.3.15: models.json support, config dedup

v0.3.14

Choose a tag to compare

@github-actions github-actions released this 16 Jul 01:33
v0.3.14: revert MAX_VERIFICATION_ROUNDS to 2

v0.3.13

Choose a tag to compare

@github-actions github-actions released this 08 Jul 09:31
feat: display full reasoning_content in live output, add thinking par…

v0.3.12

Choose a tag to compare

@github-actions github-actions released this 03 Jul 01:14
v0.3.12: Auto-execute plan after Plan mode

v0.3.10

Choose a tag to compare

@github-actions github-actions released this 02 Jul 02:59
v0.3.10: fix CI by using public npm registry in lockfile

v0.2.5

Choose a tag to compare

@github-actions github-actions released this 24 Apr 23:54
Auto-inject version from package.json via esbuild define

- Replace hardcoded AGENT_VERSION constant with __AGENT_VERSION__ placeholder
- esbuild define substitutes package.json version at build time
- Eliminates manual version sync between package.json and main.ts

Made-with: Cursor