Skip to content

v0.3.36

Latest

Choose a tag to compare

@github-actions github-actions released this 16 Sep 00:31
· 15 commits to clean-main since this release
feat: visible model switching + fallback recovery probe

Root cause:
    The wall of "Let me do it." filler in recent sessions was NOT a prompt change
    but a degenerate reasoning loop from hy3 running as the 保底 (fallback) model:
    task_1789444306501_2dspe8.json ends with model=hy3 and a single step carrying
    98,320 chars of reasoning ("Let me do it." repeated 1617 times) and exactly
    32,768 completion tokens = its whole output budget, with an empty answer that
    the loop accepted as final. Switching to the 保底 model was also silent, so
    there was no way to tell which model produced an answer. Separately, a model
    switch only swapped the provider: the context kept the previous model's window
    budget, so a 497k-token history was handed to a 256k-window model (observed:
    promptTokens 497k -> 113k with cachedTokens 0 = a forced compaction), i.e.
    "context breaks after a model switch".

Solution:
    - Report every runtime model switch as one [MODEL] line (fallback-timeout /
      primary-recovered / routing / profile) and record it in the task log's new
      modelSwitches field.
    - Supervise the primary model in the background after falling back: probe it
      with a cheap healthCheck (30s, backing off to 120s, unref'd timer) and switch
      back as soon as it answers; tunable via model_routing.fallback_probe_interval_s
      or AgentLoop.setProbeSchedule().
    - AgentContext budgets now follow the switched-to model and history is compacted
      when it no longer fits; the previous model's reasoning_content is dropped when
      moving to a non-thinking model (content/tool_calls preserved).
    - healthCheck() must not use max_tokens:1 (GPT-5.x answers 400, which made a
      healthy model look unreachable); the reachable/unreachable status policy was
      measured against the real gateway.
    - Detect degenerate reasoning loops, drop that thinking and retry with a
      corrective instruction (max 2 rounds); an empty answer is never accepted as
      the final answer. Printed THINKING is truncated to 4k chars.
    - Default 保底 model is now gpt-5.6-luna (non-thinking) in config/models templates.

Test:
    npm run test:model-switch (63 assertions: loop detection, switch reporting,
    context hygiene, fallback + recovery probe, routing resolution, health-check
    status policy, loop recovery); E2E with the installed package against a mock
    gateway (504 -> [MODEL] fallback -> probe -> [MODEL] switch back -> answer from
    the primary model); real-gateway gpt-5.6-luna healthCheck=true.

Impact area:
    code-agent CLI (agent loop, context manager, LLM providers, config templates)

Fix status:
    done