Repository navigation
v1.3.101-pre-2
Pre-release
Pre-release
·
54 commits
to main
since this release
Changelog (Current Version)
This file contains the changes for the current version only. The full history of all versions lives in docs/en/changelog.md.
v1.3.101
✨ New Features
-
Serve: Placeholder API Token Warnings
mothx serve init-configships a well-known template token, which is equivalent to publishing the API key once auth is enabled — and nothing said so. The token is now a single named constant (serve.PlaceholderAuthToken) with explicit detection (IsPlaceholderAuthToken/UsesPlaceholderAuthToken), and a replacement warning is printed wherever the template is created (mothx serve init-configand the--init-serveCLI path) and again onmothx servestartup whileapi.auth.enabledis true and the token is still in place.- The startup check only fires for an enabled auth block, so the default template (auth disabled, loopback only) stays quiet and a replaced token is never flagged.
serve init-confignow writes through the command's own stderr, so the warning travels with the command output.
-
Members Wait for Their Lead on Interactive Surfaces
- A session that can spawn members but is not bound to an expert team now keeps its run open for still-running members on interactive surfaces (TUI, Web UI, ACP), so a member's completion or question is handled within the same run instead of waiting for the next lead run. Headless and asynchronous sources (CLI, cron, WeChat/Feishu) still end the turn normally and deliver member notifications at their next iteration or next run; a bound expert team always waits.
-
Bound Teams Keep the Full Sub-Agent Tool Set
- A bound expert team always exposes the complete canonical sub-agent tool set (
subagent_spawn,subagent_status,subagent_send,subagent_wait,subagent_answer,subagent_destroy). Per-tool toggles that disable individual sub-agent tools apply only to non-team multi-agent sessions; the team capability is authoritative and never drops tools.
- A bound expert team always exposes the complete canonical sub-agent tool set (
-
New Gitee/Moark Model:
deepseek-v4.1-flash- Added
deepseek-v4.1-flashto thegiteeandmoarkproviders with a 1M context window and text+image input; no default max_tokens is sent.
- Added
🐛 Bug Fixes
-
Content-Inspected Images No Longer Kill the Session
- A provider content-policy refusal — for example DashScope/Qwen's
InternalError.Algo.DataInspectionFailed: Input image data may contain inappropriate content— arrives as an HTTP 400, but every 4xx was treated as retryable. The same rejected image was re-sent through the provider's backoff retries plus the agent's stream-failure retries (minutes of "Retrying…"), and because a refusal is permanent the run finally failed; worse, the offending image stayed in the persisted history and was replayed on every later turn, so nothing could continue. Only starting a brand-new session recovered — not even/clear, which reloads the same history. provider.IsContentRejectionErrornow classifies this narrow family (data inspection, content policy/moderation/filter, "inappropriate content") andIsRetryablereturns false for it, so the failure surfaces immediately instead of burning the retry budget. Agent Core then recovers in place: it strips the refused images — first from the current turn, then, if the refusal persists, from the whole conversation — replacing each with a model-visible note that explains the provider's content filter blocked the image and that the pixels are unavailable, and records an append-onlycontent_overridesession entry so replay (same process or after reload) never re-sends it. The run retries without the image and the session keeps working; a turn that already streamed output is healed but not re-run, so no output is duplicated.
- A provider content-policy refusal — for example DashScope/Qwen's
-
Serve: Config Saves No Longer Fail with "Access is denied" on Windows
- On Windows (including portable setups running from exFAT drives), enabling the WeChat/Feishu channel from the Web UI or saving any
serve.jsonchange failed withsync config directory: Access is denied. The atomic config writer fsynced the parent directory after the rename — a POSIX durability idiom — butFlushFileBufferson a read-only directory handle always returnsERROR_ACCESS_DENIEDon Windows, on every filesystem including exFAT. Because the failure happened after the file had already been swapped into place, the API reported an error while the new config was never applied to the runtime. - The post-rename directory flush is now skipped on Windows (matching etcd/bolt practice); the config file itself is still fsynced before the rename, so durability is preserved and Unix behavior is unchanged.
- On Windows (including portable setups running from exFAT drives), enabling the WeChat/Feishu channel from the Web UI or saving any
-
MCP: Image Tool Results Reach the Model Instead of a Placeholder
- An MCP tool that returned image content reached the model as the literal string
[image content: image/png]. The base64 payload was already present in the MCP response, but the client decoded every content block into text only, and the tool returned a text-only result, so a screenshot-style MCP server could report coordinates while the model never saw the picture.resources/readbinary resources were worse: theblobfield had no matching struct field, so the payload was dropped during decode and not even the placeholder appeared. tools/callandresources/readnow project image blocks intotools.ToolResult.Contentsas real provider image content, reusing the same provider-aware preprocessing as thereadandbrowserscreenshot tools, and decodeblob/uriresource fields. Results without images keep the historical text-only shape, so existing text MCP tools are unchanged. Malformed, oversized, or excess images (capped at 4 per result, matching the ACP projection limit) degrade to a text note instead of failing the call. The image capability gate in Agent Core still decides whether a non-vision model receives images at all.
- An MCP tool that returned image content reached the model as the literal string
-
SQLite: Transient Busy Transaction Begins Are Retried
- Several processes opening one session directory race onto the single writer lock. The DSN begins non-read-only transactions with
BEGIN IMMEDIATE, so a begin can outlast the connection'sbusy_timeoutwhile other processes keep committing undersynchronous(FULL), and a healthy database failed withdatabase is locked (5). - The retry policy now lives in
internal/dbnext to the DSN ownership:BeginTx(Bun),BeginSQLTx(raw*sql.DB), andRunInTxretry onlySQLITE_BUSY/SQLITE_LOCKEDinside a bounded budget (90 s) with exponential backoff (200 ms up to a 2 s cap); non-transient errors are returned unchanged and a caller's context deadline still wins. - DAO
Begin/BeginTx/RunInTx(including the bindings helpers),internal/db.Write, and the session schema initialization/migration boundary all use it, so concurrent startup and ordinary writes no longer turn a transient writer conflict into a hard failure.
- Several processes opening one session directory race onto the single writer lock. The DSN begins non-read-only transactions with
-
Workflow: Runaway DSL Scripts Are Bounded by a Wall-Clock Budget
- A workflow source such as
while (true) {}could pin the process forever whenever its caller passed a context without a deadline. Source evaluation — which only builds the node graph, since worker agents run natively afterwards — now runs under two bounds: the caller's context and a wall-clock budget, whichever fires first interrupts the VM. - The budgets are 30 s for a workflow run and 5 s for the interactive
workflow_lintauthoring check, which must fail fast; a timeout surfaces as the sentinelErrJSEvaluationTimeout(the lint result carries a stable, readable error), while a caller cancellation keeps returning its context error.Runner.EvalTimeoutlets callers tighten the budget, and the zero value keeps the documented defaults.
- A workflow source such as
-
Cancelled or Expired Decisions No Longer Block Forking
- A session whose only decision had actually been cancelled or timed out was still treated as having a pending decision, so forking it was rejected as
source session is active. Every decision-ledger reader now shares one vocabulary, so cancelled and timed-out decisions (and the legacy channel request name) clear correctly and the fork proceeds. The durable decision event name and its{"decision": …}envelope also gained a single owner each, so cross-entry decision recovery reads the same records regardless of which surface wrote them.
- A session whose only decision had actually been cancelled or timed out was still treated as having a pending decision, so forking it was rejected as
-
Mid-Stream Network Failures Retry Automatically Instead of Ending the Reply
- A provider stream that died with a transient transport error (
connection reset by peer, unexpected EOF, gateway 5xx, ...) after text or thinking had already been streamed failed the whole run withstream read error: ...: provider-level retries only cover streams that break before any visible output, and the agent-level retry only covered idle-stream timeouts. - The agent loop now performs a bounded continuation retry (up to 2 attempts) for such transient errors. Already-streamed partial output is persisted into history and a continuation instruction quoting the exact suffix is injected, so the model resumes from the interruption point instead of duplicating what the user already saw; with no visible output yet, the turn simply re-runs. Turns with an already-emitted tool call, context overflow (dedicated compaction recovery), and idle-stream timeouts (dedicated timeout retry) keep their existing behavior, and Responses remote-state turns keep their existing failover path.
- A provider stream that died with a transient transport error (
🔧 Improvements
- SQLite: Three-Phase Write-Pressure Reduction for the Session Database
- The connection durability policy moves from
synchronous(FULL)to the WAL-recommendedsynchronous(NORMAL): a commit no longer fsyncs while holding the single writer lock (the fsync moves to checkpoint time), so writer-lock occupancy across processes sharing one session directory shrinks from fsync scale to page-cache scale, largely eliminating the recorded "another process keeps committing until begin exceeds busy_timeout and reports database is locked" scenario. Process crashes still lose nothing; an OS crash or power loss can roll back the seconds of commits since the last checkpoint (the database stays consistent, and a missing run terminal state converges through the existing lease-expiry → orphan → bounded-recovery path).MOTHX_SQLITE_SYNCHRONOUS=FULLrestores the legacy durability per process, and mixed old/new processes sharing one database file is safe. - Tool results now persist in batches: the session domain gained
AppendMessages, writing one agent iteration's tool results as a parent-chained single transaction (capped at 64 entries per transaction, chunked above that) with the lease fence and the leaf-conflict check still inside the write transaction; tool-heavy rounds drop from N+3 write transactions to about 3. The assistant message still persists before tool side effects, and a failed batch still fails the run withsession_save. - Lease heartbeats are coalesced: instead of one goroutine committing a renewal transaction per active lease every 3 seconds, a single scheduler per session directory renews all of this process's leases for that database in one transaction (the per-lease owner/epoch/token CAS fence is unchanged, so a displaced or released lease still only loses itself); steady-state heartbeat writes drop from N transactions/3s to 1 transaction/3s/process. TTL, heartbeat interval, retry budget, and the 30-second bounded-recovery guarantee are untouched, and the scheduler retires after the last lease is released.
internal/dbgained process-wide busy-retry and transaction-begin wait metrics (BusyRetryStats/BeginWaitStats), published through expvar asmothx_sqliteand readable on the--debugpprof server's/debug/vars, making cross-process writer contention observable.- Full plan, multi-process reasoning, and load-test matrix:
docs/proposal/sqlite-write-pressure-reduction-proposal.md.
- The connection durability policy moves from
✅ Tests
- Database:
internal/dbpins the begin-retry policy — only SQLITE_BUSY/SQLITE_LOCKED are retried, other driver errors and an expiring context are surfaced unchanged, driver codes are classified through the error's ownCode()method, andRunInTxkeeps commit/rollback semantics. - Workflow: a runaway script is interrupted by a 50 ms budget (previously that test hung), a successful evaluation behaves exactly as before, and the lint path reports an invalid result with the timeout message.
- Serve: the generated template keeps the same constant the detection uses, a whitespace-padded placeholder is still recognized while a real or empty token is not, and the
serve init-configoutput must contain the warning. - Agent loop: ten real
bash echocalls run through one parallel batch (the rendezvous only closes once all ten workers are live), and ordered-start coverage pins the launch semantics — starts in the model's declared order, an in-flight call unaffected by an earlier approval wait, queued calls released when an earlier call fails, and the same ordered handle for background tool calls. - Runtime: the cross-process takeover test now retries the expire-then-takeover pair inside a bounded budget and gives helper startup a load-tolerant window, so a loaded machine fails setup with a clear diagnosis instead of failing the invariant under test.
- Decision ledger: round-trip and replay coverage for the shared event name, envelope, and loaders; the legacy channel event name still decodes; and a cancelled decision no longer blocks a fork.
- Member wait: an interactive (TUI) non-team lead holds its run open for a running member, while a headless (CLI) one ends normally.
- Architecture: a guard rejects new use of the legacy session run/lease APIs or a low-level
agent.Newin adapter test files, with a documented allowlist for the remaining fixtures. systeminit: the shared/systeminitprompt is pinned — interactive-only question guidance, trimmed extra instructions before the final note, blank extra input ignored, and determinism.- Stream-failure recovery: a mid-stream connection reset resumes through the continuation retry — partial output persists and continues from the exact suffix, no visible output re-runs the turn fresh, and exhausting the budget surfaces the original error — while a turn with an already-emitted tool call or a non-retryable error never retries.
- SQLite write pressure:
internal/dbpins the synchronous default (NORMAL), theMOTHX_SQLITE_SYNCHRONOUS=FULLoverride, and busy-retry counting that never charges permanent errors; session tests cover theAppendMessagesparent chain and replay order, whole-batch stale-writer rejection persisting no rows, chunking above the transaction cap, and sub-agent table isolation; the heartbeat scheduler test covers one scheduler per directory, batched renewal, a displaced lease losing only itself while the survivor renews, and retirement after the last release. Plus write-pressure load shapes A/B/C (multi-process distinct-session writers, mixed FULL/NORMAL deployment, single-process many sessions with leases) reporting busy/begin contention metrics, scalable throughMOTHX_WRITE_PRESSURE_SCALEfor baseline runs.