Skip to content

[0.7.14] — 2026-09-19

Latest

Choose a tag to compare

@brijesh1100 brijesh1100 released this 19 Sep 17:06

[0.7.14] — 2026-09-19

Hardens Autonomous mode so runs terminate for a named reason, do not loop on compaction/recall, enforce response_format, and give the Critic the same evidence the generator saw. Ships first-class WebSearchTool in core. Two provider patches for bugs that only those SDKs hit.

Coordinated package versions:

  • nucleusiq 0.7.14
  • nucleusiq-ollama 0.2.2 — tool-call arguments as a mapping on the official ollama SDK
  • nucleusiq-anthropic 0.2.2 — additionalProperties: false on every object in structured output

Unchanged: nucleusiq-openai 0.7.1, nucleusiq-openai-compatible 0.1.0, nucleusiq-gemini 0.3.1, nucleusiq-groq 0.1.1, nucleusiq-mcp 0.1.1. Provider floors stay as they were (nucleusiq>=0.7.12, openai-compatible >=0.7.13).

Added — core nucleusiq

  • WebSearchTool — first-class core web search (nucleusiq.tools.WebSearchTool). Adapter pattern: one tool, swappable engines via provider=. Default is DuckDuckGoSearch (free, no API key; ddgs is a core dependency). Paid/official engines ask for config at construction: google (API key + CSE id), brave, tavily, serper. bing uses the free Bing index through ddgs — Microsoft retired the Bing Web Search API in 2025. Custom adapters register with WebSearchBackendFactory.register.
  • Search safety — results are title / URL / snippet only (no page fetch). Public http(s) URLs only (blocks javascript:, file:, localhost, private IPs). HTML and control characters are stripped from snippets. Constructor-only allowed_domains / blocked_domains (hostname suffix match + site: query hint). Tavily also receives include_domains when an allowlist is set.

Autonomous harness hardening (design: docs/design/AUTONOMOUS_HARNESS_HARDENING.md)

  • AgentResult.termination_reason — every run now says why it stopped: completed, structured_output, tool_budget, context_tool_budget, no_progress, emergency_compaction, deadline, llm_timeout, context_overflow, critic_abstain, schema_invalid, plugin_halt, empty_response, refusal, error. Set at every exit of the sync and streaming tool loops, the Autonomous dispatcher and Agent.execute(); the last exit wins so a retried-then-successful run reads completed.
  • AgentResult.diagnostics (RunReport, also agent.last_run_report) — always-on, summary-level run record with no task text or payloads: resolved configuration (window, reserve, working budget, caps, tool counts), decisions (classification), counters (rounds, LLM calls, business vs context-management tool calls, dedup banners, recall errors, stalled rounds, compactions, emergency count, write-backs), children, coverage, termination record and a bounded timeline. report.redacted() hashes identifiers for sharing; report.explain() renders a markdown block for an issue. Populated regardless of enable_tracing.
  • AgentConfig.max_context_tool_calls — context-management tools (recall_*, workspace, evidence, corpus) stay exempt from max_tool_calls but now have their own finite cap (default 2 × effective max_tool_calls). Exhaustion stops the loop with context_tool_budget instead of spinning.
  • No-progress guard — the tool loops track a per-round signature (tool names + arguments + result hashes). Two identical rounds, or two rounds made only of duplicate-call banners / [recall_error], inject a one-time nudge telling the model to answer from existing evidence; a third stops the loop with no_progress. Polling tools whose results change never trip it. A guard stop still gets the tools-free synthesis chance before the run is failed.
  • AgentConfig.enable_decomposition (default True) — False skips the Decomposer classifier LLM call in Autonomous mode and runs the task as one agent with validation, Critic and Refiner intact. For jobs where every part reads the same documents or fills one record, splitting loses coverage; this is the explicit switch (previously only max_sub_agents=1 had that effect, and it still paid the classifier call).
  • Construction-time budget validation — AgentConfig rejects an explicit ContextConfig.response_reserve >= max_context_tokens, and an explicit llm_max_output_tokens > response_reserve, with a message that explains the shared input/output window on vLLM-style servers. Defaults are never rejected.
  • execute_stream never leaks — runtime exceptions from a mode's stream become a terminal ERROR event (and a run report) instead of tearing down the consumer's generator. Configuration errors still raise at setup, as before.
  • Structured-output contract (Agent(response_format=...) is now enforced, not just requested) — every content exit of the STANDARD / streaming / Autonomous tool loops is parsed and validated against the schema (nucleusiq.agents.structured_output.contract.StructuredOutputContract). Prose or partial JSON triggers one tools-free structured finalizer call (no tools on the request, so providers with constrained decoding enforce the schema server-side; the OpenAI-compatible adapter previously only prompt-injected the schema while tools were present). A loop stopped by a budget or guard runs the finalizer instead of the prose synthesis pass. In Autonomous mode a new deterministic validation layer schema sits between L1 and L2 and feeds the validator's exact errors into the retry prompt; the Critic prompt gains an Output Contract section (JSON-for-schema is the deliverable, prose is FAIL) and the Refiner is told to return only the corrected JSON. Sub-agents never inherit response_format.
  • AgentResult.parsed / AgentResult.structured / result.schema_valid — the validated schema instance (Pydantic model, dataclass, TypedDict or dict) and the verdict {"schema", "mode", "valid", "finalizer_runs", "errors"}. result.output stays the raw text so str(result) is unchanged; the raw text is mirrored at metadata["raw_output"]. A final answer that still fails the schema keeps status=SUCCESS (the caller decides) but records termination_reason="schema_invalid" with the validator errors in structured["errors"]. Runs without response_format see no change (parsed=None, structured=None).
  • BudgetResolver (nucleusiq.agents.context.budgets, exposed as ContextEngine.budgets) — every hand-off cap in the harness now derives from the model window instead of a fixed constant tuned for 128K: per-child finding in COMPLEX synthesis (was [:2000]), Refiner candidate / tool summary (was 8 000 / 4 000), dossier gap summary, L3 LLM-review task / result, and Critic claimed_answer / evidence totals (clamp-down only — presets stay the upper bound). Each role has an overhead, a share of the prompt budget, a floor and a ceiling; on a 128K model the ceilings hold, so today's caps are the upper bound raised modestly (findings ≤ 40K chars, Refiner candidate ≤ 32K). On a 65K model three children now hand over their whole findings instead of the first 2 000 chars; on an 8K model the caps shrink instead of overflowing. RunReport.config_resolved["budgets"] shows the numbers.
  • Derived response_reserve — ContextConfig.resolve_response_reserve(config, window, max_output_tokens): an explicit value always wins; the 8 192 default is kept whenever it fits (≤ ¼ window and ≥ max_output_tokens + 512), so 128K users see no change; otherwise the reserve is derived from the window and the configured reply size. On an 8K model the old fixed 8 192 exceeded the window (utilization pinned at 100 %, emergency compaction every turn); it now resolves to 2 560. ContextEngine.resolved_response_reserve is what the ledger and the synthesis pass actually use; the report shows it under config_resolved["response_reserve"].
  • Wall clock — AgentConfig.max_execution_time (default 3600 s, 0 = unlimited) is finally enforced, gracefully: the STANDARD / streaming / Autonomous tool loops check a monotonic deadline at every loop boundary and stop with termination_reason="deadline" followed by the tools-free synthesis (or structured finalizer), so the caller gets the best answer from the evidence gathered; the Autonomous SimpleRunner refuses to start a new Critic/Refiner attempt once the budget is spent and returns the best candidate instead of abstaining. COMPLEX sub-agents inherit the parent's remaining seconds.
  • Opt-in per-call timeouts — AgentConfig.llm_call_timeout and step_timeout had defaults for years without being enforced. They are now enforced only when set explicitly (defaults stay inert so reasoning models and long tools keep working). A slow LLM call raises LLMTimeoutError (termination_reason="llm_timeout"); a slow tool becomes an Error: Tool 'x' timed out after Ns (step_timeout) tool result the model can reason about, and the loop continues. Children inherit explicit values.
  • ContextLengthError recovery — a provider 400 for an over-long prompt is no longer terminal on the first try. call_llm forces the emergency compaction tier and writes the reduced transcript back (the server is the authority when our estimate under-counts), then on a second rejection shrinks max_output_tokens so prompt + reply fits the window; only then does the run fail — with termination_reason="context_overflow" and the numbers (window, prompt≈, max_output_tokens) in the report instead of a bare 400. context_overflow, deadline and llm_timeout are sticky: the generic error recorded by execute() never overwrites them.
  • Preflight fitness — before the first LLM call the agent measures the system prompt and tool-schema tokens and computes working = window − response_reserve − system − tool_schemas. For AUTONOMOUS, fewer than preflight_min_working_tokens (16 000) means the Critic/Refiner hand-offs cannot fit: the run is downgraded to STANDARD (logged, result.mode == "standard", diagnostics.decisions["preflight"]["action"] == "downgraded_to_standard"); set preflight_downgrade=False to force Autonomous. Below preflight_marginal_working_tokens (32 000) the run proceeds and is flagged marginal. STANDARD warns below preflight_standard_min_working_tokens (8 000). The full breakdown lands in decisions["preflight"] and config_resolved["preflight"].
  • Undeclared context window is flagged — when neither ContextConfig.max_context_tokens nor a provider override of get_context_window() gives a real number, the agent warns once (Context window unknown … assuming 128000) and the report carries config_resolved["window_is_fallback"] = true, so a 65K model sized as 128K is diagnosable. MockLLM(context_window=…) declares its window for tests.
  • LLMTimeoutError (nucleusiq.llms.errors) — new LLMError subclass raised by the harness when an explicit llm_call_timeout expires.
  • Task.resources — typed list of the documents / files / records a task must cover (context["resources"] is read as a fallback; the typed field wins). task.effective_resources() normalises either. In Autonomous mode the Decomposer grounds its classification on them, every sub-task must claim a slice, and (PR F) unprocessed resources are reported and followed up.
  • Task.context is now rendered — MessageBuilder emits one bounded ## Task Context / ## Resources (N) block as a user message immediately before the objective (window-derived cap via BudgetResolver, explicit [... task context truncated] marker, resources beyond 60 summarised as a count). Previously Task.context was accepted and silently dropped, so callers pasted it into objective — where it also reached the classifier, the Critic and the validator. Tasks without context / resources produce exactly the same messages as before.
  • Grounded 4-gate classifier — the Decomposer prompt now lists the task's resources, the documents already indexed in the run corpus and the business tool names, and adds Gate 4 — separable sources, separate outputs: a split where every child would re-read the same sources, or where the deliverable is one record, is SIMPLE. With resources present each sub-task must declare "resources": [...]. Replies that omit gate4 are judged on the original three gates.
  • Coverage contract on COMPLEX splits — Decomposer._parse_analysis enforces, when Task.resources is set: every resource owned by at least one sub-task, no resource-less sub-task, and no resource owned by more than AgentConfig.decomposition_max_owners_per_resource (default 1) sub-tasks (matching is case-insensitive and accepts a trailing path component, a.pdf ↔ docs/a.pdf). Any violation downgrades to SIMPLE; the gate vector, downgrade_reason and resource counts land in diagnostics.decisions["classification"] with a decomposition_downgraded timeline event. Malformed sub-task entries are dropped before counting.
  • Children receive the parent task — COMPLEX sub-agents now execute a Task carrying the parent's attachments, context (minus the resources key), metadata and their own resources slice (Decomposer.build_child_task); a child with no declared slice inherits the full list. Findings carry resources so reconciliation can compute coverage. Previously children saw only {"id", "objective"}.
  • Shared evidence between parent and children (nucleusiq.agents.context.shared_evidence) — a COMPLEX sub-agent no longer starts from an empty store, corpus and dossier. It receives a read-only ParentEvidenceView and runs on LayeredContentStore / LayeredDocumentCorpus / LayeredEvidenceDossier: reads (recall_tool_result, search_document_corpus, get_document_chunk, evidence queries) fall through to the parent, writes stay local, the parent's telemetry is untouched. When there is something to find the child's system prompt says so (Shared evidence: N document(s) are already indexed — use search_document_corpus … before calling any read or fetch tool). After each child finishes, merge_child_evidence copies its local entries into the parent (idempotent by key / document id; MergeReport per child, evidence_merged timeline event), so the synthesis step and any follow-up can recall what the children read. Previously three children on the same nine documents each re-read them and the parent saw none of it after gather.
  • SubTaskFinding hand-off — child findings now carry status, termination_reason, touched_resources, refs (the child's offloaded tool results, recallable by the parent after the merge) and merged, and the synthesis prompt renders them under each sub-task. The {"id", "objective", "result"} shape is preserved. Per-child summaries land in diagnostics.decisions["sub_task_findings"] and diagnostics.children (with the child's resolved max_tool_calls, window, reserve and counters), so a child spawned with a smaller window or hidden call cap is visible in the report.
  • Resource touch tracking (nucleusiq.agents.context.coverage.ResourceTouchTracker) — every tool call is matched (case-insensitive, trailing path component, bounded scan of the result head) against Task.resources; the parent unions its children's touches. touched / unprocessed feed the gather-first hint now and coverage reconciliation in PR F.
  • AgentConfig.decomposition_gather_first (default False) — opt-in pre-pass for COMPLEX runs with Task.resources: one read-only "gather" child (idempotent tools only, max_tool_calls = min(parent budget, 2 × len(resources)), no synthesis pass, parent's remaining wall clock) reads every resource into the shared stores before the analysis children start; they are told already fetched: …; do not fetch them again. Skipped with a recorded reason when no resources or no idempotent tools are present (decisions["gather_first"], gather_first_completed / gather_first_skipped events).
  • ContextEngine(store=...) — optional injected ContentStore (used for the layered child store).
  • Coverage reconciliation — every run with Task.resources now ends with unprocessed = resources − touched (tool traffic ∪ children ∪ indexed corpus documents) in result.metadata["coverage"] and diagnostics.coverage (resources, touched, unprocessed, acknowledged = unprocessed but named in the answer, unaccounted, complete, followup, enforce, blocked), with a coverage_gap timeline event when something was missed. Before the answer is accepted, the harness spends exactly one bounded follow-up on the gap (AgentConfig.coverage_followup, default True, only with business tools present): in COMPLEX a coverage-followup child runs on exactly the unprocessed slice (max_tool_calls = min(parent, 2 × len(unprocessed)), shared evidence view, appended as a finding so synthesis sees it; decisions["coverage"], coverage_followup event); in SIMPLE Autonomous a validated answer that neither touched nor mentioned a declared resource gets one retry prompt naming only the missing resources (coverage_retry event) — never on the last attempt, so a gap cannot turn the only answer into nothing. STANDARD records coverage without retrying.
  • Resource gate — AgentConfig.evidence_gate_enforce=True now also applies to resources: a residual gap after the follow-up returns status=ABSTAINED, abstention_code="coverage_incomplete", the unprocessed names in abstention_reason, and the output kept for inspection. Default stays False (record, don't block).
  • RunReport.redacted() hashes nested resource lists (coverage followup, children resources / touched_resources) and recomputes findings from the redacted data so no evidence field can leak a resource name.
  • Run analyzer — nucleusiq.agents.diagnostics.analyzer: deterministic, offline rules over RunReport (counters, resolved config, decisions, children, coverage, termination — never task text or payloads). Every AgentResult.diagnostics now carries findings (sorted critical → warning → info, each with code, title, evidence naming the exact report fields used, cause, fix_now, fix_release) and de-duplicated recommendations. Catalog: LOOP_RECALL_DEADLOCK, LOOP_NO_PROGRESS, EMERGENCY_THRASH, WINDOW_FALLBACK, CHILD_WINDOW_MISMATCH, CHILD_HIDDEN_CAP, DECOMP_SHARED_SOURCE, DECOMP_COVERAGE_GAP / COVERAGE_GAP, HANDOFF_TRUNCATED, DECOMP_DOWNGRADED, CHILDREN_FAILED, PREFLIGHT_UNFIT / PREFLIGHT_MARGINAL, RESERVE_LT_MAX_TOKENS, CONTEXT_OVERFLOW_400, TOOL_BUDGET_EXHAUSTED / CONTEXT_TOOL_BUDGET_EXHAUSTED, DEADLINE / LLM_TIMEOUT, CRITIC_ABSTAIN, SYNTH_VS_STRUCTURED, SCHEMA_NOT_SATISFIED, EMPTY_RESPONSES, GATHER_SKIPPED. Public API: analyze(report), attach_findings(report), recommendations(findings), load_report(path | json | dict), registered_rules(), RunReport.analyze(); new rules are added with the @rule decorator. AgentResult.display() shows the top three findings; RunReport.explain(redact=False) renders every finding with cause / fix / evidence plus one line per child; the run-end log line ends with findings=[…]. CLI: python -m nucleusiq.agents.diagnostics report.json [--json] [--redact] (accepts - for stdin and a whole AgentResult.summary()), exit status 2 / 1 / 0 for critical / warning / clean.
  • Report data for the analyzer — RunCounters.synthesis_runs (prose synthesis passes), decisions["synthesis_handoff"] (per_finding_chars, findings, truncated child ids), and children entries now carry resources / touched_resources as lists plus result_chars.
  • Verifier parity — design invariant I-10: the verifier sees at least what the generator saw, or knows it is seeing less. Live runs showed the Critic/Refiner loop turning a correct answer into a wrong one whenever the Critic was fed less evidence than the generator (gpt-oss 120B: "the evidence only contains three invoices" → six correct records deleted). Four changes close it:
    • Budget-derived synthesis package — Agent.build_synthesis_package / _build_synthesis_messages_from_context size the curated package from BudgetResolver per consumer (new role synthesis_package for the synthesis pass, critic_evidence_total for the Critic) instead of a fixed 12_000 chars, and the workspace-notes section is no longer a fixed 2_500-char slice: every note stays present (previews shrink first) up to the derived budget. On a 65K model the Critic's package grows from 12K to 40K chars; on an 8K model it shrinks instead of overflowing.
    • Honest omission — the package degrades by dropping whole items and appending [N more … omitted for space — this list is INCOMPLETE; absence here is not absence of evidence], never by cutting a list mid-item (which reads as a shorter complete list). SynthesisPackage.metadata carries complete, included_items, omitted_items, cut_items; package.visibility_note() renders the deficit for the model. A new required Resources Processed section states, from tool traffic, how many declared resources were read and names them — a harness-verified fact every downstream role now sees.
    • The Critic gets the map and the territory — CriticRunner always includes the raw tool trace (rehydrated, capped per result from the resolved window minus the package it sits next to) alongside the package, computes a CriticView (package_complete, raw_trace_complete, caps, sizes → decisions["critic_views"]), and when the combined view is partial prepends an ## EVIDENCE VISIBILITY notice. A FAIL reached on a partial view is downgraded to UNCERTAIN (CritiqueResult.original_verdict, evidence_view, downgraded), so a partially-sighted verifier can request another pass but cannot, on its own, condemn an answer as ungrounded or force abstention. Counters critic_partial_views / critic_fail_downgraded, timeline events, explain() line, and analyzer finding CRITIC_PARTIAL_VIEW (warning).
    • The Refiner honours the view — its evidence summary now leads with the same Resources Processed facts, and when the critique came from a partial view the prompt says so and forbids removing records the Critic called unsupported unless the tool results contradict them.
    • Parity in the COMPLEX path — a synthesis over sub-agent findings usually makes no tool calls of its own, so the Critic received no trace, an empty package, and verified nine records against nothing (Gemma 4 live: pass 1.00 with CRITIC_PARTIAL_VIEW). The hand-off the synthesizer reads (Decomposer.build_findings_section, the exact text under the same window-derived per-finding cap) is now published as agent._generator_inputs and shown to the Critic (## MATERIAL THE GENERATOR WAS GIVEN) and the Refiner. CriticView.complete is defined as parity: the tool trace is shown whole (or there was none) or the package is a complete map — a tool-less generator whose inputs are all shown is a complete view, not a partial one.
  • Every receipt shape is rehydrated for verifiers — extract_raw_trace (Critic trace, Refiner summarize_tool_results) and ContextEngine.prepare_for_synthesis recognised only the post-response masker's [observation consumed] marker; results moved by Tier-1 / adaptive offload leave [context_ref: key] receipts and reached the Critic, Refiner and synthesis pass as previews. nucleusiq.agents.context.store.receipt_store_key recognises both and all three readers use it.
  • Unseen evidence is never evicted silently — a round whose tool results exceed the working budget before any assistant turn has read them used to go straight to emergency eviction (a real Gemma-4 run then invented vendors and totals for all nine files). Tier-1 now has an adaptive squeeze: when utilisation is at or past compaction_trigger, raw tool results (seen first, then unseen, largest first) are offloaded to the store as recallable previews with an explicit [offloaded to fit the context window before you saw it — call recall_tool_result(ref=...)…] hint until the tool share is back under tool_compaction_trigger. The emergency tier squeezes an oversized current-turn tail the same way instead of no-op'ing, and counts what it still had to drop: CompactionEvent.unseen_evicted, RunCounters.unseen_evidence_evicted, analyzer finding EVIDENCE_EVICTED_UNSEEN (critical).
  • Critic / Refiner caps use the resolved window — compute_per_tool_cap for both roles was fed llm.get_context_window() (a flat provider default such as Ollama's 8192) instead of the resolved window (ContextConfig → engine → provider → fallback). With a user-configured 65K window the Critic was starved to the 500-char floor per result and the Refiner's total evidence cap resolved to 0.

Fixed — nucleusiq-ollama 0.2.2

  • Multi-turn tool histories sent tool_calls[].function.arguments as a JSON string; the ollama client's pydantic Message requires a mapping, so every autonomous run with a tool round-trip failed validation on the second call. _normalize_tool_call_entry now always emits mapping arguments (dict passthrough, None → {}, JSON string decoded, undecodable → {"_raw": …}).

Fixed — nucleusiq-anthropic 0.2.2

  • Structured output built from Pydantic / dataclass / TypedDict schemas set additionalProperties: false only on the root object; Claude's grammar requires it on every object, so any nested model was rejected with a 400. The builder now closes every object recursively (raw dict schemas are passed through untouched).

Fixed — core nucleusiq 0.7.14

  • Synthesis vs response_format — the tools-free prose synthesis pass no longer runs when a structured schema is set. JSON/Pydantic answers are almost always under the word threshold, so the old "write the full deliverable" nudge replaced them with markdown. Same skip in STANDARD, streaming, and Autonomous (they share this loop).
  • Autonomous emergency-compaction loop — prepare() used to shrink only the copy sent to the LLM. The live transcript stayed fat, so the next turn re-triggered the same compaction (emergency could loop toward Autonomous's 300-call default). Any reduced prepare() view is now written back as the continuing transcript. The tool loop still stops after 3 emergency events in one run as a last-resort circuit breaker.
  • Eviction vs recall addressability — conversation compaction (Tier 2) and emergency reduction (Tier 3) both used to drop store-backed receipts with the rest of the old groups. Dedup / masker then told the model to recall_tool_result(ref=...) with no ref left in context, which produced [recall_error] and a non-convergent re-ask loop. Both tiers now emit one chat-template-safe [evidence available for recall] catalog ([observation consumed] and [context_ref:] receipts; raw tool results are offloaded to the store first). Dedup points at the catalog or list_recalled_evidence when no live marker is visible.
  • Autonomous COMPLEX sub-agent window — Decomposer.create_sub_agent used a blank AgentConfig(STANDARD), so the parent's ContextConfig (e.g. Gemma 65K) was dropped and the child engine sized against the openai-compatible 8192 last-resort floor. Sub-agents now inherit the parent's ContextConfig, or the parent's already-resolved engine window when that field is unset. Set AgentConfig.sub_agent_context only when the child should use a different window. The provider DEFAULT_CONTEXT_WINDOW = 8192 stays as the last resort for an LLM with no declared / probed window — it is not an Agent default.
  • Autonomous COMPLEX sub-agent call limit — children were also capped by a hardcoded ModelCallLimitPlugin(max_calls=15), independent of AgentConfig.max_tool_calls. That plugin is gone. Sub-agents inherit the parent's max_tool_calls (and max_retries). If the parent left max_tool_calls unset, the child uses the STANDARD default (80), not Autonomous 300.
  • Compaction evicted the task itself — both compaction tiers pinned only the first user message. Prompt templates routinely emit several leading user messages (a preamble, the Task.resources block, the objective), so under pressure the objective and the resource list were the first things dropped and the model "forgot" what it still had to process. The whole task head — leading system messages plus every user message before the first assistant turn — is now pinned (_split_task_head); memory-first transcripts keep the old first-user pin.
  • Evidence catalog was anonymous — the [evidence available for recall] catalog written by conversation and emergency compaction listed ref and tool name only. The model's own tool_calls are evicted with the group, so it could not tell which input a ref belonged to and re-fetched it, which is what fed the dedup-banner loop. Each catalog line now carries the call's arguments (args={"path": "..."}), taken from the evicted assistant message by tool_call_id or from a masked receipt's args: line.
  • Packaging — nucleusiq.agents.diagnostics was missing from [tool.setuptools] packages, so a wheel build would have shipped without the run report / analyzer package (caught by tests/unit/test_verify_scripts.py).
  • Provider-side structured parse failure ended the run — in NATIVE mode every adapter parses the reply into the schema class inside llm.call; a model that returns two JSON objects back to back, JSON followed by prose, or a field that fails validation makes that parse raise StructuredOutputError before the output contract (repair → finalizer → schema_invalid) sees the text, and the run ended in ResultStatus.ERROR (GPT-4.1-mini live: Extra data: line 2 column 1). BaseExecutionMode.call_llm now catches it once, strips the schema type while keeping the provider wire format (response_format=(wire_format, None), the tuple every adapter already accepts) and retries; the raw reply then flows through the normal contract. Recorded as timeline event structured_parse_fallback.
  • Coverage follow-up child died in preflight — Agent.execute treats max_tool_calls as the maximum number of user tools an agent may carry. The follow-up child is budgeted at 2 × unprocessed calls, so with four resources left and 19 business tools inherited from the parent it failed with "has 19 tools but STANDARD mode allows max 8" before its first LLM call, the run reported CHILDREN_FAILED, and the finding was empty because str(AgentResult) of an error is "". Auxiliary children (gather, follow-up) now get a budget that is never below the tool list they were handed (Decomposer._aux_tool_budget; the child's own context-management tools are exempt). Child failures are no longer silent: SubTaskFinding.absorb_result keeps error and renders Error: … as the finding text, report.children[].error carries it, and CHILDREN_FAILED evidence names it.

Tests — core nucleusiq

  • Office scenario replay — tests/agents/integration/test_office_scenario.py reproduces the production failure offline: Autonomous mode, 65K shared window, nine declared documents, 19 idempotent tools, response_format set, driven by a behavioural fake model (reads whatever is still unread in its context, answers classifier / Critic prompts cooperatively, emits schema-valid JSON, has no memory outside the transcript). Eight runs assert the incident invariants — termination_reason == completed, ≤ 1 emergency compaction, zero [recall_error], 9/9 resources touched with no follow-up needed, output validates, analyzer reports no LOOP_* / critical finding — across the SIMPLE path, the COMPLEX path (children inherit 65K / 40, one owner per document, evidence merged), streaming, and display() / explain(redact=True) / offline re-analysis. Under pressure: a 12K window with 14K-char documents completes with zero dedup banners (task head pinned, catalog names arguments); a stubborn model that ignores markers and banners is stopped by the no-progress guard with the tool executed once per document, a schema-valid finalizer output and LOOP_NO_PROGRESS + COVERAGE_GAP in the report; a lazy model that stops after five documents is completed by the single coverage retry, is reported as COVERAGE_GAP without it, and is withheld with ABSTAINED / coverage_incomplete when evidence_gate_enforce=True. Added this round: a literal Critic (the observed gpt-oss behaviour — counts the invoice headers in its own prompt and fails any answer with more records) passes first time at 65K because it now sees all nine, and at 12K with 9K-char documents the view is recorded honestly and the nine records survive; a greedy model that fetches all nine documents in one turn (4× the working budget) completes with zero emergency compactions and nothing evicted unseen; the same run with offloading disabled is reported as EVIDENCE_EVICTED_UNSEEN; a provider that under-reports its window (8192) no longer truncates the Critic's evidence when the user configured 65K; a COMPLEX split into nine one-document sub-tasks (capped at five children) is completed by the follow-up child with all 19 tools and no CHILDREN_FAILED; an adapter that raises SchemaParseError on the structured reply is retried once without provider-side parsing and the run completes schema-valid; the COMPLEX Critic prompt carries the three sub-agent findings and the view is complete.
  • Verifier parity unit tests — tests/agents/unit/test_critic_evidence_view.py (complete view keeps the model's FAIL; partial view downgrades it with the notice in the prompt and the counters / CRITIC_PARTIAL_VIEW in the report; PASS is never touched; the per-result cap shrinks by the package size; the Refiner prompt carries the keep-instruction and the coverage facts; a tool-less synthesis is shown the sub-agent hand-off and counts as a complete view, with or without its own tool trace; the Refiner summary leads with the same hand-off), item-level omission and coverage-section tests for the synthesis package, rehydration of [context_ref:] receipts in extract_raw_trace / prepare_for_synthesis, SubTaskFinding.absorb_result keeping a failed child's error text, and Decomposer._aux_tool_budget.
  • Live verification (not part of the suite; run with the keys in .env) — the office scenario replayed in Autonomous mode with response_format against Gemma 4 31B (65K, and 12K under 4× pressure), Gemini 2.5 Flash, GPT-4.1-mini, Claude Haiku 4.5 and gpt-oss 120B: every run terminates with completed, zero loops, schema-valid output, 9/9 documents; gpt-oss, which previously lost six of nine records to the Critic/Refiner loop, now returns 9/9 with critic verdicts pass 1; GPT-4.1-mini, which previously ended in ResultStatus.ERROR (follow-up child preflight + provider parse failure), now completes 9/9 through the follow-up child and the parse fallback; Gemma 4 COMPLEX no longer reports CRITIC_PARTIAL_VIEW.