[0.7.14] — 2026-09-19
Hardens Autonomous mode so runs terminate for a named reason, do not loop on compaction/recall, enforce
response_format, and give the Critic the same evidence the generator saw. Ships first-classWebSearchToolin core. Two provider patches for bugs that only those SDKs hit.Coordinated package versions:
nucleusiq0.7.14nucleusiq-ollama0.2.2 — tool-callargumentsas a mapping on the officialollamaSDKnucleusiq-anthropic0.2.2 —additionalProperties: falseon every object in structured outputUnchanged:
nucleusiq-openai0.7.1,nucleusiq-openai-compatible0.1.0,nucleusiq-gemini0.3.1,nucleusiq-groq0.1.1,nucleusiq-mcp0.1.1. Provider floors stay as they were (nucleusiq>=0.7.12, openai-compatible>=0.7.13).
Added — core nucleusiq
WebSearchTool— first-class core web search (nucleusiq.tools.WebSearchTool). Adapter pattern: one tool, swappable engines viaprovider=. Default isDuckDuckGoSearch(free, no API key;ddgsis a core dependency). Paid/official engines ask for config at construction:google(API key + CSE id),brave,tavily,serper.binguses the free Bing index throughddgs— Microsoft retired the Bing Web Search API in 2025. Custom adapters register withWebSearchBackendFactory.register.- Search safety — results are title / URL / snippet only (no page fetch). Public
http(s)URLs only (blocksjavascript:,file:, localhost, private IPs). HTML and control characters are stripped from snippets. Constructor-onlyallowed_domains/blocked_domains(hostname suffix match +site:query hint). Tavily also receivesinclude_domainswhen an allowlist is set.
Autonomous harness hardening (design: docs/design/AUTONOMOUS_HARNESS_HARDENING.md)
AgentResult.termination_reason— every run now says why it stopped:completed,structured_output,tool_budget,context_tool_budget,no_progress,emergency_compaction,deadline,llm_timeout,context_overflow,critic_abstain,schema_invalid,plugin_halt,empty_response,refusal,error. Set at every exit of the sync and streaming tool loops, the Autonomous dispatcher andAgent.execute(); the last exit wins so a retried-then-successful run readscompleted.AgentResult.diagnostics(RunReport, alsoagent.last_run_report) — always-on, summary-level run record with no task text or payloads: resolved configuration (window, reserve, working budget, caps, tool counts), decisions (classification), counters (rounds, LLM calls, business vs context-management tool calls, dedup banners, recall errors, stalled rounds, compactions, emergency count, write-backs), children, coverage, termination record and a bounded timeline.report.redacted()hashes identifiers for sharing;report.explain()renders a markdown block for an issue. Populated regardless ofenable_tracing.AgentConfig.max_context_tool_calls— context-management tools (recall_*, workspace, evidence, corpus) stay exempt frommax_tool_callsbut now have their own finite cap (default 2 × effectivemax_tool_calls). Exhaustion stops the loop withcontext_tool_budgetinstead of spinning.- No-progress guard — the tool loops track a per-round signature (tool names + arguments + result hashes). Two identical rounds, or two rounds made only of duplicate-call banners /
[recall_error], inject a one-time nudge telling the model to answer from existing evidence; a third stops the loop withno_progress. Polling tools whose results change never trip it. A guard stop still gets the tools-free synthesis chance before the run is failed. AgentConfig.enable_decomposition(defaultTrue) —Falseskips the Decomposer classifier LLM call in Autonomous mode and runs the task as one agent with validation, Critic and Refiner intact. For jobs where every part reads the same documents or fills one record, splitting loses coverage; this is the explicit switch (previously onlymax_sub_agents=1had that effect, and it still paid the classifier call).- Construction-time budget validation —
AgentConfigrejects an explicitContextConfig.response_reserve >= max_context_tokens, and an explicitllm_max_output_tokens > response_reserve, with a message that explains the shared input/output window on vLLM-style servers. Defaults are never rejected. execute_streamnever leaks — runtime exceptions from a mode's stream become a terminalERRORevent (and a run report) instead of tearing down the consumer's generator. Configuration errors still raise at setup, as before.- Structured-output contract (
Agent(response_format=...)is now enforced, not just requested) — every content exit of the STANDARD / streaming / Autonomous tool loops is parsed and validated against the schema (nucleusiq.agents.structured_output.contract.StructuredOutputContract). Prose or partial JSON triggers one tools-free structured finalizer call (notoolson the request, so providers with constrained decoding enforce the schema server-side; the OpenAI-compatible adapter previously only prompt-injected the schema while tools were present). A loop stopped by a budget or guard runs the finalizer instead of the prose synthesis pass. In Autonomous mode a new deterministic validation layerschemasits between L1 and L2 and feeds the validator's exact errors into the retry prompt; the Critic prompt gains an Output Contract section (JSON-for-schema is the deliverable, prose is FAIL) and the Refiner is told to return only the corrected JSON. Sub-agents never inheritresponse_format. AgentResult.parsed/AgentResult.structured/result.schema_valid— the validated schema instance (Pydantic model, dataclass, TypedDict or dict) and the verdict{"schema", "mode", "valid", "finalizer_runs", "errors"}.result.outputstays the raw text sostr(result)is unchanged; the raw text is mirrored atmetadata["raw_output"]. A final answer that still fails the schema keepsstatus=SUCCESS(the caller decides) but recordstermination_reason="schema_invalid"with the validator errors instructured["errors"]. Runs withoutresponse_formatsee no change (parsed=None,structured=None).BudgetResolver(nucleusiq.agents.context.budgets, exposed asContextEngine.budgets) — every hand-off cap in the harness now derives from the model window instead of a fixed constant tuned for 128K: per-child finding in COMPLEX synthesis (was[:2000]), Refiner candidate / tool summary (was 8 000 / 4 000), dossier gap summary, L3 LLM-review task / result, and Criticclaimed_answer/ evidence totals (clamp-down only — presets stay the upper bound). Each role has an overhead, a share of the prompt budget, a floor and a ceiling; on a 128K model the ceilings hold, so today's caps are the upper bound raised modestly (findings ≤ 40K chars, Refiner candidate ≤ 32K). On a 65K model three children now hand over their whole findings instead of the first 2 000 chars; on an 8K model the caps shrink instead of overflowing.RunReport.config_resolved["budgets"]shows the numbers.- Derived
response_reserve—ContextConfig.resolve_response_reserve(config, window, max_output_tokens): an explicit value always wins; the 8 192 default is kept whenever it fits (≤ ¼ window and ≥max_output_tokens + 512), so 128K users see no change; otherwise the reserve is derived from the window and the configured reply size. On an 8K model the old fixed 8 192 exceeded the window (utilization pinned at 100 %, emergency compaction every turn); it now resolves to 2 560.ContextEngine.resolved_response_reserveis what the ledger and the synthesis pass actually use; the report shows it underconfig_resolved["response_reserve"]. - Wall clock —
AgentConfig.max_execution_time(default 3600 s,0= unlimited) is finally enforced, gracefully: the STANDARD / streaming / Autonomous tool loops check a monotonic deadline at every loop boundary and stop withtermination_reason="deadline"followed by the tools-free synthesis (or structured finalizer), so the caller gets the best answer from the evidence gathered; the Autonomous SimpleRunner refuses to start a new Critic/Refiner attempt once the budget is spent and returns the best candidate instead of abstaining. COMPLEX sub-agents inherit the parent's remaining seconds. - Opt-in per-call timeouts —
AgentConfig.llm_call_timeoutandstep_timeouthad defaults for years without being enforced. They are now enforced only when set explicitly (defaults stay inert so reasoning models and long tools keep working). A slow LLM call raisesLLMTimeoutError(termination_reason="llm_timeout"); a slow tool becomes anError: Tool 'x' timed out after Ns (step_timeout)tool result the model can reason about, and the loop continues. Children inherit explicit values. ContextLengthErrorrecovery — a provider 400 for an over-long prompt is no longer terminal on the first try.call_llmforces the emergency compaction tier and writes the reduced transcript back (the server is the authority when our estimate under-counts), then on a second rejection shrinksmax_output_tokenssoprompt + replyfits the window; only then does the run fail — withtermination_reason="context_overflow"and the numbers (window,prompt≈,max_output_tokens) in the report instead of a bare 400.context_overflow,deadlineandllm_timeoutare sticky: the genericerrorrecorded byexecute()never overwrites them.- Preflight fitness — before the first LLM call the agent measures the system prompt and tool-schema tokens and computes
working = window − response_reserve − system − tool_schemas. For AUTONOMOUS, fewer thanpreflight_min_working_tokens(16 000) means the Critic/Refiner hand-offs cannot fit: the run is downgraded to STANDARD (logged,result.mode == "standard",diagnostics.decisions["preflight"]["action"] == "downgraded_to_standard"); setpreflight_downgrade=Falseto force Autonomous. Belowpreflight_marginal_working_tokens(32 000) the run proceeds and is flaggedmarginal. STANDARD warns belowpreflight_standard_min_working_tokens(8 000). The full breakdown lands indecisions["preflight"]andconfig_resolved["preflight"]. - Undeclared context window is flagged — when neither
ContextConfig.max_context_tokensnor a provider override ofget_context_window()gives a real number, the agent warns once (Context window unknown … assuming 128000) and the report carriesconfig_resolved["window_is_fallback"] = true, so a 65K model sized as 128K is diagnosable.MockLLM(context_window=…)declares its window for tests. LLMTimeoutError(nucleusiq.llms.errors) — newLLMErrorsubclass raised by the harness when an explicitllm_call_timeoutexpires.Task.resources— typed list of the documents / files / records a task must cover (context["resources"]is read as a fallback; the typed field wins).task.effective_resources()normalises either. In Autonomous mode the Decomposer grounds its classification on them, every sub-task must claim a slice, and (PR F) unprocessed resources are reported and followed up.Task.contextis now rendered —MessageBuilderemits one bounded## Task Context/## Resources (N)block as a user message immediately before the objective (window-derived cap viaBudgetResolver, explicit[... task context truncated]marker, resources beyond 60 summarised as a count). PreviouslyTask.contextwas accepted and silently dropped, so callers pasted it intoobjective— where it also reached the classifier, the Critic and the validator. Tasks withoutcontext/resourcesproduce exactly the same messages as before.- Grounded 4-gate classifier — the Decomposer prompt now lists the task's resources, the documents already indexed in the run corpus and the business tool names, and adds Gate 4 — separable sources, separate outputs: a split where every child would re-read the same sources, or where the deliverable is one record, is SIMPLE. With resources present each sub-task must declare
"resources": [...]. Replies that omitgate4are judged on the original three gates. - Coverage contract on COMPLEX splits —
Decomposer._parse_analysisenforces, whenTask.resourcesis set: every resource owned by at least one sub-task, no resource-less sub-task, and no resource owned by more thanAgentConfig.decomposition_max_owners_per_resource(default 1) sub-tasks (matching is case-insensitive and accepts a trailing path component,a.pdf↔docs/a.pdf). Any violation downgrades to SIMPLE; the gate vector,downgrade_reasonand resource counts land indiagnostics.decisions["classification"]with adecomposition_downgradedtimeline event. Malformed sub-task entries are dropped before counting. - Children receive the parent task — COMPLEX sub-agents now execute a
Taskcarrying the parent'sattachments,context(minus the resources key),metadataand their ownresourcesslice (Decomposer.build_child_task); a child with no declared slice inherits the full list. Findings carryresourcesso reconciliation can compute coverage. Previously children saw only{"id", "objective"}. - Shared evidence between parent and children (
nucleusiq.agents.context.shared_evidence) — a COMPLEX sub-agent no longer starts from an empty store, corpus and dossier. It receives a read-onlyParentEvidenceViewand runs onLayeredContentStore/LayeredDocumentCorpus/LayeredEvidenceDossier: reads (recall_tool_result,search_document_corpus,get_document_chunk, evidence queries) fall through to the parent, writes stay local, the parent's telemetry is untouched. When there is something to find the child's system prompt says so (Shared evidence: N document(s) are already indexed — use search_document_corpus … before calling any read or fetch tool). After each child finishes,merge_child_evidencecopies its local entries into the parent (idempotent by key / document id;MergeReportper child,evidence_mergedtimeline event), so the synthesis step and any follow-up can recall what the children read. Previously three children on the same nine documents each re-read them and the parent saw none of it aftergather. SubTaskFindinghand-off — child findings now carrystatus,termination_reason,touched_resources,refs(the child's offloaded tool results, recallable by the parent after the merge) andmerged, and the synthesis prompt renders them under each sub-task. The{"id", "objective", "result"}shape is preserved. Per-child summaries land indiagnostics.decisions["sub_task_findings"]anddiagnostics.children(with the child's resolvedmax_tool_calls, window, reserve and counters), so a child spawned with a smaller window or hidden call cap is visible in the report.- Resource touch tracking (
nucleusiq.agents.context.coverage.ResourceTouchTracker) — every tool call is matched (case-insensitive, trailing path component, bounded scan of the result head) againstTask.resources; the parent unions its children's touches.touched/unprocessedfeed the gather-first hint now and coverage reconciliation in PR F. AgentConfig.decomposition_gather_first(defaultFalse) — opt-in pre-pass for COMPLEX runs withTask.resources: one read-only "gather" child (idempotent tools only,max_tool_calls = min(parent budget, 2 × len(resources)), no synthesis pass, parent's remaining wall clock) reads every resource into the shared stores before the analysis children start; they are toldalready fetched: …; do not fetch them again. Skipped with a recorded reason when no resources or no idempotent tools are present (decisions["gather_first"],gather_first_completed/gather_first_skippedevents).ContextEngine(store=...)— optional injectedContentStore(used for the layered child store).- Coverage reconciliation — every run with
Task.resourcesnow ends withunprocessed = resources − touched(tool traffic ∪ children ∪ indexed corpus documents) inresult.metadata["coverage"]anddiagnostics.coverage(resources,touched,unprocessed,acknowledged= unprocessed but named in the answer,unaccounted,complete,followup,enforce,blocked), with acoverage_gaptimeline event when something was missed. Before the answer is accepted, the harness spends exactly one bounded follow-up on the gap (AgentConfig.coverage_followup, defaultTrue, only with business tools present): in COMPLEX acoverage-followupchild runs on exactly the unprocessed slice (max_tool_calls = min(parent, 2 × len(unprocessed)), shared evidence view, appended as a finding so synthesis sees it;decisions["coverage"],coverage_followupevent); in SIMPLE Autonomous a validated answer that neither touched nor mentioned a declared resource gets one retry prompt naming only the missing resources (coverage_retryevent) — never on the last attempt, so a gap cannot turn the only answer into nothing. STANDARD records coverage without retrying. - Resource gate —
AgentConfig.evidence_gate_enforce=Truenow also applies to resources: a residual gap after the follow-up returnsstatus=ABSTAINED,abstention_code="coverage_incomplete", the unprocessed names inabstention_reason, and the output kept for inspection. Default staysFalse(record, don't block). RunReport.redacted()hashes nested resource lists (coveragefollowup, childrenresources/touched_resources) and recomputes findings from the redacted data so no evidence field can leak a resource name.- Run analyzer —
nucleusiq.agents.diagnostics.analyzer: deterministic, offline rules overRunReport(counters, resolved config, decisions, children, coverage, termination — never task text or payloads). EveryAgentResult.diagnosticsnow carriesfindings(sorted critical → warning → info, each withcode,title,evidencenaming the exact report fields used,cause,fix_now,fix_release) and de-duplicatedrecommendations. Catalog:LOOP_RECALL_DEADLOCK,LOOP_NO_PROGRESS,EMERGENCY_THRASH,WINDOW_FALLBACK,CHILD_WINDOW_MISMATCH,CHILD_HIDDEN_CAP,DECOMP_SHARED_SOURCE,DECOMP_COVERAGE_GAP/COVERAGE_GAP,HANDOFF_TRUNCATED,DECOMP_DOWNGRADED,CHILDREN_FAILED,PREFLIGHT_UNFIT/PREFLIGHT_MARGINAL,RESERVE_LT_MAX_TOKENS,CONTEXT_OVERFLOW_400,TOOL_BUDGET_EXHAUSTED/CONTEXT_TOOL_BUDGET_EXHAUSTED,DEADLINE/LLM_TIMEOUT,CRITIC_ABSTAIN,SYNTH_VS_STRUCTURED,SCHEMA_NOT_SATISFIED,EMPTY_RESPONSES,GATHER_SKIPPED. Public API:analyze(report),attach_findings(report),recommendations(findings),load_report(path | json | dict),registered_rules(),RunReport.analyze(); new rules are added with the@ruledecorator.AgentResult.display()shows the top three findings;RunReport.explain(redact=False)renders every finding with cause / fix / evidence plus one line per child; the run-end log line ends withfindings=[…]. CLI:python -m nucleusiq.agents.diagnostics report.json [--json] [--redact](accepts-for stdin and a wholeAgentResult.summary()), exit status 2 / 1 / 0 for critical / warning / clean. - Report data for the analyzer —
RunCounters.synthesis_runs(prose synthesis passes),decisions["synthesis_handoff"](per_finding_chars,findings,truncatedchild ids), and children entries now carryresources/touched_resourcesas lists plusresult_chars. - Verifier parity — design invariant I-10: the verifier sees at least what the generator saw, or knows it is seeing less. Live runs showed the Critic/Refiner loop turning a correct answer into a wrong one whenever the Critic was fed less evidence than the generator (gpt-oss 120B: "the evidence only contains three invoices" → six correct records deleted). Four changes close it:
- Budget-derived synthesis package —
Agent.build_synthesis_package/_build_synthesis_messages_from_contextsize the curated package fromBudgetResolverper consumer (new rolesynthesis_packagefor the synthesis pass,critic_evidence_totalfor the Critic) instead of a fixed12_000chars, and the workspace-notes section is no longer a fixed2_500-char slice: every note stays present (previews shrink first) up to the derived budget. On a 65K model the Critic's package grows from 12K to 40K chars; on an 8K model it shrinks instead of overflowing. - Honest omission — the package degrades by dropping whole items and appending
[N more … omitted for space — this list is INCOMPLETE; absence here is not absence of evidence], never by cutting a list mid-item (which reads as a shorter complete list).SynthesisPackage.metadatacarriescomplete,included_items,omitted_items,cut_items;package.visibility_note()renders the deficit for the model. A new required Resources Processed section states, from tool traffic, how many declared resources were read and names them — a harness-verified fact every downstream role now sees. - The Critic gets the map and the territory —
CriticRunneralways includes the raw tool trace (rehydrated, capped per result from the resolved window minus the package it sits next to) alongside the package, computes aCriticView(package_complete,raw_trace_complete, caps, sizes →decisions["critic_views"]), and when the combined view is partial prepends an## EVIDENCE VISIBILITYnotice. A FAIL reached on a partial view is downgraded to UNCERTAIN (CritiqueResult.original_verdict,evidence_view,downgraded), so a partially-sighted verifier can request another pass but cannot, on its own, condemn an answer as ungrounded or force abstention. Counterscritic_partial_views/critic_fail_downgraded, timeline events,explain()line, and analyzer findingCRITIC_PARTIAL_VIEW(warning). - The Refiner honours the view — its evidence summary now leads with the same Resources Processed facts, and when the critique came from a partial view the prompt says so and forbids removing records the Critic called unsupported unless the tool results contradict them.
- Parity in the COMPLEX path — a synthesis over sub-agent findings usually makes no tool calls of its own, so the Critic received no trace, an empty package, and verified nine records against nothing (Gemma 4 live:
pass 1.00withCRITIC_PARTIAL_VIEW). The hand-off the synthesizer reads (Decomposer.build_findings_section, the exact text under the same window-derived per-finding cap) is now published asagent._generator_inputsand shown to the Critic (## MATERIAL THE GENERATOR WAS GIVEN) and the Refiner.CriticView.completeis defined as parity: the tool trace is shown whole (or there was none) or the package is a complete map — a tool-less generator whose inputs are all shown is a complete view, not a partial one.
- Budget-derived synthesis package —
- Every receipt shape is rehydrated for verifiers —
extract_raw_trace(Critic trace, Refinersummarize_tool_results) andContextEngine.prepare_for_synthesisrecognised only the post-response masker's[observation consumed]marker; results moved by Tier-1 / adaptive offload leave[context_ref: key]receipts and reached the Critic, Refiner and synthesis pass as previews.nucleusiq.agents.context.store.receipt_store_keyrecognises both and all three readers use it. - Unseen evidence is never evicted silently — a round whose tool results exceed the working budget before any assistant turn has read them used to go straight to emergency eviction (a real Gemma-4 run then invented vendors and totals for all nine files). Tier-1 now has an adaptive squeeze: when utilisation is at or past
compaction_trigger, raw tool results (seen first, then unseen, largest first) are offloaded to the store as recallable previews with an explicit[offloaded to fit the context window before you saw it — call recall_tool_result(ref=...)…]hint until the tool share is back undertool_compaction_trigger. The emergency tier squeezes an oversized current-turn tail the same way instead of no-op'ing, and counts what it still had to drop:CompactionEvent.unseen_evicted,RunCounters.unseen_evidence_evicted, analyzer findingEVIDENCE_EVICTED_UNSEEN(critical). - Critic / Refiner caps use the resolved window —
compute_per_tool_capfor both roles was fedllm.get_context_window()(a flat provider default such as Ollama's 8192) instead of the resolved window (ContextConfig→ engine → provider → fallback). With a user-configured 65K window the Critic was starved to the 500-char floor per result and the Refiner's total evidence cap resolved to 0.
Fixed — nucleusiq-ollama 0.2.2
- Multi-turn tool histories sent
tool_calls[].function.argumentsas a JSON string; theollamaclient's pydanticMessagerequires a mapping, so every autonomous run with a tool round-trip failed validation on the second call._normalize_tool_call_entrynow always emits mapping arguments (dict passthrough,None → {}, JSON string decoded, undecodable →{"_raw": …}).
Fixed — nucleusiq-anthropic 0.2.2
- Structured output built from Pydantic / dataclass / TypedDict schemas set
additionalProperties: falseonly on the root object; Claude's grammar requires it on every object, so any nested model was rejected with a 400. The builder now closes every object recursively (raw dict schemas are passed through untouched).
Fixed — core nucleusiq 0.7.14
- Synthesis vs
response_format— the tools-free prose synthesis pass no longer runs when a structured schema is set. JSON/Pydantic answers are almost always under the word threshold, so the old "write the full deliverable" nudge replaced them with markdown. Same skip in STANDARD, streaming, and Autonomous (they share this loop). - Autonomous emergency-compaction loop —
prepare()used to shrink only the copy sent to the LLM. The live transcript stayed fat, so the next turn re-triggered the same compaction (emergency could loop toward Autonomous's 300-call default). Any reducedprepare()view is now written back as the continuing transcript. The tool loop still stops after 3 emergency events in one run as a last-resort circuit breaker. - Eviction vs recall addressability — conversation compaction (Tier 2) and emergency reduction (Tier 3) both used to drop store-backed receipts with the rest of the old groups. Dedup / masker then told the model to
recall_tool_result(ref=...)with no ref left in context, which produced[recall_error]and a non-convergent re-ask loop. Both tiers now emit one chat-template-safe[evidence available for recall]catalog ([observation consumed]and[context_ref:]receipts; raw tool results are offloaded to the store first). Dedup points at the catalog orlist_recalled_evidencewhen no live marker is visible. - Autonomous COMPLEX sub-agent window —
Decomposer.create_sub_agentused a blankAgentConfig(STANDARD), so the parent'sContextConfig(e.g. Gemma 65K) was dropped and the child engine sized against the openai-compatible 8192 last-resort floor. Sub-agents now inherit the parent'sContextConfig, or the parent's already-resolved engine window when that field is unset. SetAgentConfig.sub_agent_contextonly when the child should use a different window. The providerDEFAULT_CONTEXT_WINDOW = 8192stays as the last resort for an LLM with no declared / probed window — it is not an Agent default. - Autonomous COMPLEX sub-agent call limit — children were also capped by a hardcoded
ModelCallLimitPlugin(max_calls=15), independent ofAgentConfig.max_tool_calls. That plugin is gone. Sub-agents inherit the parent'smax_tool_calls(andmax_retries). If the parent leftmax_tool_callsunset, the child uses the STANDARD default (80), not Autonomous 300. - Compaction evicted the task itself — both compaction tiers pinned only the first user message. Prompt templates routinely emit several leading user messages (a preamble, the
Task.resourcesblock, the objective), so under pressure the objective and the resource list were the first things dropped and the model "forgot" what it still had to process. The whole task head — leading system messages plus every user message before the first assistant turn — is now pinned (_split_task_head); memory-first transcripts keep the old first-user pin. - Evidence catalog was anonymous — the
[evidence available for recall]catalog written by conversation and emergency compaction listedrefand tool name only. The model's owntool_callsare evicted with the group, so it could not tell which input a ref belonged to and re-fetched it, which is what fed the dedup-banner loop. Each catalog line now carries the call's arguments (args={"path": "..."}), taken from the evicted assistant message bytool_call_idor from a masked receipt'sargs:line. - Packaging —
nucleusiq.agents.diagnosticswas missing from[tool.setuptools] packages, so a wheel build would have shipped without the run report / analyzer package (caught bytests/unit/test_verify_scripts.py). - Provider-side structured parse failure ended the run — in NATIVE mode every adapter parses the reply into the schema class inside
llm.call; a model that returns two JSON objects back to back, JSON followed by prose, or a field that fails validation makes that parse raiseStructuredOutputErrorbefore the output contract (repair → finalizer →schema_invalid) sees the text, and the run ended inResultStatus.ERROR(GPT-4.1-mini live:Extra data: line 2 column 1).BaseExecutionMode.call_llmnow catches it once, strips the schema type while keeping the provider wire format (response_format=(wire_format, None), the tuple every adapter already accepts) and retries; the raw reply then flows through the normal contract. Recorded as timeline eventstructured_parse_fallback. - Coverage follow-up child died in preflight —
Agent.executetreatsmax_tool_callsas the maximum number of user tools an agent may carry. The follow-up child is budgeted at2 × unprocessedcalls, so with four resources left and 19 business tools inherited from the parent it failed with "has 19 tools but STANDARD mode allows max 8" before its first LLM call, the run reportedCHILDREN_FAILED, and the finding was empty becausestr(AgentResult)of an error is"". Auxiliary children (gather, follow-up) now get a budget that is never below the tool list they were handed (Decomposer._aux_tool_budget; the child's own context-management tools are exempt). Child failures are no longer silent:SubTaskFinding.absorb_resultkeepserrorand rendersError: …as the finding text,report.children[].errorcarries it, andCHILDREN_FAILEDevidence names it.
Tests — core nucleusiq
- Office scenario replay —
tests/agents/integration/test_office_scenario.pyreproduces the production failure offline: Autonomous mode, 65K shared window, nine declared documents, 19 idempotent tools,response_formatset, driven by a behavioural fake model (reads whatever is still unread in its context, answers classifier / Critic prompts cooperatively, emits schema-valid JSON, has no memory outside the transcript). Eight runs assert the incident invariants —termination_reason == completed, ≤ 1 emergency compaction, zero[recall_error], 9/9 resources touched with no follow-up needed, output validates, analyzer reports noLOOP_*/ critical finding — across the SIMPLE path, the COMPLEX path (children inherit 65K / 40, one owner per document, evidence merged), streaming, anddisplay()/explain(redact=True)/ offline re-analysis. Under pressure: a 12K window with 14K-char documents completes with zero dedup banners (task head pinned, catalog names arguments); a stubborn model that ignores markers and banners is stopped by the no-progress guard with the tool executed once per document, a schema-valid finalizer output andLOOP_NO_PROGRESS+COVERAGE_GAPin the report; a lazy model that stops after five documents is completed by the single coverage retry, is reported asCOVERAGE_GAPwithout it, and is withheld withABSTAINED/coverage_incompletewhenevidence_gate_enforce=True. Added this round: a literal Critic (the observed gpt-oss behaviour — counts the invoice headers in its own prompt and fails any answer with more records) passes first time at 65K because it now sees all nine, and at 12K with 9K-char documents the view is recorded honestly and the nine records survive; a greedy model that fetches all nine documents in one turn (4× the working budget) completes with zero emergency compactions and nothing evicted unseen; the same run with offloading disabled is reported asEVIDENCE_EVICTED_UNSEEN; a provider that under-reports its window (8192) no longer truncates the Critic's evidence when the user configured 65K; a COMPLEX split into nine one-document sub-tasks (capped at five children) is completed by the follow-up child with all 19 tools and noCHILDREN_FAILED; an adapter that raisesSchemaParseErroron the structured reply is retried once without provider-side parsing and the run completes schema-valid; the COMPLEX Critic prompt carries the three sub-agent findings and the view is complete. - Verifier parity unit tests —
tests/agents/unit/test_critic_evidence_view.py(complete view keeps the model's FAIL; partial view downgrades it with the notice in the prompt and the counters /CRITIC_PARTIAL_VIEWin the report; PASS is never touched; the per-result cap shrinks by the package size; the Refiner prompt carries the keep-instruction and the coverage facts; a tool-less synthesis is shown the sub-agent hand-off and counts as a complete view, with or without its own tool trace; the Refiner summary leads with the same hand-off), item-level omission and coverage-section tests for the synthesis package, rehydration of[context_ref:]receipts inextract_raw_trace/prepare_for_synthesis,SubTaskFinding.absorb_resultkeeping a failed child's error text, andDecomposer._aux_tool_budget. - Live verification (not part of the suite; run with the keys in
.env) — the office scenario replayed in Autonomous mode withresponse_formatagainst Gemma 4 31B (65K, and 12K under 4× pressure), Gemini 2.5 Flash, GPT-4.1-mini, Claude Haiku 4.5 and gpt-oss 120B: every run terminates withcompleted, zero loops, schema-valid output, 9/9 documents; gpt-oss, which previously lost six of nine records to the Critic/Refiner loop, now returns 9/9 withcritic verdicts pass 1; GPT-4.1-mini, which previously ended inResultStatus.ERROR(follow-up child preflight + provider parse failure), now completes 9/9 through the follow-up child and the parse fallback; Gemma 4 COMPLEX no longer reportsCRITIC_PARTIAL_VIEW.