Repository navigation
v0.27.0
Fixed
-
Compact real-app controllers retain editable-field, resource-ID and window semantics across
post-action observations, plus input verification/submission status. Sending guidance now
distinguishes an IME submit action or actual app Send control from keyboard Enter and
text-selection controls, avoiding draft/retype loops without automatically sending content. -
A confirmed pre-dispatch selector miss can recover without invalidating an evidence-valid
scenario: it requires AUA's no-action refusal with a fresh observation, zero unknown outcomes
and a later matching host terminal claim. Such journals never produce a replay candidate.
Caller-owned terminal claims are not mistaken for failed cleanup; uncertain actions, missing
claims and actual cleanup failures remain blocking. -
Judge evidence selection now preserves observed screen families, selection/checkpoint
changes and post-restart provenance instead of silently truncating to eight text frames.
Four bounded images prioritize a same-screen state pair and omit redundant final captures;
host-owned image/ref/action mappings make lifecycle evidence attributable. -
Added a price-capped, data-collection-denied open GPT-5.6 Luna vision/tool fallback profile
with reasoning disabled. A saved-evidence audit validated two consistent structured votes
while preserving an unavailable independent-system fact as unverified; no device acceptance
is implied by this audit and vote timeouts remain unchanged. -
Judge replies with oversized optional
satisfied/unsatisfiedsummaries now retain only
the schema-bounded prefix, with content-free normalization diagnostics. Every item must
already be valid; required fields, verdicts and criterion identities/results/evidence
remain strict and unchanged. -
Optional primary-flow preview/export failures now produce a sanitized QA warning and no
replay candidate, without rewriting product evidence or falsely marking cleanup failed.
Corrupt/incomplete execution journals and actual lifecycle cleanup failures remain blocking. -
An explicit forced-tool capability rejection no longer consumes the relaxed answer's time
allowance. The relaxed request receives up to 45 seconds, still capped by the unchanged
absolute vote deadline, with renewed-budget diagnostics and normal cancellation handling. -
After a provider forces relaxed tool choice, judges may recover a single strict JSON object
from message content (optionally one JSON fence), under the same compact schema and evidence
validation. Prose/ambiguous output is rejected, native calls take precedence, and sanitized
schema-repair diagnostics explain failures without retaining private model text. -
Judge tool replies identify contract bullets by compact zero-based indexes instead of
repeating long criterion strings. The host validates identities and restores exact authored
labels in source order; duplicates/out-of-range values require repair, and omissions stay unverified. -
Reasoning-only judge responses advance immediately even when providers label them normal
completions. Remaining routes share the vote deadline fairly, and judges request a separate
2048-token reasoning budget without changing controller profiles. Reported upstream charges
count toward spend limits and totals when the provider's top-level cost is zero. -
Judge requests now share a 45-second route deadline across transport retries/backoff and schema
repairs, with a 90-second deadline per vote. Stalled routes advance promptly; cancellation
preserves cleanup and result artifacts, and requests without usage mark reported cost incomplete. -
Real-app judging keeps routing retries separate from each model's schema-repair budget,
escalates reasoning-only exhaustion, and normalizes unambiguous criterion formatting while
leaving missing evidence unverified. Failed judgement spend remains in report totals. Safe
primary-flow previews may accompany an unverified evidence gap without promoting the verdict. -
The real-app harness supports opt-in private database setup proof through its existing
read-only database API. A device-clock boundary and SQL-level expected-label allowlist keep
stale evidence and private fields out of reports; final revalidation revokes obsolete proof. -
A failed provision claim retains its exact boot identity for a same-worker retry. When the
fallback acquires that same target, unattended session cleanup retires the original boot;
replacement boots and different worker, owner, or cache scopes are never adopted. -
Explicit teardown discard can archive a lost target's undo records while its own process-bound
lease is still alive, allowing cleanup and release without reconnecting to the missing target.
Other owners and sibling worker scopes remain protected. -
The real-app harness can prove setup from a current-run log mark with whitespace-tolerant
regular expressions and latest-value comparison, without exposing captured log fields. It
rechecks before cleanup so later state changes cannot inherit an earlier positive proof. -
Opt-in primary-flow export previews the exact clean controller action suffix into the run's
output directory. Incomplete or unsaveable journals cannot produce a passing export, and
observation-only runs never borrow setup actions or write global flow memory. -
The real-app harness offers a separate forbidden-foreground-package guard: installed sibling
apps are allowed, but setup, controller, and final evidence entering a forbidden app aborts the
run independently of model judgment. The existing installed-package exclusion remains available.