Skip to content

v0.27.0

Choose a tag to compare

@github-actions github-actions released this 16 Sep 13:42
· 216 commits to main since this release

Fixed

  • Compact real-app controllers retain editable-field, resource-ID and window semantics across
    post-action observations, plus input verification/submission status. Sending guidance now
    distinguishes an IME submit action or actual app Send control from keyboard Enter and
    text-selection controls, avoiding draft/retype loops without automatically sending content.

  • A confirmed pre-dispatch selector miss can recover without invalidating an evidence-valid
    scenario: it requires AUA's no-action refusal with a fresh observation, zero unknown outcomes
    and a later matching host terminal claim. Such journals never produce a replay candidate.
    Caller-owned terminal claims are not mistaken for failed cleanup; uncertain actions, missing
    claims and actual cleanup failures remain blocking.

  • Judge evidence selection now preserves observed screen families, selection/checkpoint
    changes and post-restart provenance instead of silently truncating to eight text frames.
    Four bounded images prioritize a same-screen state pair and omit redundant final captures;
    host-owned image/ref/action mappings make lifecycle evidence attributable.

  • Added a price-capped, data-collection-denied open GPT-5.6 Luna vision/tool fallback profile
    with reasoning disabled. A saved-evidence audit validated two consistent structured votes
    while preserving an unavailable independent-system fact as unverified; no device acceptance
    is implied by this audit and vote timeouts remain unchanged.

  • Judge replies with oversized optional satisfied/unsatisfied summaries now retain only
    the schema-bounded prefix, with content-free normalization diagnostics. Every item must
    already be valid; required fields, verdicts and criterion identities/results/evidence
    remain strict and unchanged.

  • Optional primary-flow preview/export failures now produce a sanitized QA warning and no
    replay candidate, without rewriting product evidence or falsely marking cleanup failed.
    Corrupt/incomplete execution journals and actual lifecycle cleanup failures remain blocking.

  • An explicit forced-tool capability rejection no longer consumes the relaxed answer's time
    allowance. The relaxed request receives up to 45 seconds, still capped by the unchanged
    absolute vote deadline, with renewed-budget diagnostics and normal cancellation handling.

  • After a provider forces relaxed tool choice, judges may recover a single strict JSON object
    from message content (optionally one JSON fence), under the same compact schema and evidence
    validation. Prose/ambiguous output is rejected, native calls take precedence, and sanitized
    schema-repair diagnostics explain failures without retaining private model text.

  • Judge tool replies identify contract bullets by compact zero-based indexes instead of
    repeating long criterion strings. The host validates identities and restores exact authored
    labels in source order; duplicates/out-of-range values require repair, and omissions stay unverified.

  • Reasoning-only judge responses advance immediately even when providers label them normal
    completions. Remaining routes share the vote deadline fairly, and judges request a separate
    2048-token reasoning budget without changing controller profiles. Reported upstream charges
    count toward spend limits and totals when the provider's top-level cost is zero.

  • Judge requests now share a 45-second route deadline across transport retries/backoff and schema
    repairs, with a 90-second deadline per vote. Stalled routes advance promptly; cancellation
    preserves cleanup and result artifacts, and requests without usage mark reported cost incomplete.

  • Real-app judging keeps routing retries separate from each model's schema-repair budget,
    escalates reasoning-only exhaustion, and normalizes unambiguous criterion formatting while
    leaving missing evidence unverified. Failed judgement spend remains in report totals. Safe
    primary-flow previews may accompany an unverified evidence gap without promoting the verdict.

  • The real-app harness supports opt-in private database setup proof through its existing
    read-only database API. A device-clock boundary and SQL-level expected-label allowlist keep
    stale evidence and private fields out of reports; final revalidation revokes obsolete proof.

  • A failed provision claim retains its exact boot identity for a same-worker retry. When the
    fallback acquires that same target, unattended session cleanup retires the original boot;
    replacement boots and different worker, owner, or cache scopes are never adopted.

  • Explicit teardown discard can archive a lost target's undo records while its own process-bound
    lease is still alive, allowing cleanup and release without reconnecting to the missing target.
    Other owners and sibling worker scopes remain protected.

  • The real-app harness can prove setup from a current-run log mark with whitespace-tolerant
    regular expressions and latest-value comparison, without exposing captured log fields. It
    rechecks before cleanup so later state changes cannot inherit an earlier positive proof.

  • Opt-in primary-flow export previews the exact clean controller action suffix into the run's
    output directory. Incomplete or unsaveable journals cannot produce a passing export, and
    observation-only runs never borrow setup actions or write global flow memory.

  • The real-app harness offers a separate forbidden-foreground-package guard: installed sibling
    apps are allowed, but setup, controller, and final evidence entering a forbidden app aborts the
    run independently of model judgment. The existing installed-package exclusion remains available.