From 19d05f486dfaa5a2f10c07750f744d6e55367fc9 Mon Sep 17 00:00:00 2001 From: akoita Date: Sat, 5 Sep 2026 02:39:37 +0200 Subject: [PATCH] docs(pilot): record tenth consented observation (#298) --- docs/consented-pilot-v0.9.md | 32 +++++++++++++++++++++++--------- docs/roadmap.md | 9 ++++++--- 2 files changed, 29 insertions(+), 12 deletions(-) diff --git a/docs/consented-pilot-v0.9.md b/docs/consented-pilot-v0.9.md index f55af60..9974e8f 100644 --- a/docs/consented-pilot-v0.9.md +++ b/docs/consented-pilot-v0.9.md @@ -1,8 +1,8 @@ # v0.9 consented workflow-parity result **Status:** Indeterminate
-**Recorded:** 2026-09-04
-**Cases:** Nine bounded observations of one consented, anonymized case +**Recorded:** 2026-09-05
+**Cases:** Ten bounded observations of one consented, anonymized case ## Result @@ -119,6 +119,19 @@ violation paths were preserved without leaking candidate prose or private data. No draft was accepted, so the OpenAI critic was not invoked, and no adjudication, revision, approval, or export occurred. +The tenth observation under issue #298 used the declared 1,200,000 ms timeout +and standing candidate authorization. The first Anthropic author attempt failed +local factual-invariant validation on three claim paths (`sections.1.blocks.0.claims.3.text`, +`sections.1.blocks.0.claims.5.text`, `sections.1.blocks.0.claims.6.text`), classified +as `factual-invariant-rejection`. The second attempt exceeded the output-token budget +(`output_token_budget_exceeded`). The third attempt applied concise-output feedback +and resolved the token excess plus two of the prior factual-invariant violations, but +failed one remaining claim path (`sections.1.blocks.0.claims.5.text`) under +`factual-invariant-rejection`. The run exhausted its three author attempts after 412 +seconds of active provider time, inside the 20-minute cap. No draft was accepted, so +the OpenAI critic was not called, and no adjudication, revision, approval, or export +occurred. + ## Predeclared comparison gate | Dimension | Status | @@ -152,9 +165,10 @@ unresolved findings as accepted facts. authentication failure plus two default-timeout revision attempts before a new artifact existed; the eighth reached author success on attempt three and critic success on attempt one, then exhausted three invalid-response revision - attempts with the declared timeout in force; and the ninth exhausted three - author attempts under the newly classified `factual-invariant-rejection` - stage. Provider-reported cost remained unavailable. + attempts with the declared timeout in force; the ninth exhausted three + author attempts under `factual-invariant-rejection`; and the tenth exhausted + three author attempts across token-budget excess and `factual-invariant-rejection`. + Provider-reported cost remained unavailable. - Human approval and export were not completed. Review time, edit count, and user confidence were therefore unavailable. - The existing private manual CV was retained as the human baseline. It @@ -180,7 +194,7 @@ observation, while #275 records the post-citation-completion observation. The sixth observation is recorded by #286, #287 removed its typed-history storage blocker, and #290 records the confirmed but exhausted adjudication continuation. Issue #291 records the declared-timeout observation, #293 -delivered content-free failure-stage classification, and #296 records the -classified author-exhaustion observation. Issue #75 stays open, and -release-preparation issue #250 remains blocked. No approval or export is -authorized by this result. +delivered content-free failure-stage classification, #296 records the +classified author-exhaustion observation, and #298 records the tenth observation. +Issue #75 stays open, and release-preparation issue #250 remains blocked. No +approval or export is authorized by this result. diff --git a/docs/roadmap.md b/docs/roadmap.md index 235afa3..7cab70f 100644 --- a/docs/roadmap.md +++ b/docs/roadmap.md @@ -177,7 +177,7 @@ applications. | Previous | Integration hardening and outcome validation ([v0.6.0](https://github.com/akoita/draft-loop/releases/tag/v0.6.0)) | Released; validation failed | Preserve a reproducible integrated baseline without overstating application readiness | Failed representative result carried into v0.7; see [stage evidence](stage-evidence-v0.6.0.md) | | Previous | Evidence-backed CV drafting (v0.7 program) | [Released alpha.5 checkpoint](stage-evidence-v0.7.0-alpha.5.md); implementation history carried forward; outcome not validated | Produce a complete factual, source-traceable application draft | v0.8 candidate evidence now covers the bounded drafting and review vertical | | Previous | Usable CV MVP ([v0.8.0-alpha.1](https://github.com/akoita/draft-loop/releases/tag/v0.8.0-alpha.1)) | [Released alpha](stage-evidence-v0.8.0-alpha.1.md); 17/17 issues closed; representative outcome not recorded | Produce one complete, factual, reviewed, human-approved, ATS-readable CV | Representative outcome evidence remains without overstating DOCX visual coverage | -| Now | Workflow parity and release ([milestone v0.9.0](https://github.com/akoita/draft-loop/milestone/4)) | [Nine observations of one consented case are indeterminate](consented-pilot-v0.9.md); parity not validated | Demonstrate the complete application-grade workflow and publish evidence | Nine observations under #75 remain indeterminate; #296 classified author factual-invariant exhaustion; #75 and #250 remain blocked | +| Now | Workflow parity and release ([milestone v0.9.0](https://github.com/akoita/draft-loop/milestone/4)) | [Ten observations of one consented case are indeterminate](consented-pilot-v0.9.md); parity not validated | Demonstrate the complete application-grade workflow and publish evidence | Ten observations under #75 remain indeterminate; #298 exhausted author attempts across token excess and factual invariants; #75 and #250 remain blocked | | Later | Retrieval and provider quality | Integrated lexical baseline; partial components | Improve evidence selection and dependable live runs | Vector/hybrid comparison, cancellation, and provider recovery in the packaged path | | Later | Broader real-application pilot | Implemented harness; not outcome-validated | Test factuality, quality, and effort across more cases | Consented cases, calibrated measures, and recorded limitations | | Later | Production-ready beta | Partial implementation; not production-validated | Distribute a safe, dependable desktop application | Signed installers, safe migrations, recovery, accessibility, and platform evidence | @@ -431,8 +431,10 @@ exercised that classification during a fresh same-scope author attempt: all three author calls returned before timeout but failed local factual-invariant checks, correctly classified under `factual-invariant-rejection`. The run exhausted after 313 seconds of active provider time without accepted author -output, so the critic was not called. Keep #75 unvalidated and leave #250 -blocked. +output, so the critic was not called. The tenth observation under #298 also +exhausted three author attempts across token-budget excess and +`factual-invariant-rejection` after 412 active seconds. Keep #75 unvalidated and +leave #250 blocked. **Exit criterion:** The representative comparison records no factual-invariant violations or unsupported model-added facts, preserves required sections and @@ -534,6 +536,7 @@ issues retain implementation chronology. | Date | Decision | Product implication | | ---------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| 2026-09-05 | Recorded #298 as indeterminate after the fresh observation exhausted three author attempts across token-budget excess and `factual-invariant-rejection`. | The run exhausted after 412 active seconds within the 1,200,000 ms budget; attempt two exceeded output tokens, while attempts one and three failed factual-invariant checks without leaking private data. No proposal reached the critic. Parity comparison #75 and release prep #250 remain unvalidated. | | 2026-09-04 | Recorded #296 as indeterminate after the fresh observation exhausted three author attempts under `factual-invariant-rejection`. | The newly delivered failure-stage classification accurately identified factual invariant violations across all three author attempts without exposing candidate prose or private data. The run exhausted after 313 active seconds within the 1,200,000 ms budget; no proposal reached the critic. Parity comparison #75 and release prep #250 remain unvalidated. | | 2026-09-04 | Delivered #293 content-free failure-stage classification for adjudicated-revision validation. | The runtime and provider error contracts distinguish transport parsing, response-schema validation, artifact-schema validation, and factual-invariant rejection without exposing prose or private data; sanitized deterministic tests verify the ten-finding carrier shape and safe recovery. Parity comparison #75 and release prep #250 remain unvalidated pending the next live pilot observation. | | 2026-09-02 | Recorded #291 as indeterminate after the declared request timeout enabled a complete initial author/critic round but three confirmed-adjudication revision responses failed validation. | The exact two-accept, two-reject, six-nuance package and two accepted effects persisted before provider execution. The run exhausted after 959 active seconds without timeout, authentication, credit, or quota failure. #293 owns safe failure-stage classification and correction; #75 and #250 remain blocked. |