Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 23 additions & 9 deletions docs/consented-pilot-v0.9.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
# v0.9 consented workflow-parity result

**Status:** Indeterminate<br>
**Recorded:** 2026-09-04<br>
**Cases:** Nine bounded observations of one consented, anonymized case
**Recorded:** 2026-09-05<br>
**Cases:** Ten bounded observations of one consented, anonymized case

## Result

Expand Down Expand Up @@ -119,6 +119,19 @@ violation paths were preserved without leaking candidate prose or private data.
No draft was accepted, so the OpenAI critic was not invoked, and no
adjudication, revision, approval, or export occurred.

The tenth observation under issue #298 used the declared 1,200,000 ms timeout
and standing candidate authorization. The first Anthropic author attempt failed
local factual-invariant validation on three claim paths (`sections.1.blocks.0.claims.3.text`,
`sections.1.blocks.0.claims.5.text`, `sections.1.blocks.0.claims.6.text`), classified
as `factual-invariant-rejection`. The second attempt exceeded the output-token budget
(`output_token_budget_exceeded`). The third attempt applied concise-output feedback
and resolved the token excess plus two of the prior factual-invariant violations, but
failed one remaining claim path (`sections.1.blocks.0.claims.5.text`) under
`factual-invariant-rejection`. The run exhausted its three author attempts after 412
seconds of active provider time, inside the 20-minute cap. No draft was accepted, so
the OpenAI critic was not called, and no adjudication, revision, approval, or export
occurred.

## Predeclared comparison gate

| Dimension | Status |
Expand Down Expand Up @@ -152,9 +165,10 @@ unresolved findings as accepted facts.
authentication failure plus two default-timeout revision attempts before a
new artifact existed; the eighth reached author success on attempt three and
critic success on attempt one, then exhausted three invalid-response revision
attempts with the declared timeout in force; and the ninth exhausted three
author attempts under the newly classified `factual-invariant-rejection`
stage. Provider-reported cost remained unavailable.
attempts with the declared timeout in force; the ninth exhausted three
author attempts under `factual-invariant-rejection`; and the tenth exhausted
three author attempts across token-budget excess and `factual-invariant-rejection`.
Provider-reported cost remained unavailable.
- Human approval and export were not completed. Review time, edit count, and
user confidence were therefore unavailable.
- The existing private manual CV was retained as the human baseline. It
Expand All @@ -180,7 +194,7 @@ observation, while #275 records the post-citation-completion observation. The
sixth observation is recorded by #286, #287 removed its typed-history storage
blocker, and #290 records the confirmed but exhausted adjudication
continuation. Issue #291 records the declared-timeout observation, #293
delivered content-free failure-stage classification, and #296 records the
classified author-exhaustion observation. Issue #75 stays open, and
release-preparation issue #250 remains blocked. No approval or export is
authorized by this result.
delivered content-free failure-stage classification, #296 records the
classified author-exhaustion observation, and #298 records the tenth observation.
Issue #75 stays open, and release-preparation issue #250 remains blocked. No
approval or export is authorized by this result.
9 changes: 6 additions & 3 deletions docs/roadmap.md
Original file line number Diff line number Diff line change
Expand Up @@ -177,7 +177,7 @@ applications.
| Previous | Integration hardening and outcome validation ([v0.6.0](https://github.com/akoita/draft-loop/releases/tag/v0.6.0)) | Released; validation failed | Preserve a reproducible integrated baseline without overstating application readiness | Failed representative result carried into v0.7; see [stage evidence](stage-evidence-v0.6.0.md) |
| Previous | Evidence-backed CV drafting (v0.7 program) | [Released alpha.5 checkpoint](stage-evidence-v0.7.0-alpha.5.md); implementation history carried forward; outcome not validated | Produce a complete factual, source-traceable application draft | v0.8 candidate evidence now covers the bounded drafting and review vertical |
| Previous | Usable CV MVP ([v0.8.0-alpha.1](https://github.com/akoita/draft-loop/releases/tag/v0.8.0-alpha.1)) | [Released alpha](stage-evidence-v0.8.0-alpha.1.md); 17/17 issues closed; representative outcome not recorded | Produce one complete, factual, reviewed, human-approved, ATS-readable CV | Representative outcome evidence remains without overstating DOCX visual coverage |
| Now | Workflow parity and release ([milestone v0.9.0](https://github.com/akoita/draft-loop/milestone/4)) | [Nine observations of one consented case are indeterminate](consented-pilot-v0.9.md); parity not validated | Demonstrate the complete application-grade workflow and publish evidence | Nine observations under #75 remain indeterminate; #296 classified author factual-invariant exhaustion; #75 and #250 remain blocked |
| Now | Workflow parity and release ([milestone v0.9.0](https://github.com/akoita/draft-loop/milestone/4)) | [Ten observations of one consented case are indeterminate](consented-pilot-v0.9.md); parity not validated | Demonstrate the complete application-grade workflow and publish evidence | Ten observations under #75 remain indeterminate; #298 exhausted author attempts across token excess and factual invariants; #75 and #250 remain blocked |
| Later | Retrieval and provider quality | Integrated lexical baseline; partial components | Improve evidence selection and dependable live runs | Vector/hybrid comparison, cancellation, and provider recovery in the packaged path |
| Later | Broader real-application pilot | Implemented harness; not outcome-validated | Test factuality, quality, and effort across more cases | Consented cases, calibrated measures, and recorded limitations |
| Later | Production-ready beta | Partial implementation; not production-validated | Distribute a safe, dependable desktop application | Signed installers, safe migrations, recovery, accessibility, and platform evidence |
Expand Down Expand Up @@ -431,8 +431,10 @@ exercised that classification during a fresh same-scope author attempt: all
three author calls returned before timeout but failed local factual-invariant
checks, correctly classified under `factual-invariant-rejection`. The run
exhausted after 313 seconds of active provider time without accepted author
output, so the critic was not called. Keep #75 unvalidated and leave #250
blocked.
output, so the critic was not called. The tenth observation under #298 also
exhausted three author attempts across token-budget excess and
`factual-invariant-rejection` after 412 active seconds. Keep #75 unvalidated and
leave #250 blocked.

**Exit criterion:** The representative comparison records no factual-invariant
violations or unsupported model-added facts, preserves required sections and
Expand Down Expand Up @@ -534,6 +536,7 @@ issues retain implementation chronology.

| Date | Decision | Product implication |
| ---------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 2026-09-05 | Recorded #298 as indeterminate after the fresh observation exhausted three author attempts across token-budget excess and `factual-invariant-rejection`. | The run exhausted after 412 active seconds within the 1,200,000 ms budget; attempt two exceeded output tokens, while attempts one and three failed factual-invariant checks without leaking private data. No proposal reached the critic. Parity comparison #75 and release prep #250 remain unvalidated. |
| 2026-09-04 | Recorded #296 as indeterminate after the fresh observation exhausted three author attempts under `factual-invariant-rejection`. | The newly delivered failure-stage classification accurately identified factual invariant violations across all three author attempts without exposing candidate prose or private data. The run exhausted after 313 active seconds within the 1,200,000 ms budget; no proposal reached the critic. Parity comparison #75 and release prep #250 remain unvalidated. |
| 2026-09-04 | Delivered #293 content-free failure-stage classification for adjudicated-revision validation. | The runtime and provider error contracts distinguish transport parsing, response-schema validation, artifact-schema validation, and factual-invariant rejection without exposing prose or private data; sanitized deterministic tests verify the ten-finding carrier shape and safe recovery. Parity comparison #75 and release prep #250 remain unvalidated pending the next live pilot observation. |
| 2026-09-02 | Recorded #291 as indeterminate after the declared request timeout enabled a complete initial author/critic round but three confirmed-adjudication revision responses failed validation. | The exact two-accept, two-reject, six-nuance package and two accepted effects persisted before provider execution. The run exhausted after 959 active seconds without timeout, authentication, credit, or quota failure. #293 owns safe failure-stage classification and correction; #75 and #250 remain blocked. |
Expand Down