Repository navigation
Workflow study 2026-10-05 - no production promotion
Research/data release for the frozen 2026-10-05 workflow study. Not a new PASR product version; v0.4.0 remains the latest product release and its published artifacts are unchanged.
Result: no production retrieval promotion
The selected evidence-first candidate used 18.26% more cumulative provider tokens than unchanged split-4 on Luna: mean ratio 1.18258, paired 95% interval 1.04594–1.35231. Strict compound-rubric/source support was 4/40 versus 3/40, with 3 versus 5 material-error answers. These support scores are not ordinary answer accuracy. The frozen token objective failed; two requests retain unknown provider usage. Production retrieval defaults remain unchanged.
Evidence
- All 10 PASR methods: 420 successful local MCP-text output measurements, separately from provider usage and answer quality.
- 664 attempted answering trajectories: 192 original development, 24 corrected Serena, 48 across two frozen candidates and 400 confirmation attempts on 40 new questions.
- Actual Aider RepoMap in development, Repomix and Serena under the shared answer harness, not full coding-product equivalence.
- All 400 held-out semantic reviews validated. AI-mediated review, compound criteria, source-snippet qualifications and cross-repository limitations are retained.
- The explicit continuation approval completed only the remaining 128 attempts, without retrying or replacing failures. All original 272 rows and 2482 prior ledger entries remained unchanged. Full unknown reservations and the $5 ceiling were preserved; automatic promotion was forbidden.
Accounting
3081 answering requests: 3079 complete usage records and two unknown requests. Known published-rate charges: $0.636401240. Unknown full reservations: $0.008003750. Conservative bound: $0.644404990. These are not invoices. Authoring/coding/review assistant-session tokens are excluded and unmeasured.
Downloads and reproduction
The ZIP has 285 hash-verified members, including its manifest: protocols, cases, 33 exact licensed production-source snapshots, review packets and decisions, method measurements, bootstrap correction, failures, continuation approval and request ledger. Public trajectory rows omit duplicated request histories and encrypted replay state and retain original-row hashes. Each microbenchmark export redacts 12 personal-home prefixes after measurement; original token counters and original text/file hashes remain. Those redacted strings need not retokenize identically.
pasr-workflow-study-20261005-evidence.zip SHA-256:
5fd78f28078e665c9e02abd46339177e023aac38236007d1c9813c2d6c7449a8
Use SHA256SUMS.txt to verify both data assets. The standalone JSON contains all stage summaries, method-mode statistics, the frozen gate, accounting and limitations. The ZIP manifest binds the actual distributed bytes. Original protocols retain workstation-specific paths: install the pinned components and create a new relocated protocol rather than rewriting this freeze. New model runs cost money and need not reproduce sampled answers.
Original development implementation: 020c226; corrected Serena: 6858acb; approved continuation: 95987cf. This tag retains that history.
Verification and report
577 local source tests passed, four Windows symlink checks skipped; Ruff passed for all 114 files. The actual remaining 128 trajectories ran, all 400 reviews validated, all archive member hashes were checked after reopening, and the public JSON passed denominator/accounting/unknown-usage checks. The updated evidence section was rendered in Chromium at desktop and mobile widths without horizontal overflow or broken images.