Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 33 additions & 0 deletions docs/BUILD_LOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -350,3 +350,36 @@ Evidence: `docs/v0.3/results/R7_DEVELOPMENT_AUDIT.md` and immutable JSON receipt
- Generated all 18 blind rating packets before outcomes, with index `54d78382b3ddbe15cba1f8153275e8149d32ddaa5192163f99ca5f43d903e8fe`, and verified zero remaining audit containers/volumes.

Evidence: `docs/v0.3/results/R7_HELD_OUT_AUDIT.md`, immutable execution receipts, and blind packet bundle. Full R7 remains blocked on two independent experienced TypeScript ratings and adjudication; R5/R6 have not started.

## 2026-08-01 — R&D integration and Executable Operator Model proposal

- Merged protected PR #18 into `codex/shadow-cockpit-rnd` after all five GitHub checks passed, then deleted the short-lived R7 head locally and from origin.
- Retargeted the stale Jules infrastructure PR #8 away from `main` and into R&D, updated it from the current base, passed all five checks, merged it, and deleted the short-lived remote head. `main` remained untouched; origin now has only `main` and the single persistent coding branch `codex/shadow-cockpit-rnd`.
- Audited the original product goal against current runtime evidence. The live agent plane, v0.3 cockpit, cold relay, delayed-transfer study, and skill-preservation claim remain incomplete; the automatic compiler gate is not substituted for product completion.
- Proposed ADR-007: every meaningful agent checkpoint can produce tested software, an executable `observe → actuate → recover` surface, and a local Executable Operator Model containing only behaviorally supported causal claims.
- Added the draft post-R7 specification with deterministic invalidation, shadow-control, context-starved relay, control-dividend, privacy, isolation, and ablation requirements.
- Ran the spec-driven strict validator after correcting its required heading and traceability format: 98/100, no errors. Its only warning expects an HTTP method/path; the spec explicitly keeps this boundary local and injected in the extension host instead of inventing a network API.
- Updated the PRD and v0.3 index so attention is selected from operator-model divergence and uncovered recovery routes rather than random functions, timer prompts, or question counts.

Evidence: merged PRs [#18](https://github.com/yava-code/PureFlow/pull/18) and [#8](https://github.com/yava-code/PureFlow/pull/8); `docs/v0.3/ADR-007-EXECUTABLE-OPERATOR-MODEL.md`; `docs/v0.3/OPERATOR_MODEL_SPEC.md`; `docs/v0.3/GOAL_COMPLETION_AUDIT.md`. ADR-007 remains proposed and implementation remains gated by two independent R7 expert ratings.

## 2026-08-01 — Evidence-Carrying Generation proposal

- Closed the remaining architecture gap between critical-seam takeover evidence and accountability for the rest of a large generated change.
- Proposed ADR-008: a bidirectional Intent Ledger reconciles every changed line to a bounded semantic unit and labels its attribution `supported`, `claimed`, `unattributed`, `contradicted`, or `stale`.
- Kept agent-emitted intent references untrusted, required independent structural/evidence joins, and prohibited ledger coverage from updating human readiness.
- Added deterministic handling for generated artifacts, total changed-line reconciliation, negative trace-washing rules, a compact change-account UX, and a held-out ablation against raw diffs, AI summaries, and AST navigation.
- Distinguished the proposal from domain DSLs, requirements traceability, proof-carrying code, and generated explanations. No product-effectiveness or skill-retention result is claimed.

Evidence: `docs/v0.3/ADR-008-EVIDENCE-CARRYING-GENERATION.md`; updated PRD, thesis, project state, completion audit, and v0.3 index. Implementation remains gated by the complete two-rater R7 expert audit.

## 2026-08-01 — Decision Futures and Takeover Envelope proposal

- Closed the remaining architecture gap between off-path takeover practice and real live engineering authority.
- Proposed ADR-009: agents speculatively implement viable alternatives and continue unrelated work while a bounded, precommitted Decision Future offers one high-leverage human choice at a low-cost breakpoint.
- Distinguished on-time live influence from late counterfactual practice, autonomous default, skip, and integrity failure. A late agreement can never be upgraded into production-decision evidence.
- Added false-fork abstention, immutable deadline/default authority, comparable-alternative rules, interruption-cost scheduling, and executable integration evidence.
- Defined separate Autonomy and Takeover Envelopes so full agentic coding remains available while uncovered or stale human-control seams stay explicit.
- Added a three-arm ablation against autonomous explanation and shadow-only prediction. No product-effectiveness, live-influence, or skill-retention result is claimed.

Evidence: `docs/v0.3/ADR-009-DECISION-FUTURES.md`; updated PRD, thesis, research, project state, completion audit, and v0.3 index. Implementation remains gated by the complete two-rater R7 expert audit.
19 changes: 11 additions & 8 deletions docs/PROJECT_STATE.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

Last updated: 2026-08-01

## Current branch milestone — R7 corpus frozen; compiler audit pending
## Current branch milestone — R7 automatic audit passed; expert gate pending

Branch `codex/shadow-cockpit-rnd` resets the product R&D thesis around **Dual-Control Development**.

Expand Down Expand Up @@ -34,15 +34,17 @@ Branch `codex/shadow-cockpit-rnd` resets the product R&D thesis around **Dual-Co
- Protected PR #17 run `30674334938` passed `extension`, `extension-windows`, `contract`, `web`, and `jules-rnd-policy`. The Linux extension job explicitly provisioned the exact digest and passed the real Docker backend suite; Windows independently passed the deterministic contract suite. The R7 sandbox implementation gate is complete.
- `CONCEPT_LAB_CONTROLLABILITY.md` records a post-R7 category extension: compile an executable `observe → actuate → recover` human control surface, select takeover cut sets, and let a context-starved agent continue writing code from human-selected evidence and directives. Dissent cases and control dividends remain hypotheses with explicit falsifiers, not implemented features.
- The preregistered R7 collector froze 30 eligible patches from six repositories after evaluating 457 bounded eligibility records. Manifest `a4ef6cbfa48c66cb9d384bcc2834ecbfae8ff08810abfd1863b395b8aa47d149` contains 12 development and 18 held-out patches; `docs/v0.3/results/R7_CORPUS_COLLECTION.md` reports repository and first-match exclusion counts. No compiler or human outcome influenced selection.
- R7 has not passed. The recovery/probe compiler audit, three-run evidence, protected parity/adversarial runs, and two independent human ratings remain pending; R5/R6 stay gated.
- The frozen R7 automatic audit passed its preregistered automatic threshold: 17/18 held-out identities compiled and 16/18 were valid end-to-end. The frozen blind expert packet set and deterministic rating join exist, but two independent ratings and adjudication remain pending. Full R7 has not passed; R5/R6 stay gated.
- ADR-007 proposes an Executable Operator Model and shadow-control protocol. ADR-008 adds a bidirectional Intent Ledger for artifact accountability. ADR-009 adds Decision Futures and a Takeover Envelope so an on-time pre-reveal human commitment can determine a live integrated branch while agents retain implementation. Together they cover artifact accountability, demonstrated control, and real decision authority; none is implementation evidence.
- `GOAL_COMPLETION_AUDIT.md` maps the original product goal to current evidence. It explicitly records that the live agent plane, v0.3 cockpit, cold relay, delayed-transfer study, and skill-preservation claim remain incomplete.
- The readiness ledger and v0.3 cockpit do not exist yet. R0–R4.5 remain a closed reviewed-fixture mechanism and do not execute arbitrary participant or workspace code.
- No skill-retention or speed metric has been measured. Values in the PRD are predeclared R&D targets.
- A new implementation audit found five R0 ambiguities: candidate-diff identity, pre-store fixture blobs, runtime identity, check IDs, and Git object format. The normative contract closes them with structured diffs, catalog-owned blobs, standalone Node `v22.17.0`, declared test IDs, and SHA-1 Git initialization; R0a/R0b now implement and verify that complete substrate.
- A guarded Jules dispatcher and PR policy are defined as a finite R0→R4 queue. They create at most one session after a successful preflight, stop after merged R4, remain inert unless dispatch is explicitly enabled, and keep plan approval on by default. Merges remain manual because the current project tests are not an independent immutable verifier. Full scheduled continuation still requires the dispatcher workflow to be reviewed into the default branch.
- The R&D branch is published at `origin/codex/shadow-cockpit-rnd`. Its first Jules workflow run was correctly skipped because `JULES_RND_LOOP_ENABLED` is not enabled; no Jules session was created.
- A guarded Jules dispatcher and PR policy are defined as a finite R0→R4 queue. They create at most one session after a successful preflight, stop after merged R4, remain inert unless dispatch is explicitly enabled, and keep plan approval on by default. Merges remain manual because the current project tests are not an independent immutable verifier. The workflow source is now present in R&D, but GitHub schedules and manual dispatch require the workflow file on the default branch; no automatic Jules loop is active.
- The R&D branch is published at `origin/codex/shadow-cockpit-rnd`. No Jules session was created by the guarded workflow.
- `Protect main` is active: PR, conversation resolution, strict `extension`/`contract`/`web` checks, up-to-date base, deletion protection, and force-push protection are enforced with zero required approvals for the sole owner.
- `Protect R&D integration` is active on exact branch `codex/shadow-cockpit-rnd`: PR-only updates, conversation resolution, strict `extension`, `extension-windows`, `contract`, `web`, and `jules-rnd-policy` GitHub Actions checks, up-to-date base, deletion protection, and force-push protection. Codex and Jules now integrate through short-lived heads.
- Superseded `codex/v0.2-ownership-compiler` and the obsolete Jules vibe-gate branch were preserved as dated archive tags and deleted as branches. `codex/shadow-cockpit-rnd` is the only persistent coding branch; draft PR [#8](https://github.com/yava-code/PureFlow/pull/8) temporarily retains the default-branch scheduler until the owner explicitly authorizes its merge into `main`.
- Superseded `codex/v0.2-ownership-compiler` and the obsolete Jules vibe-gate branch were preserved as dated archive tags and deleted as branches. `codex/shadow-cockpit-rnd` is the only persistent coding branch. PR [#8](https://github.com/yava-code/PureFlow/pull/8) was retargeted from `main` to R&D, merged, and its short-lived head was deleted; `main` remains untouched.

The v0.1 runtime below remains released evidence and a reusable IDE shell. Its Mentor, Quiz, and Focus behavior is not the v0.3 product core.

Expand Down Expand Up @@ -135,7 +137,7 @@ The repository contains no verified evidence that the owner submitted the final
| The first live adapter is selected but no accessible Codex CLI is configured for this checkout | ADR-006 selects Codex App Server over local stdio, but the Microsoft Store packaged executable discovered here returns `Access denied` when launched from the repository shell | Keep replay R&D independent; the live spike must preflight a separately accessible, exact-version user-installed Codex CLI and fail closed when unavailable |
| R7 expert audit is not complete | Frozen held-out automatic audit passed at 16/18 end-to-end, but independent causal-relevance ratings are not yet measured | Give `docs/v0.3/results/held-out-rater-packets/` to two experienced TypeScript raters using `docs/v0.3/R7_EXPERT_RATING.md`, then adjudicate and report agreement |
| Human participants are not recruited | Takeover and delayed-transfer claims cannot be tested | Complete the technical gate, then recruit for the preregistered pilot |
| Default-branch Jules scheduler awaits explicit merge approval | Scheduled/manual continuation is not installed on `main`; draft PR #8 remains isolated and the enable variable stays off | Owner explicitly says `merge #8`; then merge through protected `main`, remove the temporary infrastructure branch, and run one guarded canary through the protected R&D branch |
| Default-branch Jules dispatcher awaits explicit authorization | The workflow exists on R&D, but GitHub will not schedule or manually dispatch it until a workflow file exists on `main`; the enable variable stays off | Create a new minimal workflow-only PR into protected `main` only after the owner explicitly authorizes that exact action, then run one guarded canary targeting R&D |

No external input blocks the repository-owned fixture R0–R4.5 mechanism in `docs/v0.3/AGENT_EXECUTION.md`. It cannot execute or judge arbitrary human/agent code and must not be represented as a general-project takeover product.

Expand All @@ -153,8 +155,9 @@ No external input blocks the repository-owned fixture R0–R4.5 mechanism in `do
## Next ordered actions

1. Obtain and adjudicate two independent blind expert ratings for packet index `54d78382b3ddbe15cba1f8153275e8149d32ddaa5192163f99ca5f43d903e8fe`.
2. Add the local readiness ledger and minimal cockpit only after the complete R7 expert gate passes.
3. Run the preregistered delayed-transfer pilot before making any skill-retention claim.
2. If full R7 passes, review ADR-007, ADR-008, and ADR-009 together, then freeze the Operator Model, Intent Ledger, Decision Future, and Takeover Envelope schemas, hash domains, projections, evidence authorities, timing policy, and selector ablations before implementation.
3. Add the local readiness ledger and minimal cockpit only after the complete R7 expert gate passes.
4. Run the preregistered delayed-transfer pilot before making any skill-retention claim.

## Recent milestone commits

Expand Down
2 changes: 2 additions & 0 deletions docs/v0.3/ADR-001-DUAL-CONTROL.md
Original file line number Diff line number Diff line change
Expand Up @@ -141,6 +141,8 @@ The router accepts the user's attention budget and readiness evidence. In v0.3 i

After recovery episodes validate the mechanism, the next router experiment is speculative live steering: fork two viable designs, collect comparable executable evidence, and let the human's choice determine the integrated branch without requiring them to type the implementation.

ADR-009 formalizes that later experiment as a Decision Future with a precommitted autonomous default, bounded integration deadline, false-fork abstention, live-influence evidence, and a Takeover Envelope. It remains gated and does not change the implemented R0–R4.5 path.

## Options considered

### Option 1: Extend the v0.1 Mentor and Focus Rep
Expand Down
Loading
Loading