Repository navigation
Releases: macayaven/agent-harness-path
Release list
v0.6.0-rc.2 — Fixture corrections
This course testing release corrects misleading fixtures and includes newly recorded offline comparisons. It is marked Latest at the maintainer's request; learner acceptance testing remains open.
What changed
- Corrected S04 ticket-generation controls and preserved parse/shape/meaning evidence across retries. Endpoint failures now stop the batch without discarding earlier results; truncated completions are not accepted.
- Fixed misleading S02, S05, S07 and S09 controls, including empty rejection/omission evidence and a repair prompt missing required fields.
- Added exact menu names to protected-shift context and clarified scope/billing house rules.
- Re-recorded all ten governed lab tapes and retained the nine deliberate naïve baselines. The lab set contains 19 files / 36 responses.
- Added four optional real-model notebook comparisons (S04/S05/S07/S09): 16 responses that replay offline. Authored defaults, predict-first work and learner attempt skeletons remain intact.
- Includes the README contributor/model/hardware acknowledgments and code/content licensing added since rc.1.
Verified
Python 3.11 and 3.12: 151 content contracts, 29 lab contracts, all twelve notebooks, the lab app, all exact recorded replays, canonical marimo, generated HTML, links and SOTA sources. Local verification ran with the socket guard. Independent review approved the changes, and both required GitHub checks passed before PR #15 was merged. Fresh synthetic recordings used locally served NVIDIA Nemotron 3.5 Lightning 30B A3B on DGX Spark. The review documents what each example actually establishes; a 9/9 lab score is not a general quality claim.
Detailed fixture review and provenance
Start testing
Check out v0.6.0-rc.2, run uv sync --frozen, then start with sessions/index.html. No provider, API key or model hardware is needed for the default course. After making your predictions and attempts, uv run python tools/record_fixtures.py replays the optional model comparisons.
Remaining acceptance work
Copilot tutor instructions remain guidance, not an enforceable answer boundary: it may reveal too much. Close chat during independent attempts, then discuss your own work. The cross-editor/Codespaces behavior bank and real learner pilot remain pending; this release does not claim broad public readiness. The previously documented unfinished context-probe accounting limitation also remains; this fixture repair does not change that counter. Jev integration remains optional future work. See companion guidance.
v0.6.0-rc.1 — Course testing release
This release candidate is available for maintainer testing and early testers. It is marked Latest at the maintainer's request; that designation does not mean the pending learner acceptance work has passed.
What changed
- Restored consent checks through the course artifact; corrected context measurement, grounded grading and budget evidence.
- Protected independent predictions/labels and answer reveals, retained schema-attempt evidence, and cleared stale notebook results.
- Fixed lab Run selections and refreshed all 19 recordings. Strict replay reproduces the observed naive 2/9 and reference 9/9 baseline.
- Added the selected AHP Tutor role and mirrored Cursor/Copilot policy, with clearer editor and offline onboarding.
- Repaired lesson navigation, notebook text/output/figure rendering and narrow-screen tables; updated sourced curriculum context and prerequisites.
Verified
132 course contracts and 20 lab contracts passed on each of Python 3.11 and 3.12. All 12 toy scripts, the lab app, strict replay, canonical formatting, generated HTML, 140 relative references and 99 sourced SOTA rows passed. Live endpoint/doctor, all 12 notebook scripts and synthetic S12 readiness/reveal/reset gates were exercised. GitHub CI and independent review were required before merging.
Known limitations and testing needed
Copilot can still reveal too much. In local VS Code 1.138.0 / Copilot Chat 0.66.0, AHP Tutor with Auto (GPT-5.6 Luna) refused code but then described enough steps to reconstruct the S01 answer. The focused policy clarification did not resolve this. Read/search-only tools do not enforce an answer boundary in chat. Close chat during independent attempts, then use it to discuss your own work; see companion guidance.
The full cross-editor/Codespaces behavior bank and real learner pilot remain pending. A context probe interrupted during tool continuation can also be omitted from the attempted-probe counter; unfinished/context-limit outcomes remain visible. This RC does not claim broad public readiness. Jev integration remains optional future work.
Start testing
Check out v0.6.0-rc.1, run uv sync --frozen, and start with sessions/index.html. Notebooks default to the offline stub; hard-path labs default to replay. Start without keys or a live endpoint. Test onboarding, concept-to-lab transfer, tutor boundaries and any confusing evidence before enabling a live model.
v0.5.0
v0.5.0 supersedes v0.4.0. Hard-path labs re-skinned trivia to cafe (one domain everywhere), lessons unlink source files to plain code, Codespaces verified end to end. CI stays zero network, zero keys, zero cost.
Highlights
- One domain everywhere: labs/trivia_host is now labs/cafe_host with propose_order, pull_item, settle_item, close_shift over a 12-item menu; all 12 lab.md protocols, 14 companions, and docs follow. Cassettes re-recorded.
- Lessons: toy.py / lab.md references build to plain code, not links (preview servers hand source out as plain text); lesson-to-lesson bridges still navigate, covered by a build regression test.
- Codespaces proven by a 13-step live checklist: pinned env, offline suite green in-container, marimo notebooks open by default, lessons render in embedded preview, Copilot loads the companion instructions.
- S12 judge calibration corpus fixed, twelve judge calls priced on the routing ledger; S05 promise and lab orientation swept.
Take it
Open the folder in Cursor, read docs/COMPANION.md, start sessions/s01-agent-loop/lesson.html with @sessions/s01-agent-loop/companion.md in chat.
Full notes: CHANGELOG.
v0.4.0
v0.4.0 supersedes v0.3.0. One directory per session, offline-by-default notebooks with static figures, English throughout, and the stale per-session videos removed. CI stays zero network, zero keys, zero cost.
Highlights
- One directory per session:
sessions/sNN-slug/holdslesson.md(builds tolesson.htmlin place),toy.py,lab.md,companion.md, andpublic/diagrams/. get_client()is offline by default: unsetCOURSE_MODEreturns the deterministic stub everywhere;COURSE_MODE=liveopts into the learner's endpoint. Notebooks embed committed SVGs (nomo.mermaid); the static-diagram pipeline is pinned with a freshness receipt.- English everywhere: café menu, customer lines, prompts, refusal texts, detection policy, and judge corpus. Per-session Video Overviews removed (they lagged the lessons); the S00 course overview stays as the single optional preview.
- Repo hygiene: superseded maps, plans, and release tooling removed; study protocols rewritten for the Cursor companion path; live transport unified (
labs/client.pyshares_post_chat_completionswithcafe/model.py,--replayoutput byte-identical).
Take it
git clone https://github.com/macayaven/agent-harness-path.git
cd agent-harness-path
uv sync --frozenOpen the folder in Cursor, read docs/COMPANION.md, start sessions/s01-agent-loop/lesson.html with @sessions/s01-agent-loop/companion.md in chat.
Full notes: CHANGELOG.
The Agent Harness Path v0.3.0
v0.3.0 supersedes v0.2.0. Take the course from this clone in Cursor, with session bridges and a complete trivia host. CourseWeave packaging is removed.
Highlights
- Keep v0.2.0 lessons, notebooks, labs contracts, study overlay, and CI.
- Ship a complete
labs/trivia_host/so--replayruns without filling stubs. - Add
bridges/s01.md–s14.md,.cursor/rules/ahp-companion.mdc, anddocs/COMPANION.md(local OpenAI-compatible tutor = override base URL + key; lab--liveuses separate shellOPENAI_*). lessons/*.htmlis the reader;lessons/src/is authoring source. S13/S14 remain unaided.
Take it
git clone https://github.com/macayaven/agent-harness-path.git
cd agent-harness-path
uv sync --frozenOpen the folder in Cursor, read docs/COMPANION.md, start lessons/S01-agent-loop.html with @bridges/s01.md.
Full notes: CHANGELOG.
The Agent Harness Path v0.2.0 — full-course student edition
The Agent Harness Path v0.2.0 extends the existing course in this same repository with the full guided student edition. The course retains fourteen lessons, twelve runnable notebooks, and S13/S14 as authored optional practical protocols. The course remains usable on its native route without CourseWeave or a provider key.
Start with the course README, then lessons/index.html in your local checkout. The README covers setup, the reading → prediction/attempt/observation → self-check route, optional labs/videos, assistant boundaries and the course feedback form.
The complete changelog and versioned release notes describe the changes. Contributor guidance, reproducible checks, security reporting and release documentation are included.
Verified source and download
This release is tagged at c90aa1d84d3b1b1c21a5d263de03c4cb198f55c6. Both main CI jobs passed. A fresh public clone followed the documented uv sync --frozen path and executed all twelve notebooks under Python 3.14.7 on macOS without unhandled errors; all 154 code cells and output-free source notebooks were preserved.
The public agent-harness-path-v0.2.0.tar.gz download matched SHA256SUMS and all 198 regular source files. Its artifact bytes remained unchanged through prerelease verification and promotion. Earlier tags are retained.
Optional guided application
The separately versioned CourseWeave v0.2.0 release owns the optional macOS student bundle. Its exact public download passed all fourteen rendered lessons, twelve executed/saved notebooks, 35 guided checks/105 options, native self-checks/navigation, protocol restrictions, continuous conversation and explicit sharing, restart/export/reset, privacy and process cleanup. See its per-module acceptance report.
S13/S14 require real prerequisites and learner activity; software interaction checks used synthetic records and do not claim completion of an unaided audit, holdout acceptance or human pilot. Teacher/author workflows, hosted operation and learning gains are outside this student acceptance. Three earlier real-provider text observations are distinguished from the synthetic public-download checks, including one S11 wording limitation. GitHub rendered page/form inspection was blocked by administrator browser policy; source/template checks are reported separately. No invented feedback was submitted.
License
The course keeps its Apache-2.0 software and CC BY 4.0 educational-content licensing. CourseWeave is a separate source-available application under PolyForm Shield 1.0.0; its licensing guide explains evaluation and competing-product boundaries. Existing course and third-party grants are unchanged.