Releases: ThePyProgrammer/turing
Release list
v4.8.1 — Cwd-Robust uv Hooks
What's new
Turing's generated Claude Code hooks are now robust to Claude changing directories during a session. Scaffolded hook commands point at absolute script paths, so stop and post-train hooks keep finding the intended harness files instead of accidentally resolving paths relative to whatever cwd Claude happens to hold.
The generated harness is also uv-first end to end. Hook scripts run Python through uv run python, scaffold setup uses uv sync, and the user-facing instructions now teach the same execution contract the scripts enforce.
Cwd-robust hook commands
- Generated
.claude/settings.local.jsonhook commands now use absolute script paths - Stop hooks continue to preserve convergence exit code propagation
- Post-train hooks preserve fallback
run.loghandling when invoked from a different cwd
uv-first scaffold runtime
- Added a shared hook helper for
uv run python - Missing
uvnow fails loudly instead of silently falling back to ambient Python - Scaffold setup, templates, agents, and command guidance consistently use
uv syncanduv run
Numbers
- 2041 tests
- 74 commands
- 93 scaffold scripts
- 4 implementation commits since v4.8.0
Full changelog
- fix: make Turing hooks cwd-robust and uv-first
- docs: align Turing instructions with uv execution
- chore: prepare v4.8.1 release metadata
- chore: release v4.8.1
v4.8.0 — Skills as Command Source
What's new
Turing's modern Claude Code layout is now the real source of truth. The editable command surface lives in skills/turing/, while the legacy commands/ tree is generated as a compatibility output for packaging, docs links, and transition tooling.
This release also enables model invocation for the command skills, so the /turing router can route directly to focused sub-command behavior instead of treating every sub-command as model-disabled.
Skills as command source
- Made
skills/turing/the editable command source - Generated the legacy
commands/compatibility tree from skill sources - Added
src/sync-commands-layout.jsas the primary sync script - Kept
sync:skillsas a backward-compatible alias
Runtime compatibility preserved
- Installer still deploys to
.claude/commands/turing - Public slash-command paths remain unchanged
config/commands.yamlremains canonical for command names and metadata- Package still includes both
skills/andcommands/
Model invocation contract
- Enabled model invocation for Turing command skills
- Removed stale model-disabled frontmatter assumptions
- Updated router and registry tests around direct sub-command handling
Numbers
- 2037 tests
- 74 commands
- 93 scaffold scripts
- 5 implementation commits since v4.7.0
Full changelog
- chore: generate commands compatibility from skills
- fix: install commands from skill sources
- docs: document skills as command source
- chore: enable model invocation for turing skills
- test: align turing router invocation contract
v4.7.0 — Modern Skills Layout Mirror
What's new
Turing v4.7.0 ships a modern Claude Code skills/turing/ package mirror while preserving the existing /turing and /turing:<command> runtime contract. The editable source remains commands/, the registry remains canonical, and the npm package now exposes both the proven legacy-compatible command layout and the modern SKILL.md layout conventions.
This is intentionally a staged migration: users get the modern package shape immediately, while install and verification behavior continue to land under .claude/commands/turing exactly as before.
Modern skills mirror
- Added generated
skills/turing/SKILL.mdandskills/turing/<command>/SKILL.mdmirrors for all 74 registered commands - Added
skills/turing/rules/loop-protocol.mdas the mirrored command-rule source - Added
npm run sync:skillsfor maintainers to refresh the generated mirror fromcommands/ - Added
node src/sync-skills-layout.js --checkfor CI-style drift detection
Package and docs polish
- Included
skills/in the npm package allowlist - Excluded generated Python cache/build artifacts from package dry-runs
- Updated contributor and architecture docs to explain the staged source-of-truth model
- Updated docs changelog, install output, and version badges for v4.7.0
Numbers
- 2032 tests
- 74 commands
- 93 template scripts
- 5 commits since v4.6.0
Verification
node src/sync-skills-layout.js --checkpython -m pytest tests/test_skills_layout.py tests/test_command_registry.py tests/test_manifest_consistency.py -v- temporary
HOMEinstall + verify npm pack --dry-run- full pytest suite: 2032 passed
Full changelog
feat: add skills layout sync utilitychore: package generated skills mirrortest: cover skills layout driftdocs: document staged skills layout migrationchore: release v4.7.0
v4.5.0 — Release Reliability
What's new
Turing v4.5.0 is a reliability release for the parts that make the harness usable outside this checkout: the public documentation site, npm/template installation path, scaffold verification, and Claude Code hook schema.
The theme is boring infrastructure. New projects should scaffold reliably from installed packages, generated settings should match Claude Code's current hook schema, and existing projects with legacy Stop hooks now get a doctor diagnosis and safe migration path.
Documentation site
- Published the Zensical documentation site with homepage, getting-started pages, architecture docs, complete command reference, changelog page, and philosophy section.
- Added docs deployment configuration and workflow.
- Polished README/docs prose, diagrams, homepage layout, visual styling, and command reference links.
Installer and scaffold reliability
turing-initnow installs and verifies templates reliably across npm, global command-pack, project, plugin, and node_modules layouts.- Scaffolding now documents derived placeholders and verifies generated projects without false missing-value failures.
- Added npm publish safety via
.npmignoreand removed private/planning/output paths from tracking.
Claude hook schema hardening
- Stop hooks now use Claude Code hook groups instead of legacy bare command objects.
/turing:doctornow checks.claude/settings.local.jsonhook schema./turing:doctor --fixsafely migrates known legacy bare command hooks into hook groups with backups, preserving valid hooks.
Numbers
- 2010 tests (+24 since v4.4.0)
- 74 commands
- 93 scripts
- 68 commits since v4.4.0
Full changelog
- Published the new documentation site.
- Added Zensical docs configuration and deployment workflow.
- Fixed Turing init template discovery and verification across installed layouts.
- Documented derived scaffold placeholders.
- Fixed Stop hook generation for Claude Code's current schema.
- Added doctor validation and healing for legacy malformed hook settings.
- Polished README, philosophy pages, homepage layout, diagrams, logo, and visual styling.
v4.4.0 — Operational Intelligence (Final Phase)
What's new in v4.4.0
The final phase. All 29 phases of the Turing roadmap are now complete — 84 implementation items, from the first hypothesis queue to the last research planner. Turing is now self-aware enough to diagnose its own failures, prescribe its own repairs, and plan its own campaigns.
/turing:postmortem — Automated Failure Postmortem
When experiments stop improving, stop guessing and start diagnosing. Analyzes the failure streak and scores 5 root cause hypotheses: search space exhaustion (micro-tuning params that don't matter), systematic config error (all experiments share a bad setting), data issue (all model types fail identically), metric ceiling (near theoretical maximum), and noise floor (improvements within seed variance). Each diagnosis includes scored evidence and ranked actionable recommendations that link to the right /turing: command to fix it.
/turing:doctor — Harness Self-Diagnosis
Is Turing healthy? Seven checks: Python environment, dependency imports, config.yaml validity, experiment log integrity (corrupt line detection), script AST parsing, disk space, and git state (uncommitted changes to evaluation-critical files). --fix mode auto-repairs safe issues — removes corrupt log lines with automatic backup, reclaims disk space. Reports HEALTHY / DEGRADED / UNHEALTHY with a pass/warn/fail score.
/turing:plan — Research Planning Assistant
Design the next N experiments strategically, not randomly. Computes per-family ROI from experiment history, adjusts strategy priorities based on project state and goal ("maximize F1 for production"), and allocates budget across 5 strategies: feature engineering (highest-ROI direction), model search, ensemble & composition, production readiness (calibration + compression), and verification. Generates a phased plan with specific experiment descriptions.
Integration
- Postmortem diagnosis and active research plan appear in
/turing:briefresearch briefing - All commands registered in router, installer, verifier, and scaffold
Numbers
| Metric | v4.3.0 | v4.4.0 | Delta |
|---|---|---|---|
| Tests | 1876 | 1986 | +110 |
| Commands | 71 | 74 | +3 |
| Scripts | 90 | 93 | +3 |
Roadmap Complete
All 29 phases (84 implementation items) from v1.0.0 through v4.4.0 are done. 74 commands, 93 scripts, 1986 tests, 2 specialized agents, 16 ADRs.
v4.3.0 — Model Lifecycle
What's new in v4.3.0
From "best experiment" to governed, versioned, production-tracked model. Turing now manages the full model lifecycle — incremental updates without retraining, a formal registry with promotion gates, and enhanced model cards with fairness analysis.
/turing:update — Incremental Model Update
Add new data to an existing model without starting from scratch. Model-specific strategies: continued boosting for XGBoost/LightGBM (add N rounds), fine-tuning with replay buffer for neural networks (configurable old-data ratio to prevent catastrophic forgetting), and partial_fit/warm_start for scikit-learn. Automatically checks for forgetting — if accuracy on old data degrades beyond tolerance, warns and offers rollback.
/turing:registry — Model Registry
Track which model is production, staging, candidate, or archived. Four-stage lifecycle with automated promotion gates:
- candidate → staging: requires regression check PASS + seed study
- staging → production: requires audit PASS + calibration check
- Demotion and archiving with reason tracking
- Full promotion/demotion history with timestamps and gate results
Enhanced /turing:card
- New
--include fairnessflag: demographic parity and equal opportunity metrics across protected groups - Registry status section: shows current stage, version, gates passed
- Both integrated automatically when data is available
Integration
- Registry status and update history appear in
/turing:briefresearch briefing - All commands registered in router, installer, verifier, and scaffold
Numbers
| Metric | v4.2.0 | v4.3.0 | Delta |
|---|---|---|---|
| Tests | 1740 | 1876 | +136 |
| Commands | 69 | 71 | +2 |
| Scripts | 88 | 90 | +2 |
One phase remaining: 29 (Operational Intelligence).
v4.2.0 — What-If Analysis
What's new in v4.2.0
Answer hypotheticals without running experiments. Turing now predicts outcomes from existing data — scaling laws, ablation results, sensitivity curves, ensemble correlations — and tells you which experiments are worth running before you spend the compute.
/turing:whatif — Counterfactual Experiment Simulation
Ask "what if I had 2x more data?" or "what if I removed class 3?" and get an estimate with confidence. Routes questions to 7 estimators automatically: scaling law extrapolation, ablation study data, sensitivity interpolation, ensemble correlation, pruning sweep interpolation, pipeline stitch estimation, and budget allocation. Each answer includes confidence level (HIGH/MED/LOW) and the command to verify.
/turing:counterfactual — Input-Level Counterfactual Explanations
For a given prediction, find the smallest input change that would flip the outcome. Two methods: greedy feature perturbation (change one feature at a time) and prototype-based search (find nearest training sample from the target class). Handles numeric and categorical features, supports batch mode for all misclassified samples. Useful for debugging individual predictions and regulatory explanations.
/turing:simulate — Experiment Outcome Prediction
Before running a sweep of N experiments, predict the likely outcome distribution. Builds a weighted k-NN surrogate model from experiment history, applies a novelty penalty for configs far from the training distribution, and ranks proposed configs by predicted metric. Auto-filters: only queue experiments predicted to beat the current best. Budget savings typically >50%.
Integration
- What-if and simulation results integrated into
/turing:briefresearch briefing - All three commands registered in router, installer, verifier, and scaffold
Numbers
| Metric | v4.1.0 | v4.2.0 | Delta |
|---|---|---|---|
| Tests | 1576 | 1740 | +164 |
| Commands | 66 | 69 | +3 |
| Scripts | 85 | 88 | +3 |
Two phases remaining: 28 (Model Lifecycle), 29 (Operational Intelligence).
v4.1.0 — Collaboration
What's new in v4.1.0
Research is a team sport. Make results shareable and reviewable.
/turing:onboard — Project Onboarding
Generate a 5-minute walkthrough for new collaborators. Summarizes task, what's been tried (grouped by family with productivity status), key decisions, where the project is heading, and how to get started. Audience adaptation: researcher (full technical detail), engineer (deployment path), stakeholder (business metrics).
/turing:share — Experiment Packaging
Package experiments into self-contained portable archives with config, metrics, seed studies, annotations, and reproduction instructions. Generates manifest.yaml and README.md. Optional includes: model artifacts, code snapshots, figures, data hashes.
/turing:review — Peer Review Simulation
Simulate a conference reviewer before submission. 10 checks (baselines, error bars, ablation, SOTA, calibration, compute cost, diversity, leakage, reproducibility). Venue calibration (NeurIPS/ICML/general). --harsh mode. Each weakness links to the /turing: command that fixes it. Scores 1-10 with verdict.
Numbers
| Metric | v4.0.0 | v4.1.0 | Delta |
|---|---|---|---|
| Tests | 1566 | 1576 | +10 |
| Commands | 63 | 66 | +3 |
| Scripts | 82 | 85 | +3 |
Three phases remaining: 27 (What-If Analysis), 28 (Model Lifecycle), 29 (Operational Intelligence).
v4.0.0 — Research Communication
What's new in v4.0.0
The v4.0 milestone. The roadmap is complete.
Turing goes from research tool to research-to-communication pipeline. Every result becomes a shareable artifact — citations tracked, presentations generated, progress communicated. With all 25 phases implemented across 72 features, Turing is a complete autonomous ML research harness from first hypothesis to final presentation.
/turing:cite — Citation & Attribution Manager
Track which papers, codebases, datasets, and methods influenced each experiment. Add citations with DOI/arXiv links, audit for missing attributions (scans experiment configs against 22 common ML methods), and generate BibTeX. Stored in experiments/citations.yaml.
/turing:present — Presentation Figure Generation
Generate presentation-ready figure specifications from experiment data: training curves with best-so-far trajectory, model family comparison bars, ablation delta tables, Pareto scatter with frontier line, and sensitivity heatmaps. Three style presets (light/dark/poster) with customizable palettes. Output as structured JSON specs in paper/figures/.
/turing:changelog — Model Changelog Generation
Auto-generate a human-readable changelog from experiment history. Detects version boundaries (significant metric jumps), groups improvements within each version, and formats as narrative with deltas. Two audiences: technical (experiment IDs, configs) and stakeholder (plain English, percentages, no jargon). Output to paper/CHANGELOG.md.
The Complete Journey: v1.0.0 → v4.0.0
| Version | Phase | Commands |
|---|---|---|
| v1.0.0 | 1–9: Core loop, hypotheses, novelty, statistics, families, tree-search | 14 |
| v1.3.0 | 10: Statistical rigor (seed, reproduce) | 19 |
| v1.4.0 | 11: Experiment intelligence (diagnose, ablate, frontier) | 22 |
| v1.5.0 | 12: Performance (profile, checkpoint) | 24 |
| v2.0.0 | 13: Deployment bridge (export) | 25 |
| v2.1.0 | 14: Research workflow (lit, paper) | 27 |
| v2.2.0 | 15: Orchestration (queue, retry, fork) | 30 |
| v2.3.0 | 16: Deep analysis (diff, watch, regress) | 33 |
| v2.4.0 | 17: Model composition (ensemble, stitch, warm) | 36 |
| v2.5.0 | 18: Scaling & efficiency (scale, budget, distill) | 39 |
| v3.0.0 | 19: Meta-intelligence (transfer, audit) | 41 |
| v3.1.0 | 20: Pre-training intelligence (sanity, baseline, leak) | 44 |
| v3.2.0 | 21: Model debugging (xray, sensitivity, calibrate) | 47 |
| v3.3.0 | 22: Feature & training intelligence (feature, curriculum) | 49 |
| v3.4.0 | 23: Model surgery (prune, quantize, merge, surgery) | 53 |
| v3.5.0 | 24: Experiment archaeology (trend, flashback, archive, annotate, search, template, replay) | 60 |
| v4.0.0 | 25: Research communication (cite, present, changelog) | 63 |
Numbers
| Metric | v1.0.0 | v4.0.0 | Growth |
|---|---|---|---|
| Tests | 257 | 1566 | +1309 (6.1x) |
| Commands | 14 | 63 | +49 (4.5x) |
| Scripts | 23 | 82 | +59 (3.6x) |
| Phases | 9 | 25 | +16 |
| Versions | 1 | 17 | +16 |
Full changelog
feat:citation_manager.py — citation tracking with add/list/check/bib, BibTeX generationfeat:generate_figures.py — presentation figure specs for 5 chart types with 3 stylesfeat:generate_changelog.py — model changelog with version detection and audience adaptationfeat:/turing:cite, /turing:present, /turing:changelog command skillstest:19 new tests across 3 test filesdocs:All 72 roadmap items marked DONE
v3.5.0 — Experiment Archaeology
What's new in v3.5.0
Manage a research project that spans months, not just sessions.
Phase 24 is the largest release yet — seven commands that transform Turing from a single-session tool into a long-term research companion. It adds strategic trend analysis for 100+ experiment projects, instant context restoration after time away, disk space management, human annotations, natural language search, reusable templates, and experiment replay to test if old ideas work better now.
/turing:trend — Long-Term Trend Analysis
Strategic view over your full experiment history: improvement velocity per time window, family ROI ranking (which experiment families are productive vs exhausted), diminishing returns detection via log-curve fitting, and phase transition detection. The "state of the research" report that's invisible in individual experiment logs.
/turing:flashback — Session Context Restoration
Come back after a week and start working in 10 seconds. Shows current best, last session's experiments with outcomes, pending hypotheses, annotations, budget status, and suggested next action. Eliminates the 30-minute log-reading ritual.
/turing:archive — Experiment Lifecycle Cleanup
After 200+ experiments, reclaim gigabytes by compressing old artifacts while preserving a queryable summary index. Protects Pareto-optimal, best, and recent experiments. Supports --dry-run to preview before committing.
/turing:annotate — Retrospective Annotations
Add human notes to any experiment: "This only worked because the data was pre-sorted." Tags for categorization, search by content or tag. The knowledge that automated metrics can't capture.
/turing:search — Natural Language Experiment Search
Query 200+ experiments with natural language plus structured filters: "LightGBM high accuracy" --filter "accuracy>0.85". Combines keyword ranking with metric/status/family/date filters.
/turing:template — Experiment Template Library
Save winning configs as reusable templates that persist across projects at ~/.turing/templates/. Strip project-specific paths automatically. Apply to new projects with one command.
/turing:replay — Experiment Replay
Re-run a historical experiment with current infrastructure. Different from /turing:reproduce (which verifies the same result) — replay tests whether an old approach does better now because data, preprocessing, or libraries have improved.
Numbers
| Metric | v3.4.0 | v3.5.0 | Delta |
|---|---|---|---|
| Tests | 1494 | 1547 | +53 |
| Commands | 53 | 60 | +7 |
| Scripts | 72 | 79 | +7 |
| Commits | — | 26 | — |
Full changelog
feat:7 new scripts: trend_analysis, session_flashback, experiment_archive, experiment_annotations, experiment_search, experiment_templates, experiment_replayfeat:7 new command skills: /turing:trend, /turing:flashback, /turing:archive, /turing:annotate, /turing:search, /turing:template, /turing:replaytest:53 new tests across 7 test filesdocs:ROADMAP Phase 24 (items 63-69) marked DONE, README updated
One phase remaining: Phase 25 (Research Communication, v4.0.0).