Skip to content

Releases: ThePyProgrammer/turing

v4.8.1 — Cwd-Robust uv Hooks

Choose a tag to compare

@ThePyProgrammer ThePyProgrammer released this 10 May 16:58

What's new

Turing's generated Claude Code hooks are now robust to Claude changing directories during a session. Scaffolded hook commands point at absolute script paths, so stop and post-train hooks keep finding the intended harness files instead of accidentally resolving paths relative to whatever cwd Claude happens to hold.

The generated harness is also uv-first end to end. Hook scripts run Python through uv run python, scaffold setup uses uv sync, and the user-facing instructions now teach the same execution contract the scripts enforce.

Cwd-robust hook commands

  • Generated .claude/settings.local.json hook commands now use absolute script paths
  • Stop hooks continue to preserve convergence exit code propagation
  • Post-train hooks preserve fallback run.log handling when invoked from a different cwd

uv-first scaffold runtime

  • Added a shared hook helper for uv run python
  • Missing uv now fails loudly instead of silently falling back to ambient Python
  • Scaffold setup, templates, agents, and command guidance consistently use uv sync and uv run

Numbers

  • 2041 tests
  • 74 commands
  • 93 scaffold scripts
  • 4 implementation commits since v4.8.0

Full changelog

  • fix: make Turing hooks cwd-robust and uv-first
  • docs: align Turing instructions with uv execution
  • chore: prepare v4.8.1 release metadata
  • chore: release v4.8.1

v4.8.0 — Skills as Command Source

Choose a tag to compare

@ThePyProgrammer ThePyProgrammer released this 10 May 15:33

What's new

Turing's modern Claude Code layout is now the real source of truth. The editable command surface lives in skills/turing/, while the legacy commands/ tree is generated as a compatibility output for packaging, docs links, and transition tooling.

This release also enables model invocation for the command skills, so the /turing router can route directly to focused sub-command behavior instead of treating every sub-command as model-disabled.

Skills as command source

  • Made skills/turing/ the editable command source
  • Generated the legacy commands/ compatibility tree from skill sources
  • Added src/sync-commands-layout.js as the primary sync script
  • Kept sync:skills as a backward-compatible alias

Runtime compatibility preserved

  • Installer still deploys to .claude/commands/turing
  • Public slash-command paths remain unchanged
  • config/commands.yaml remains canonical for command names and metadata
  • Package still includes both skills/ and commands/

Model invocation contract

  • Enabled model invocation for Turing command skills
  • Removed stale model-disabled frontmatter assumptions
  • Updated router and registry tests around direct sub-command handling

Numbers

  • 2037 tests
  • 74 commands
  • 93 scaffold scripts
  • 5 implementation commits since v4.7.0

Full changelog

  • chore: generate commands compatibility from skills
  • fix: install commands from skill sources
  • docs: document skills as command source
  • chore: enable model invocation for turing skills
  • test: align turing router invocation contract

v4.7.0 — Modern Skills Layout Mirror

Choose a tag to compare

@ThePyProgrammer ThePyProgrammer released this 10 May 14:50

What's new

Turing v4.7.0 ships a modern Claude Code skills/turing/ package mirror while preserving the existing /turing and /turing:<command> runtime contract. The editable source remains commands/, the registry remains canonical, and the npm package now exposes both the proven legacy-compatible command layout and the modern SKILL.md layout conventions.

This is intentionally a staged migration: users get the modern package shape immediately, while install and verification behavior continue to land under .claude/commands/turing exactly as before.

Modern skills mirror

  • Added generated skills/turing/SKILL.md and skills/turing/<command>/SKILL.md mirrors for all 74 registered commands
  • Added skills/turing/rules/loop-protocol.md as the mirrored command-rule source
  • Added npm run sync:skills for maintainers to refresh the generated mirror from commands/
  • Added node src/sync-skills-layout.js --check for CI-style drift detection

Package and docs polish

  • Included skills/ in the npm package allowlist
  • Excluded generated Python cache/build artifacts from package dry-runs
  • Updated contributor and architecture docs to explain the staged source-of-truth model
  • Updated docs changelog, install output, and version badges for v4.7.0

Numbers

  • 2032 tests
  • 74 commands
  • 93 template scripts
  • 5 commits since v4.6.0

Verification

  • node src/sync-skills-layout.js --check
  • python -m pytest tests/test_skills_layout.py tests/test_command_registry.py tests/test_manifest_consistency.py -v
  • temporary HOME install + verify
  • npm pack --dry-run
  • full pytest suite: 2032 passed

Full changelog

  • feat: add skills layout sync utility
  • chore: package generated skills mirror
  • test: cover skills layout drift
  • docs: document staged skills layout migration
  • chore: release v4.7.0

v4.5.0 — Release Reliability

Choose a tag to compare

@ThePyProgrammer ThePyProgrammer released this 10 May 07:56

What's new

Turing v4.5.0 is a reliability release for the parts that make the harness usable outside this checkout: the public documentation site, npm/template installation path, scaffold verification, and Claude Code hook schema.

The theme is boring infrastructure. New projects should scaffold reliably from installed packages, generated settings should match Claude Code's current hook schema, and existing projects with legacy Stop hooks now get a doctor diagnosis and safe migration path.

Documentation site

  • Published the Zensical documentation site with homepage, getting-started pages, architecture docs, complete command reference, changelog page, and philosophy section.
  • Added docs deployment configuration and workflow.
  • Polished README/docs prose, diagrams, homepage layout, visual styling, and command reference links.

Installer and scaffold reliability

  • turing-init now installs and verifies templates reliably across npm, global command-pack, project, plugin, and node_modules layouts.
  • Scaffolding now documents derived placeholders and verifies generated projects without false missing-value failures.
  • Added npm publish safety via .npmignore and removed private/planning/output paths from tracking.

Claude hook schema hardening

  • Stop hooks now use Claude Code hook groups instead of legacy bare command objects.
  • /turing:doctor now checks .claude/settings.local.json hook schema.
  • /turing:doctor --fix safely migrates known legacy bare command hooks into hook groups with backups, preserving valid hooks.

Numbers

  • 2010 tests (+24 since v4.4.0)
  • 74 commands
  • 93 scripts
  • 68 commits since v4.4.0

Full changelog

  • Published the new documentation site.
  • Added Zensical docs configuration and deployment workflow.
  • Fixed Turing init template discovery and verification across installed layouts.
  • Documented derived scaffold placeholders.
  • Fixed Stop hook generation for Claude Code's current schema.
  • Added doctor validation and healing for legacy malformed hook settings.
  • Polished README, philosophy pages, homepage layout, diagrams, logo, and visual styling.

v4.4.0 — Operational Intelligence (Final Phase)

Choose a tag to compare

@ThePyProgrammer ThePyProgrammer released this 02 Apr 06:21

What's new in v4.4.0

The final phase. All 29 phases of the Turing roadmap are now complete — 84 implementation items, from the first hypothesis queue to the last research planner. Turing is now self-aware enough to diagnose its own failures, prescribe its own repairs, and plan its own campaigns.

/turing:postmortem — Automated Failure Postmortem

When experiments stop improving, stop guessing and start diagnosing. Analyzes the failure streak and scores 5 root cause hypotheses: search space exhaustion (micro-tuning params that don't matter), systematic config error (all experiments share a bad setting), data issue (all model types fail identically), metric ceiling (near theoretical maximum), and noise floor (improvements within seed variance). Each diagnosis includes scored evidence and ranked actionable recommendations that link to the right /turing: command to fix it.

/turing:doctor — Harness Self-Diagnosis

Is Turing healthy? Seven checks: Python environment, dependency imports, config.yaml validity, experiment log integrity (corrupt line detection), script AST parsing, disk space, and git state (uncommitted changes to evaluation-critical files). --fix mode auto-repairs safe issues — removes corrupt log lines with automatic backup, reclaims disk space. Reports HEALTHY / DEGRADED / UNHEALTHY with a pass/warn/fail score.

/turing:plan — Research Planning Assistant

Design the next N experiments strategically, not randomly. Computes per-family ROI from experiment history, adjusts strategy priorities based on project state and goal ("maximize F1 for production"), and allocates budget across 5 strategies: feature engineering (highest-ROI direction), model search, ensemble & composition, production readiness (calibration + compression), and verification. Generates a phased plan with specific experiment descriptions.

Integration

  • Postmortem diagnosis and active research plan appear in /turing:brief research briefing
  • All commands registered in router, installer, verifier, and scaffold

Numbers

Metric v4.3.0 v4.4.0 Delta
Tests 1876 1986 +110
Commands 71 74 +3
Scripts 90 93 +3

Roadmap Complete

All 29 phases (84 implementation items) from v1.0.0 through v4.4.0 are done. 74 commands, 93 scripts, 1986 tests, 2 specialized agents, 16 ADRs.

v4.3.0 — Model Lifecycle

Choose a tag to compare

@ThePyProgrammer ThePyProgrammer released this 01 Apr 11:34

What's new in v4.3.0

From "best experiment" to governed, versioned, production-tracked model. Turing now manages the full model lifecycle — incremental updates without retraining, a formal registry with promotion gates, and enhanced model cards with fairness analysis.

/turing:update — Incremental Model Update

Add new data to an existing model without starting from scratch. Model-specific strategies: continued boosting for XGBoost/LightGBM (add N rounds), fine-tuning with replay buffer for neural networks (configurable old-data ratio to prevent catastrophic forgetting), and partial_fit/warm_start for scikit-learn. Automatically checks for forgetting — if accuracy on old data degrades beyond tolerance, warns and offers rollback.

/turing:registry — Model Registry

Track which model is production, staging, candidate, or archived. Four-stage lifecycle with automated promotion gates:

  • candidate → staging: requires regression check PASS + seed study
  • staging → production: requires audit PASS + calibration check
  • Demotion and archiving with reason tracking
  • Full promotion/demotion history with timestamps and gate results

Enhanced /turing:card

  • New --include fairness flag: demographic parity and equal opportunity metrics across protected groups
  • Registry status section: shows current stage, version, gates passed
  • Both integrated automatically when data is available

Integration

  • Registry status and update history appear in /turing:brief research briefing
  • All commands registered in router, installer, verifier, and scaffold

Numbers

Metric v4.2.0 v4.3.0 Delta
Tests 1740 1876 +136
Commands 69 71 +2
Scripts 88 90 +2

One phase remaining: 29 (Operational Intelligence).

v4.2.0 — What-If Analysis

Choose a tag to compare

@ThePyProgrammer ThePyProgrammer released this 01 Apr 11:15

What's new in v4.2.0

Answer hypotheticals without running experiments. Turing now predicts outcomes from existing data — scaling laws, ablation results, sensitivity curves, ensemble correlations — and tells you which experiments are worth running before you spend the compute.

/turing:whatif — Counterfactual Experiment Simulation

Ask "what if I had 2x more data?" or "what if I removed class 3?" and get an estimate with confidence. Routes questions to 7 estimators automatically: scaling law extrapolation, ablation study data, sensitivity interpolation, ensemble correlation, pruning sweep interpolation, pipeline stitch estimation, and budget allocation. Each answer includes confidence level (HIGH/MED/LOW) and the command to verify.

/turing:counterfactual — Input-Level Counterfactual Explanations

For a given prediction, find the smallest input change that would flip the outcome. Two methods: greedy feature perturbation (change one feature at a time) and prototype-based search (find nearest training sample from the target class). Handles numeric and categorical features, supports batch mode for all misclassified samples. Useful for debugging individual predictions and regulatory explanations.

/turing:simulate — Experiment Outcome Prediction

Before running a sweep of N experiments, predict the likely outcome distribution. Builds a weighted k-NN surrogate model from experiment history, applies a novelty penalty for configs far from the training distribution, and ranks proposed configs by predicted metric. Auto-filters: only queue experiments predicted to beat the current best. Budget savings typically >50%.

Integration

  • What-if and simulation results integrated into /turing:brief research briefing
  • All three commands registered in router, installer, verifier, and scaffold

Numbers

Metric v4.1.0 v4.2.0 Delta
Tests 1576 1740 +164
Commands 66 69 +3
Scripts 85 88 +3

Two phases remaining: 28 (Model Lifecycle), 29 (Operational Intelligence).

v4.1.0 — Collaboration

Choose a tag to compare

@ThePyProgrammer ThePyProgrammer released this 01 Apr 08:38

What's new in v4.1.0

Research is a team sport. Make results shareable and reviewable.

/turing:onboard — Project Onboarding

Generate a 5-minute walkthrough for new collaborators. Summarizes task, what's been tried (grouped by family with productivity status), key decisions, where the project is heading, and how to get started. Audience adaptation: researcher (full technical detail), engineer (deployment path), stakeholder (business metrics).

/turing:share — Experiment Packaging

Package experiments into self-contained portable archives with config, metrics, seed studies, annotations, and reproduction instructions. Generates manifest.yaml and README.md. Optional includes: model artifacts, code snapshots, figures, data hashes.

/turing:review — Peer Review Simulation

Simulate a conference reviewer before submission. 10 checks (baselines, error bars, ablation, SOTA, calibration, compute cost, diversity, leakage, reproducibility). Venue calibration (NeurIPS/ICML/general). --harsh mode. Each weakness links to the /turing: command that fixes it. Scores 1-10 with verdict.

Numbers

Metric v4.0.0 v4.1.0 Delta
Tests 1566 1576 +10
Commands 63 66 +3
Scripts 82 85 +3

Three phases remaining: 27 (What-If Analysis), 28 (Model Lifecycle), 29 (Operational Intelligence).

v4.0.0 — Research Communication

Choose a tag to compare

@ThePyProgrammer ThePyProgrammer released this 01 Apr 08:06

What's new in v4.0.0

The v4.0 milestone. The roadmap is complete.

Turing goes from research tool to research-to-communication pipeline. Every result becomes a shareable artifact — citations tracked, presentations generated, progress communicated. With all 25 phases implemented across 72 features, Turing is a complete autonomous ML research harness from first hypothesis to final presentation.

/turing:cite — Citation & Attribution Manager

Track which papers, codebases, datasets, and methods influenced each experiment. Add citations with DOI/arXiv links, audit for missing attributions (scans experiment configs against 22 common ML methods), and generate BibTeX. Stored in experiments/citations.yaml.

/turing:present — Presentation Figure Generation

Generate presentation-ready figure specifications from experiment data: training curves with best-so-far trajectory, model family comparison bars, ablation delta tables, Pareto scatter with frontier line, and sensitivity heatmaps. Three style presets (light/dark/poster) with customizable palettes. Output as structured JSON specs in paper/figures/.

/turing:changelog — Model Changelog Generation

Auto-generate a human-readable changelog from experiment history. Detects version boundaries (significant metric jumps), groups improvements within each version, and formats as narrative with deltas. Two audiences: technical (experiment IDs, configs) and stakeholder (plain English, percentages, no jargon). Output to paper/CHANGELOG.md.

The Complete Journey: v1.0.0 → v4.0.0

Version Phase Commands
v1.0.0 1–9: Core loop, hypotheses, novelty, statistics, families, tree-search 14
v1.3.0 10: Statistical rigor (seed, reproduce) 19
v1.4.0 11: Experiment intelligence (diagnose, ablate, frontier) 22
v1.5.0 12: Performance (profile, checkpoint) 24
v2.0.0 13: Deployment bridge (export) 25
v2.1.0 14: Research workflow (lit, paper) 27
v2.2.0 15: Orchestration (queue, retry, fork) 30
v2.3.0 16: Deep analysis (diff, watch, regress) 33
v2.4.0 17: Model composition (ensemble, stitch, warm) 36
v2.5.0 18: Scaling & efficiency (scale, budget, distill) 39
v3.0.0 19: Meta-intelligence (transfer, audit) 41
v3.1.0 20: Pre-training intelligence (sanity, baseline, leak) 44
v3.2.0 21: Model debugging (xray, sensitivity, calibrate) 47
v3.3.0 22: Feature & training intelligence (feature, curriculum) 49
v3.4.0 23: Model surgery (prune, quantize, merge, surgery) 53
v3.5.0 24: Experiment archaeology (trend, flashback, archive, annotate, search, template, replay) 60
v4.0.0 25: Research communication (cite, present, changelog) 63

Numbers

Metric v1.0.0 v4.0.0 Growth
Tests 257 1566 +1309 (6.1x)
Commands 14 63 +49 (4.5x)
Scripts 23 82 +59 (3.6x)
Phases 9 25 +16
Versions 1 17 +16

Full changelog

  • feat: citation_manager.py — citation tracking with add/list/check/bib, BibTeX generation
  • feat: generate_figures.py — presentation figure specs for 5 chart types with 3 styles
  • feat: generate_changelog.py — model changelog with version detection and audience adaptation
  • feat: /turing:cite, /turing:present, /turing:changelog command skills
  • test: 19 new tests across 3 test files
  • docs: All 72 roadmap items marked DONE

v3.5.0 — Experiment Archaeology

Choose a tag to compare

@ThePyProgrammer ThePyProgrammer released this 01 Apr 07:44

What's new in v3.5.0

Manage a research project that spans months, not just sessions.

Phase 24 is the largest release yet — seven commands that transform Turing from a single-session tool into a long-term research companion. It adds strategic trend analysis for 100+ experiment projects, instant context restoration after time away, disk space management, human annotations, natural language search, reusable templates, and experiment replay to test if old ideas work better now.

/turing:trend — Long-Term Trend Analysis

Strategic view over your full experiment history: improvement velocity per time window, family ROI ranking (which experiment families are productive vs exhausted), diminishing returns detection via log-curve fitting, and phase transition detection. The "state of the research" report that's invisible in individual experiment logs.

/turing:flashback — Session Context Restoration

Come back after a week and start working in 10 seconds. Shows current best, last session's experiments with outcomes, pending hypotheses, annotations, budget status, and suggested next action. Eliminates the 30-minute log-reading ritual.

/turing:archive — Experiment Lifecycle Cleanup

After 200+ experiments, reclaim gigabytes by compressing old artifacts while preserving a queryable summary index. Protects Pareto-optimal, best, and recent experiments. Supports --dry-run to preview before committing.

/turing:annotate — Retrospective Annotations

Add human notes to any experiment: "This only worked because the data was pre-sorted." Tags for categorization, search by content or tag. The knowledge that automated metrics can't capture.

/turing:search — Natural Language Experiment Search

Query 200+ experiments with natural language plus structured filters: "LightGBM high accuracy" --filter "accuracy>0.85". Combines keyword ranking with metric/status/family/date filters.

/turing:template — Experiment Template Library

Save winning configs as reusable templates that persist across projects at ~/.turing/templates/. Strip project-specific paths automatically. Apply to new projects with one command.

/turing:replay — Experiment Replay

Re-run a historical experiment with current infrastructure. Different from /turing:reproduce (which verifies the same result) — replay tests whether an old approach does better now because data, preprocessing, or libraries have improved.

Numbers

Metric v3.4.0 v3.5.0 Delta
Tests 1494 1547 +53
Commands 53 60 +7
Scripts 72 79 +7
Commits — 26 —

Full changelog

  • feat: 7 new scripts: trend_analysis, session_flashback, experiment_archive, experiment_annotations, experiment_search, experiment_templates, experiment_replay
  • feat: 7 new command skills: /turing:trend, /turing:flashback, /turing:archive, /turing:annotate, /turing:search, /turing:template, /turing:replay
  • test: 53 new tests across 7 test files
  • docs: ROADMAP Phase 24 (items 63-69) marked DONE, README updated

One phase remaining: Phase 25 (Research Communication, v4.0.0).