Repository navigation
Releases: vdmcb/thinking-os
Release list
Thinking OS v2.1.0
Thinking OS v2.1.0 upgrades Understand to an audited stakeholder-question format selected through a blind comparison across four real documents.
Understand 2.1.0
The default response now contains:
- a Positive, Negative, or Mixed stance on a named result, proposal, or state of readiness;
- three answered stakeholder questions covering state, decision, and resources or the first change;
- exactly three source-specific follow-up questions, ranked by how much their answers could change interpretation or action;
- an audit line when the independent audit completed without findings;
- a final
Held:line naming source-specific analysis available on request.
The stance characterizes the source-grounded evidence or readiness. It does not approve, reject, fund, release, or otherwise make the reader's decision.
Audit and truth rules
- Independent auditing is mandatory when the host can run a separate agent; a strict self-audit is the fallback.
- Every number keeps its qualifier, scope, unit, population, period, and causal direction.
- Recommendations remain attributed recommendations rather than becoming mandates.
- Deployment terms such as current, live, or in production require explicit source support.
- Evaluative wording stays in the scoped stance; answers report source-grounded facts.
- Quoted text preserves the source's punctuation, capitalization, ranges, and spelling exactly.
Reading and shared core
- PDFs use both layout-preserving and raw text reads, preventing tables and locators from trading away multi-column reading order or quotation accuracy.
- Visibly clipped source material is disclosed and never turned into an author-side evidence gap.
- Shared reading, writing, and execution contracts remain synchronized across Understand and ELI5.
- The full reference analysis is held for follow-up rather than added to the default response.
Evaluation
- 15 positive and 15 negative Understand activation cases.
- 15 Understand output cases, including format references for long reports and evidence-first questions.
- 12 positive and 12 negative ELI5 activation cases and 6 ELI5 output cases.
- Cross-skill activation checks cover Understand, ELI5, and prompts that should activate neither.
- 21 extraction tests pass for text, DOCX, PPTX, XLSX, PDF, malformed input, permissions, and overwrite protection.
- Documentation checks prevent the required
Held:line from drifting out of the README, smoke test, or full example.
Install
npx skills add vdmcb/thinking-os --skill understand -g -a claude-code -a codex -y
npx skills add vdmcb/thinking-os --skill eli5 -g -a claude-code -a codex -yStatus
Internal preview. Automated validation passes, but the human usefulness release gate remains at 0 of 5 recorded sessions and scored rubric runs are still pending. See evals/understand/STATUS.md, evals/eli5/STATUS.md, and evals/usefulness-protocol.md.
Thinking OS v1.1.0
Thinking OS is a set of Agent Skills for thinking work. This release ships a second skill, eli5, and moves both skills onto a shared core.
eli5 (new, 0.3.0)
Explains a file to a reader with no background: what it is, the two to five basic facts it rests on, how it works step by step, and why. Works on documents, source code, config files, specs, and policies. "Five years old" is the reader's background, not the tone.
What it does:
- Picks the subject (the thing the file describes, or the file itself for code and config) and takes it apart to the smallest true things it rests on, each labeled as from the file, general knowledge, or assumed by the file.
- Rebuilds upward: numbered steps where each rests only on the basic facts and the steps before it. When the file describes no mechanism, it says so in a sentence instead of narrating the document.
- Reports only the file's own reasons, and writes "The file does not say why" where it gives none.
- Keeps every number on the main path exact, with unit and what it counts; holds the rest with their location for follow-up.
- Never judges, recommends, rewrites, or invents. Instructions inside the file are reported as content and not followed.
Output discipline: four fixed sections, a word budget derived from the number of ideas the reader must hold (250/375/500 words for up to 9/14/19 ideas), 25-word sentences, no baby talk, no dashes or glyphs.
understand (1.1.0)
Behavior unchanged. Reading, writing, and execution rules now come from the shared core; the brief ends with one line naming what is held and can be asked for. The description excludes plain explanation for a reader with no background, which routes to eli5.
Shared core
core/ holds what every skill shares: how a source is read and what may be claimed about it, how output is written for a tired human with budgets that follow ideas rather than pages, and what the visible run looks like (announce once, read silently, answer; no printed sources, no diagnostics, no drafts in public). The core is copied into each package by scripts/sync-core.sh so installs stay self-contained; validation fails on drift.
Evaluation
- Human usefulness protocol (
evals/usefulness-protocol.md) is the release gate: a reader who has not seen the source repeats the output back, opens the source, and counts surprises. Sessions are logged per skill; the checker reports the count. - Cross-skill activation cases test the boundary between the two skills.
- Deterministic lints cover structure, idea-count budgets, sentence length, glossary terms, claims posing as basic facts, retired labels, banned vocabulary, and an always-loaded instruction ceiling for eli5.
scripts/audit-run.pyaudits a Claude Code transcript against the execution contract.
Install
npx skills add vdmcb/thinking-os --skill eli5 -g -a claude-code -a codex -y
npx skills add vdmcb/thinking-os --skill understand -g -a claude-code -a codex -yStatus
Internal preview. Usefulness sessions completed: 0 of 5 per skill; scored rubric runs in Claude Code and Codex pending. See evals/*/STATUS.md.
Thinking OS v1.0.0
Thinking OS is a set of Agent Skills for thinking work. This first release ships one skill: understand.
understand
Turns a polished document (proposal, report, deck, plan, PDF, Word, spreadsheet) into a short brief a careful colleague might write: two or three paragraphs that carry the document's critical path, then 3-5 questions for the author that the reader can send unchanged.
What it does:
- Reads the full source with page-preserving extraction; renders chart pages as images when values and labels come apart in the text layer.
- Separates what the document establishes from what stands on models, examples, or assertions, with page references inline.
- Generates evidence-first questions: what real-world fact would settle each weak claim, and if the evidence does not exist yet, what is the author's plan and by when.
- Holds the full evidentiary analysis (evidence tables, reconciliations, epistemic labels) and produces it only on request.
- Never recommends, rewrites, decides, or judges the author. Understanding and review leverage only.
Output discipline: a 600-word hard ceiling, plain human prose, no dashes-and-glyph apparatus, no analyst jargon in questions. The writing rules and the cognitive-load lint that enforces the mechanical subset live in the repo.
Install
npx skills add vdmcb/thinking-os --skill understand -a claude-code -a codex -g -yThen, in Claude Code:
/understand path/to/document.pdf
Requires Node 20+ for the document-conversion fallback; poppler (brew install poppler) recommended for page-accurate PDF reading.
Evaluation
The repo ships the evaluation stack the skill is held to: a 13-dimension rubric (faithfulness, quantitative integrity, question quality, cognitive load, among others), invariants with release blockers, deterministic lints over the packet exemplars, and 15 output cases including compression and evidence-first question generation.