Skip to content

Thinking OS v2.1.0

Latest

Choose a tag to compare

@vdmcb vdmcb released this 28 Aug 06:48
· 9 commits to main since this release
c79df0b

Thinking OS v2.1.0 upgrades Understand to an audited stakeholder-question format selected through a blind comparison across four real documents.

Understand 2.1.0

The default response now contains:

  • a Positive, Negative, or Mixed stance on a named result, proposal, or state of readiness;
  • three answered stakeholder questions covering state, decision, and resources or the first change;
  • exactly three source-specific follow-up questions, ranked by how much their answers could change interpretation or action;
  • an audit line when the independent audit completed without findings;
  • a final Held: line naming source-specific analysis available on request.

The stance characterizes the source-grounded evidence or readiness. It does not approve, reject, fund, release, or otherwise make the reader's decision.

Audit and truth rules

  • Independent auditing is mandatory when the host can run a separate agent; a strict self-audit is the fallback.
  • Every number keeps its qualifier, scope, unit, population, period, and causal direction.
  • Recommendations remain attributed recommendations rather than becoming mandates.
  • Deployment terms such as current, live, or in production require explicit source support.
  • Evaluative wording stays in the scoped stance; answers report source-grounded facts.
  • Quoted text preserves the source's punctuation, capitalization, ranges, and spelling exactly.

Reading and shared core

  • PDFs use both layout-preserving and raw text reads, preventing tables and locators from trading away multi-column reading order or quotation accuracy.
  • Visibly clipped source material is disclosed and never turned into an author-side evidence gap.
  • Shared reading, writing, and execution contracts remain synchronized across Understand and ELI5.
  • The full reference analysis is held for follow-up rather than added to the default response.

Evaluation

  • 15 positive and 15 negative Understand activation cases.
  • 15 Understand output cases, including format references for long reports and evidence-first questions.
  • 12 positive and 12 negative ELI5 activation cases and 6 ELI5 output cases.
  • Cross-skill activation checks cover Understand, ELI5, and prompts that should activate neither.
  • 21 extraction tests pass for text, DOCX, PPTX, XLSX, PDF, malformed input, permissions, and overwrite protection.
  • Documentation checks prevent the required Held: line from drifting out of the README, smoke test, or full example.

Install

npx skills add vdmcb/thinking-os --skill understand -g -a claude-code -a codex -y
npx skills add vdmcb/thinking-os --skill eli5 -g -a claude-code -a codex -y

Status

Internal preview. Automated validation passes, but the human usefulness release gate remains at 0 of 5 recorded sessions and scored rubric runs are still pending. See evals/understand/STATUS.md, evals/eli5/STATUS.md, and evals/usefulness-protocol.md.