Releases: bdeva1975/whydunit
Release list
v0.2.0 — The AI narrative layer
Whydunit v0.2.0 — The AI narrative layer
v0.1 diagnosed pipeline failures deterministically. v0.2 adds an LLM that
explains those findings — with zero authority.
What's new
whydunit.explain— an optional narrative layer over the case file:- The model is shown only the leading hypothesis and its evidence,
with stable IDs, and must cite an ID for every factual claim. - A deterministic citation validator parses every narrative and
refuses any that cites an unknown ID or makes an uncited claim —
one corrective retry, then refusal. Invalid narratives are never shown. - Alternative hypotheses are appended by code, labelled
machine-written — the model never sees them, so it cannot misstate them.
- The model is shown only the leading hypothesis and its evidence,
- CLI:
python -m whydunit.explain --scenario retrieval_degradation --seed 42 - Console: "AI narrative report" panel on the Notes & Case File page,
plus scenario-qualified export filenames. - Optional dependency:
uv sync --extra llm+ANTHROPIC_API_KEY. The
core remains fully offline and LLM-free; CI runs without a key
(explain tests use a stubbed client).
The design that survived contact
During development the validator refused three narratives — invented ID
formats, uncited scene-setting, uncited summaries of lower-ranked
hypotheses. The third was fixed structurally, not by prompt tuning: the
model no longer sees the alternatives at all. Every refusal was the
system working; none of those narratives reached a user. Full story:
docs/narrative-layer.md
Unchanged
12/12 top-1 diagnostic accuracy, 0 FP / 0 FN, ground-truth firewall
enforced by test. 158 tests passing, ruff clean, CI green on
ubuntu/windows × py3.12/3.13.
MIT licensed.
v0.1.0 — First public release
Whydunit v0.1.0 — First public release
Find the why behind AI pipeline failures — synthetic, offline, deterministic.
What's in the box
- Simulator — 10-stage AI pipeline, 35-signal telemetry catalogue, seeded generator with diurnal seasonality, and a 12-scenario incident library (σ-shift injection with onset/ramp control). CLI:
python -m whydunit.simulator - Forensic engine — robust z-score + EWMA anomaly detection, two-window change-point detection, cross-signal correlation clustering, dependency-DAG origin reasoning, and a signature-scoring hypothesis engine with confidence bands. Fully deterministic, zero LLM calls.
- Streamlit console — ten investigation pages: dashboard, incident explorer, pipeline view, timeline, signal explorer, forensic analysis, evidence graph, incident comparison, case-file export (server-side fallback included), and ground-truth eval.
- Eval harness —
python -m whydunit.evaluationgrades diagnoses against ground truth, which is firewalled out of the forensic engine (enforced by test).
Report card (1 day @ 1-min resolution, seed 42)
- 12/12 scenarios diagnosed correctly at top-1
- 0 false positives, 0 false negatives
- Detection delays: 10–37 minutes
Under the hood
- Python ≥3.12, uv-managed; pandas 3 / numpy 2.5 / plotly / networkx / scikit-learn
- 140 tests passing, ruff clean
- CI: lint + format + tests + eval smoke on ubuntu/windows × py3.12/3.13
Roadmap
v0.2: optional LLM explanation layer (explains deterministic evidence, never generates it), more scenarios. See README for the full roadmap.
MIT licensed.