Skip to content

Releases: bdeva1975/whydunit

v0.2.0 — The AI narrative layer

Choose a tag to compare

@bdeva1975 bdeva1975 released this 25 Sep 09:50

Whydunit v0.2.0 — The AI narrative layer

v0.1 diagnosed pipeline failures deterministically. v0.2 adds an LLM that
explains those findings — with zero authority.

What's new

  • whydunit.explain — an optional narrative layer over the case file:
    • The model is shown only the leading hypothesis and its evidence,
      with stable IDs, and must cite an ID for every factual claim.
    • A deterministic citation validator parses every narrative and
      refuses any that cites an unknown ID or makes an uncited claim —
      one corrective retry, then refusal. Invalid narratives are never shown.
    • Alternative hypotheses are appended by code, labelled
      machine-written — the model never sees them, so it cannot misstate them.
  • CLI: python -m whydunit.explain --scenario retrieval_degradation --seed 42
  • Console: "AI narrative report" panel on the Notes & Case File page,
    plus scenario-qualified export filenames.
  • Optional dependency: uv sync --extra llm + ANTHROPIC_API_KEY. The
    core remains fully offline and LLM-free; CI runs without a key
    (explain tests use a stubbed client).

The design that survived contact

During development the validator refused three narratives — invented ID
formats, uncited scene-setting, uncited summaries of lower-ranked
hypotheses. The third was fixed structurally, not by prompt tuning: the
model no longer sees the alternatives at all. Every refusal was the
system working; none of those narratives reached a user. Full story:
docs/narrative-layer.md

Unchanged

12/12 top-1 diagnostic accuracy, 0 FP / 0 FN, ground-truth firewall
enforced by test. 158 tests passing, ruff clean, CI green on
ubuntu/windows × py3.12/3.13.

MIT licensed.

v0.1.0 — First public release

Choose a tag to compare

@bdeva1975 bdeva1975 released this 25 Sep 04:36

Whydunit v0.1.0 — First public release

Find the why behind AI pipeline failures — synthetic, offline, deterministic.

What's in the box

  • Simulator — 10-stage AI pipeline, 35-signal telemetry catalogue, seeded generator with diurnal seasonality, and a 12-scenario incident library (σ-shift injection with onset/ramp control). CLI: python -m whydunit.simulator
  • Forensic engine — robust z-score + EWMA anomaly detection, two-window change-point detection, cross-signal correlation clustering, dependency-DAG origin reasoning, and a signature-scoring hypothesis engine with confidence bands. Fully deterministic, zero LLM calls.
  • Streamlit console — ten investigation pages: dashboard, incident explorer, pipeline view, timeline, signal explorer, forensic analysis, evidence graph, incident comparison, case-file export (server-side fallback included), and ground-truth eval.
  • Eval harness — python -m whydunit.evaluation grades diagnoses against ground truth, which is firewalled out of the forensic engine (enforced by test).

Report card (1 day @ 1-min resolution, seed 42)

  • 12/12 scenarios diagnosed correctly at top-1
  • 0 false positives, 0 false negatives
  • Detection delays: 10–37 minutes

Under the hood

  • Python ≥3.12, uv-managed; pandas 3 / numpy 2.5 / plotly / networkx / scikit-learn
  • 140 tests passing, ruff clean
  • CI: lint + format + tests + eval smoke on ubuntu/windows × py3.12/3.13

Roadmap

v0.2: optional LLM explanation layer (explains deterministic evidence, never generates it), more scenarios. See README for the full roadmap.

MIT licensed.