Skip to content

Releases: sanlee-ys/defense-news-classifier

v3.2.1 — The bar held, the ruler grew

Choose a tag to compare

@sanlee-ys sanlee-ys released this 03 Aug 03:05
10f6348

The global-boundary prompt clause ships — the same clause v3.2.0-era ADR-023 measured and reverted at p=0.0522, re-run against a ruler twice the size and adopted at p=0.0002.

The bar never moved. The ruler did.

What changed

One bullet in the region rubric: naming a US service, command, program office, contractor, unit or official identifies the actor, not a place — so a story whose only geography is institutional is global, not americas. A second sentence bounds the fix, because ADR-020 is the measured precedent for what happens when restraint lives only in a prompt.

PATCH, not MINOR: the classifier is more correct; the {category, operational_domain, region} contract is untouched.

The measurement

All four pre-registered rules passed at effective n=595 (295 rows reused, 300 new, 5 exact duplicates dropped before grading), at a design power of 0.837.

Axis Baseline Candidate Lift cand/base McNemar p
region (target) 89.9% 94.1% +4.2 35 / 10 0.0002
category (guardrail) 93.4% 94.1% +0.7 11 / 7 0.4807
operational_domain (guardrail) 89.4% 92.9% +3.5 30 / 9 0.0011

20 of 32 named global pulls fixed, 8 correct rows dragged the other way, harness clean 595/595 on all three axes.

Published gold numbers (n=54, human-graded)

Adoption moved the prompt fingerprint a59689e8…b0202d06…, forcing a full paid re-run — workhorse and judge, because judge_region_agreement is a gated floor and had to be re-measured rather than inherited. This is the first release since v3.0.0 where these numbers move at all.

Gated metric Floor v3.2.0 v3.2.1
category_accuracy 0.83 92.6% 94.4%
category_macro_f1 0.85 0.911 0.930
domain_accuracy 0.83 92.6% 98.1%
domain_macro_f1 0.83 0.933 0.982
judge_category_agreement 0.83 92.6% 94.4%
judge_domain_agreement 0.88 98.1% 92.6%
region_accuracy 0.78 87.0% 94.4%
judge_region_agreement 0.93 100.0% 96.3%

All eight pass as written. None was moved, waived, or added. judge_region_agreement is the one with real authority: 2 judge-vs-human disagreements against a budget of 3, and the number that would have stopped this release.

On the gold set the region error mode inverted — 7 misses to 3, only one still the named cluster, the other two americas rows over-called to global. Five under-calls traded for two over-calls, and the collateral was priced before the run rather than discovered after it.

The honest lesson

ADR-023 ran at about 49% power against the effect it observed — a coin flip on whether it could detect its own effect. So its finding was underpowered, not refuted, and its verdict was correct on its date. It is amended with a dated pointer, not rewritten: honoring that threshold when it was inconvenient is the only reason this adoption means anything.

Measure-first is now seven-for-seven, and this is the first adoption after six declines.

Also in this release

  • The re-run harness goes dormant and refuses every run and report entry point, closing a quiet trap: after the adoption re-run the published gold record is produced by the clause prompt, so --gold-report would otherwise have compared the candidate arm against itself and printed a 0.0-point lift.
  • Published-marker cascade across architecture, portfolio, learning-notes and kb-agent.

Full detail: ADR-024 · CHANGELOG · reports in evals/region_clause_rerun.txt and evals/region_clause_gold_rerun.txt

v3.2.0 - The ruler shrinks

Choose a tag to compare

@sanlee-ys sanlee-ys released this 02 Aug 14:08
99545c8

The region axis had one number and an 18-point error bar. It now has a 7-point one.

v3.2.0 runs the scaled region eval: 300 DVIDS snippets graded by the Opus judge that
validated at 100.0% region agreement against the human labels — the gate
ADR-014
set before allowing the judge anywhere near this axis.

Workhorse vs judge Accuracy 95% CI (Wilson) Macro-F1
Region 88.3% (265/300) [84.2%, 91.5%] 0.904
Category 91.7% (275/300) [88.0%, 94.3%] 0.763
Domain 89.3% (268/300) [85.3%, 92.3%] 0.868

The interval is the deliverable, not the accuracy. 88.3% against the gold set's 87.0% is
corroboration; what changed is that the number is now known to within 7 points instead of 18,
so a future change to the region rubric can finally be told apart from noise.

The named global cluster is confirmed systematic. On the 54-row gold set, all seven
region misses were rows whose true label is global that the model pulled to a specific
region, inferring a theater from the US actor where the snippet names no place — but seven
rows cannot separate a behavior from a run of luck. At n=300 they can: of 70 answer-key
global rows, 17 were pulled to a region (16 to americas), which is 49% of all 35
region disagreements
. One identified failure mode, half the region error. The prompt clause
that would target it is deliberately not in this release — measuring first is the method.

Read alongside, never instead of, the human-graded n=54 figures: this is workhorse-vs-judge
agreement, and the judge's own measured disagreement with humans on region was 0/54, which is
itself a wide interval.

The shipped classifier is unchanged — same prompt, same single call, same
{category, operational_domain, region} contract. MINOR.

Also in this tag: one API error taxonomy plus a stop-reason assertion so a truncated response
is never scored (ADR-021), a paired-comparison layer with harness health reported separately,
the ADR-017 classical baseline made runnable in a browser behind a parity gate, two provenance
pins that stop a stale prompt being published or graded, and two CI lanes.

Full detail: the [3.2.0] CHANGELOG entry
and ADR-022.
Report: evals/scale_eval_v3.txt.

v3.1.0 — The autonomy ladder, built to the top

Choose a tag to compare

@sanlee-ys sanlee-ys released this 25 Jul 23:35
629f3c6

Milestone: the autonomy ladder, built to the top. L3's second rung and all of L4 land here, together with the classical-ML bake-off that finally measured whether the LLM was worth paying for at all.

Four verdicts, three of them negative. All four are kept as records rather than deleted, which is the point of the repo.

What landed

ADR-017 — the classical baseline, finally measured. TF-IDF + logistic regression trained on 300 judge-graded snippets, scored once against the human gold set. The LLM wins decisively: category 92.6% vs 72.2% (McNemar p=0.013), domain 92.6% vs 66.7% (p=0.0005). This repo had measured three things layered on top of an LLM without ever baselining the LLM itself. Now it has.

ADR-018 — L3 rung 2, the agent-driven ML loop. An agentic outer loop reads out-of-fold errors and proposes the next sklearn experiment, reusing rung 1's honesty architecture as shared code. The first live run is the headline: the agent improved the metric it could see and degraded the one that matters, and the harness caught it. Best-by-B iteration scored B 0.699 (+6.0 over baseline) while held-out C fell to 0.545 (−8.6).

Nothing cheated. The mechanism is distribution shift — splits A and B are both judge-labeled wire text, so keywords mined from A's errors genuinely generalize there and mislead on C's human-labeled mix. Every guard behaved: plateau fired, best-iteration selection read B alone, and C exposed the trade precisely because nothing was allowed to optimize it. The held-out set vetoed the loop's own best iteration. That veto is the artifact.

ADR-019 — kNN exemplars, a clean null. k=3 BM25-retrieved labeled exemplars in the prompt: category 91.0% vs 90.0% (p=0.70), domain flat, region guardrail exact. That closes the retrieval question in three shapes — neighbor documents harmful, mined keyword features harmful off-distribution, labeled exemplars inert.

ADR-020 — L4, the multi-agent pipeline. Triage produces verbatim evidence spans, the shipped classify() runs blind, and a critic with a narrow rubric-checkable charter can bounce a label backward for exactly one re-classify. Hypothesis confirmed: the backward edge fixed 6 of the 7 named region misses. Pipeline declined anyway — the all-axes critic over-challenged (57.4% against an expected ~13%, tripping the spec's own red-flag rule) and did net harm elsewhere, including the repo's first statistically significant harm (scale domain 91.3 → 86.7, p=0.016) at roughly four times the calls.

The shipped single call stays production.

Why MINOR

Everything above is eval and experiment machinery. src/api.py is untouched since v3.0.0 and the system prompt did not move, so the {category, operational_domain, region} contract holds and callers see nothing new.

The published gold numbers are unchanged and still describe what ships — regenerating evals/metrics.json moved only the version field, which is the mechanical proof.

Roadmap note

v3.0.0's release notes pencilled the scaled (n=300) region eval in at this version number. That work is unblocked but unscheduled, and none of the ladder work depended on it, so the ladder shipped first and the scaled region eval moves to v3.2.0. Surfaces that named the old slot now use version-free phrasing so the reference cannot rot again.

Full detail: CHANGELOG 3.1.0

v3.0.0 — The region field

Choose a tag to compare

@sanlee-ys sanlee-ys released this 18 Jul 14:13
d5ccb7c

Milestone: the region field (ADR-014). The roadmap's planned breaking change: output becomes {category, operational_domain, region} — six labels (indo-pacific, europe, middle-east, africa, americas, global), with global as the single catch-all for no-anchor and multi-region stories, mirroring multi on the domain axis.

The numbers (n=54, human answer key — evals/gold_eval_v3.txt)

Axis Accuracy Judge vs human
Category 92.6% 92.6%
Operational domain 92.6% 98.1%
Region 87.0% 100.0%

The 100.0% judge-vs-human region agreement clears ADR-014's gate for the scaled (n=300) region eval in a future v3.1.0. The region misses are one named cluster: all seven are gold=global rows the model pulled to a specific region — it infers a theater from the US actor where the rubric's no-guessing rule says no anchor (evals/gold_confusion_v3.md names every row).

How the gold labels earned trust

The 54-row gold set gained a hand-labeled region column, owner-reviewed row by row — and then every label on all three axes was adversarially verified against each snippet's source article (one verifier per row, adversarial skeptics on challenges, a cross-row consistency audit). Category and domain survived 108/108; region took two review corrections before anything was measured against it. The labeling conventions are recorded in data/gold/README.md.

Also in this release

  • Measured CI floors for regionregion_accuracy 0.78, judge_region_agreement 0.93, both derived from the live run (never aspirational); the gate now grades all eight floors against the committed v3 snapshot. Deliberately no region macro-F1 floor (europe support is n=1).
  • Frozen v2 records — all v3 eval outputs use new _v3 filenames; the v2 snapshots behind the published two-axis numbers are never regenerated.
  • Live-pass hardening — the strict-mode guarantee blipped three times in the wild (the API returned a tool input missing required fields despite strict: true; identical replays were clean), so the eval harness retries transient InvalidLabelErrors with bounded backoff, and interrupted runs now print a PARTIAL RUN banner instead of clean-looking fractional results.

Breaking-change note for integrators: the /classify response now carries a required region field, and the SYS-004 contract literals moved with it.

Full detail: CHANGELOG 3.0.0

v2.2.0 — Tiered routing: measured and declined

Choose a tag to compare

@sanlee-ys sanlee-ys released this 18 Jul 14:06
8975b83

Milestone: tiered model routing — measured and declined (ADR-013). The project's second measured negative result, shipped as such.

The shipped classifier forces tool use and emits no confidence signal, so the routing layer manufactures one: a routing-only tool variant requires a runner_up_category, and escalation to Opus fires exactly when the top-two are {technology, operations} — the one clustered confusion the gold set ever showed. Measured with three honesty guards (human-gold-only quality grading, zero-new-spend escalation replay, schema perturbation measured rather than assumed away):

  • Routing moved +0 rows on both axes (94.4% / 92.6%, identical to the workhorse); escalated rows read fixed 0, broke 1, unchanged 8 — the one change was Opus importing the judge's own known error.
  • Even all-Opus scores the same 94.4% category — no category router has headroom on this set.
  • ~1.97x cost per article at a 19.4% real-world escalation rate, for zero measured quality.

Routing is declined; the harness stays dormant as the reproducible record. Also fixed: bulk passes now record-and-continue on safety-layer refusals (the s151 finding — a schema change alone flipped a benign chem-bio defense story into a refusal).

Full detail: CHANGELOG 2.2.0

v2.1.0 — Scale the eval

Choose a tag to compare

@sanlee-ys sanlee-ys released this 18 Jul 14:06
f4c6f4f

Milestone: scale the eval. The Opus judge — validated against the human gold labels at 94.4% / 94.4% agreement — grades 300 fresh real DVIDS snippets, disjoint from both the corpus and the gold set, with 95% Wilson confidence intervals: category 93.3% [89.9, 95.6], domain 90.3% [86.5, 93.2] — half the n=54 CI width, corroborating the gold-set numbers rather than replacing them.

This tag also ships everything accumulated under [Unreleased] since v2.0.1:

  • Evals-as-CI capability gate (ADR-007) — a free offline gate on every push/PR grading the committed snapshots against thresholds.toml floors, plus a paid weekly live gate.
  • Rung-1 prompt-optimization loop (ADR-005, autonomy ladder L3) — agent-driven prompt revision with a structural Goodhart guard and an explicit done-signal, held for this release per the spec's sequencing.
  • Prompt refinement (tech-vs-ops + land) — measured on the gold set: category 90.7% → 94.4%, domain 90.7% → 92.6%, technology recall to 1.000.
  • BM25 grounding retired (ADR-012) — a fair same-prompt re-measure showed grounding no longer pays: neutral on category, 0 fixed / 4 broken on domain. Removed from the shipped path; kept dormant as the record.
  • OpenTelemetry tracing over classify(), the v2.0.2 orchestration-test backfill (src/ coverage 90% → 97%), and a parity Jenkinsfile.

Full detail: CHANGELOG 2.1.0

v2.0.1

Choose a tag to compare

@sanlee-ys sanlee-ys released this 05 Jul 04:34
59eb8ad

Dead-code hardening release. Removes the Kafka consumer path outright instead of carrying it as inactive weight. No change to the {category, operational_domain} output contract; the classifier's live surface remains the /classify HTTP provider.

Removed

  • Kafka consumer (src/consumer.py) and its test suite. The consumer was an event-driven alternative to the /classify HTTP path — reading NoteCreated events off a note-events topic, classifying in-process, and writing labels back as namespaced tags. It went dead once the notes consumer moved to an HTTP BackgroundTasks writeback loop: nothing publishes note-events anymore, so it was a no-op against the live system. Deleted here (src/consumer.py, tests/test_consumer.py, tests/test_consumer_integration.py); the reference implementation remains available in git history and the ADRs.
  • Unused Kafka dependencies (kafka-python, testcontainers[kafka]) dropped from pyproject.toml / requirements.txt; uv.lock regenerated.
  • CI integration-test job (the Testcontainers Kafka lane in .github/workflows/tests.yml) and the matching Jenkins stage — neither had an integration test left to run.
  • Orphaned Kafka vars in .env.example, the consumer entry in the README structure listing, and docs/integration-testing.md.

Full unit suite (68 tests), ruff, and CodeQL pass clean.

v2.0.0 — Real-text RAG eval

Choose a tag to compare

@sanlee-ys sanlee-ys released this 22 Jun 00:03
a1068e3

v2 moves the classifier off synthetic, self-graded data and onto real public-domain text, with retrieval grounding and a non-circular, human-labeled answer key. v1 measured in-distribution consistency; v2 measures real-world accuracy.

Headline results (real text, human-graded gold set, n=54)

Field Accuracy Macro-F1
Category 88.9% 0.906
Operational domain 88.9% 0.894
  • v1's worst class fixed: industry recall 0.217 → F1 1.000 on real SEC filings (caveat: n=5 clear-cut cases).
  • Judge validated: Opus judge agrees with the human labels 88.9% (category) / 94.4% (domain).
  • Grounding measured, not assumed: BM25 grounding gave +1.9% category / flat domain → lexical retrieval doesn't justify embeddings here. The negative result is the finding.

What's new

  • Real public-domain corpus (62 docs: DVIDS news wire + SEC filings) with a BM25 retriever (src/retrieve.py).
  • Human-labeled gold set (data/gold/) + labeling guide; honest eval harness with Opus-judge validation (src/gold_eval.py).
  • Retrieval-grounded classification with citations (src/classify_rag.py) + grounding-lift eval with flip analysis (src/gold_eval_rag.py).
  • README rewritten to lead with the v2 numbers.

Full details in CHANGELOG.md. Diff: v1.1.0...v2.0.0

1.1.0 — Tighten the measurement

Choose a tag to compare

@sanlee-ys sanlee-ys released this 21 Jun 02:07

Hardens the measurement around the v1.0.0 classifier. The classifier's
happy-path behavior is unchanged; this release is about trusting the numbers.

Added

  • Macro-F1 — the honest single number for an imbalanced problem
    (category macro-F1 0.765 vs 79.0% accuracy; domain 0.973).
  • Stability harness (src/stability.py) — sizes the run-to-run noise floor
    (category accuracy std 0.24 pts over 5 runs), so a config change counts only if
    it clears ~2× std. Confirms the reverted prompt experiment's 2.3-pt regression
    was real (~10× the noise), not luck.
  • Error audit (evals/error_audit.md) — all 67 misses triaged; ~90% are
    label-scheme overlap (industry/procurement/technology), not model error.
  • Enum-validation guard in classify() — re-samples once on an out-of-enum
    label, then raises. Caught a real defect: 1/300 predictions returned an
    invalid category.

Fixed

  • Corrected the docs claim that tool-use "rejects out-of-enum output at the API
    layer." It's a guided prior, not a hard constraint; enum membership is now
    validated in code.

Docs

  • docs/CASE_STUDY.md narrative writeup; docs/how-it-works.md plain-language
    one-pager with a precision/recall/F1 definition and a "Threats to validity"
    section.

Full details in CHANGELOG.md.

v1.0.0 — Initial release

Choose a tag to compare

@sanlee-ys sanlee-ys released this 20 Jun 21:25

First complete version of the defense news classifier.

  • Synthetic dataset generator (300 labeled articles)
  • LLM classifier via Anthropic API with structured JSON output
  • Eval harness with accuracy, per-label precision/recall/F1, confusion matrices, and misclassification log
  • 79.0% category accuracy / 97.3% domain accuracy on 300-article eval set
  • 33 unit tests, GitHub Actions CI, pre-commit hooks

See CHANGELOG.md for full details including a documented negative result (prompt experiment that regressed the eval).