Releases: naeyn/nobody
Release list
nobody v0.3.0 — leakage-first pipeline
nobody v0.3.0 is the first configuration in three research loops to pass every pre-registered release gate. It reduces German protocol-v2 PII token leakage to 1.51% while retaining 0.795 exact-span F1.
The model weights are byte-for-byte unchanged from v0.2.0. This is a source and production-policy release.
What changed since v0.2.0
- Expanded cue-gated German building-number fields (
Gebäude-Nr.,Gebäudenr.,Bau-Nummer,Haus-Nr.) without accepting generic address, room, order, or telephone fields. - Restored recall-first month/slash date handling after weak prepositions. Explicit operational fields remain excluded.
- Added spaced and punctuated phone exports plus German contact lead-ins.
- Prevented phone spans inside IBAN-like values, credit cards, hyphenated identifiers, and long digit runs.
- Trimmed corpus-style honorifics from person boundaries without dropping all-title spans.
- Rejected network identifiers and generated placeholders misclassified as postal addresses.
- Trimmed structured phone tails and normalized adjacent postal-address components.
- Expanded the public behavior suite from 89 to 111 passing tests.
Against the loop-6 single-model champion at identical weights, the release changes produced:
- paired ΔF1 −0.0064 [−0.0097, −0.0034];
- paired Δleakage −0.54 percentage points [−0.96, −0.20].
Both effects exclude zero. This is an explicit gate-first trade: a small, real F1 loss bought a larger, real reduction in missed PII.
Seven research loops
- Instrument and baseline: built the evaluator, leakage metric, deterministic/model pipeline, and bounded synthetic runs. High synthetic scores did not transfer to German.
- Target-shaped synthetic data: reached the first gate-passing release line on the original German set (0.662 F1 / 1.74% leakage).
- Full epochs, OpenPII, and WiSE-FT: raised research F1 to 0.783, but leakage remained 2.67%; G2 failed.
- Protocol v2 and the licensed mix: selected checkpoint 800 on development data; v0.2.0 reached 0.791 / 1.63%.
- Address, DOB, and phone coverage: reached 0.799 / 2.05%; G2 failed by 0.049pp. Broad DOB cues, street-suffix rules, and continued address training were rejected.
- Synthetic address quality, routing, and ensembles: the best ensemble reached 0.815 F1 but 2.61% leakage. Component-address models reduced some misses but lost too much exact-span precision. Every candidate failed G2.
- Fix the instrument, then the pipeline: a second disjoint development stream exposed DOB and phone leakage hidden by the first. Four deterministic changes produced 0.795 / 1.51% and the first all-gates pass.
The detailed accepted and rejected hypotheses are in RESEARCH.md.
Scientific methodology
- Candidate data, checkpoints, rules, and fixed thresholds were selected on disjoint development data.
- Each loop froze its finalist and decision rule before one canonical evaluation.
- Shipping required German F1 ≥ zero-shot + 0.05, ≤2.00% leakage on both German and synthetic sets, and cross-set F1 ≥0.55.
- Loop 7 additionally required the German leakage interval's upper bound to remain below 2.00%; it passed at 1.995%, only 0.005pp inside the bar.
- System and finalist comparisons use paired document-level bootstrap differences with 10,000 resamples, not overlapping marginal confidence intervals.
- Exact-text and n-gram contamination checks included planted-positive controls.
- Failed gates and negative experiments are reported rather than silently discarded.
Honest benchmark verdict
External German protocol-v2 benchmark
3,000 synthetic German validation documents from ai4privacy/pii-masking-400k:
| model | exact-span F1 | PII token leakage |
|---|---|---|
| nobody v0.3.0 | 0.795 [0.780–0.810] | 1.51% [1.079%–1.995%] |
| Piiranha-v1, matched pipeline | 0.840 | 0.93% |
| GLiNER base, zero-shot | 0.635 | 2.26% |
nobody does not beat Piiranha on Piiranha's home release. The paired difference is −0.0452 F1 [−0.0586, −0.0318] and +0.58pp leakage [+0.11, +1.09]. Piiranha remains a CC-BY-NC-ND research comparator, not a dependency.
Multilingual synthetic regression
nobody remains at 0.927 F1 / 0.51% leakage across German, English, and Dutch. This is in-distribution regression evidence, not real-world evidence.
Cross-release result
On a contamination-filtered OpenPII-1M German validation sample, nobody scores 0.982 versus Piiranha's 0.915, paired ΔF1 +0.067 [+0.059, +0.074]. This ranking reversal is not a universal win: each model leads on its home generator, OpenPII-1M has no DOB class, and address boundaries differ. The ≈0.105 interaction is evidence that synthetic exact-span benchmarks measure generator and annotation conventions as well as detection capability.
Privacy and artifact boundary
All published training and benchmark sources are described by their publishers as synthetic. Raw benchmark documents, leak excerpts, local paths, credentials, chat identifiers, and production redaction logs are not included. No 400k evaluation row was used to train the released weights.
The Hugging Face model card now includes structured metrics, current methodology, intended uses, limitations, provenance, and discoverability metadata. Binary checksum remains:
model.safetensors — 84cb4207ec386a2467894328bccd83127ac62ee771779c0c909047ea3cc54d06
Full change list: CHANGELOG.md
nobody v0.2.0 — initial public model
Historical source release corresponding to the first published nobody-pii-de Hugging Face artifact.
Highlights
- German-first GLiNER model plus deterministic redaction pipeline.
- Checksum validation for IBAN, credit cards, German tax IDs, and Dutch BSNs.
- Format-aware email, phone, license-plate, and numeric-date detection.
- Mixed checkpoint trained from a pinned Apache-2.0 GLiNER base on 25,909 German OpenPII-1M rows (CC-BY-4.0) plus 12,672 companion synthetic instances.
- Protocol-v2 German result: 0.791 F1 / 1.63% PII token leakage.
- Multilingual synthetic regression result: 0.927 F1 / 0.51% leakage.
The German leakage point estimate passed the original 2% gate, but its 95% bootstrap interval [1.17%, 2.17%] crossed the threshold. All benchmark corpora were synthetic; this release did not establish real-world performance.
The current model binary introduced here remains unchanged in v0.3.0:
model.safetensors — 84cb4207ec386a2467894328bccd83127ac62ee771779c0c909047ea3cc54d06