Training-data memorization auditor for Hugging Face Trainer / TRL fine-tunes.
A local, Apache-2.0 plugin that answers two questions every fine-tune in a regulated setting should document
- Membership - can an attacker with logprob access tell what was trained on?
- Regurgitation - does the model emit training content when prompted with a prefix?
memaudit injects pre-registered canaries into the raw dataset, runs a PEFT-aware pre-flight when training starts, and writes memaudit-report.json when training ends. The same engine is available as a post-hoc CLI (memaudit audit --ref auto).
It runs entirely on your machine. There is no phone-home, no account, no SaaS.
This tool produces evidence of resistance to the attacks it actually runs. It does not make you GDPR / AI Act / CNIL compliant.
Fine-tuners who have to document membership-inference and regurgitation testing (legal / compliance / security reviewers), and engineers who need the audit inside the training loop rather than a post-hoc upload.
What you are buying: a pip-installable, fully-local test layer. You get a versioned JSON report with both verdicts, negative controls, Clopper-Pearson CIs, provenance hashes, a shipped EDPB compliance_annex, and an explicit limitations statement. You do not get a compliance certificate or a SaaS dashboard.
pip install memaudit
memaudit demoPyPI — extras: "memaudit[peft]", "memaudit[trl]", or "memaudit[peft,trl]". From source: git clone then pip install -e ".[dev,peft,trl]".
This is the public run to click first: TinyLlama-1.1B-Chat, real Alpaca, LoRA r=8, README API, 100 inserted canaries / 200 controls, honest 0.874% canary budget, model-scored high_ppl canaries. Measured 2026-08-30 on Apple M3 Pro (18 GB, MPS). Full write-up: docs/case-study-alpaca.md — live site.
| Metric | Measured (examples/alpaca-powered-report.json) |
|---|---|
| Model / data | TinyLlama-1.1B-Chat-v1.0 + 20,000 tatsu-lab/alpaca rows |
| LoRA / epochs / budget | r=8 on q,k,v,o_proj — 1 epoch — 0.874% of tokens (22,228 / 2,521,431) |
| Canaries | requested high_ppl — actual model_scored_high_ppl (300/300 in PPL band) |
| Inserted / held-out | 100 / 200 (include_prob=1.0 so n is not a coin flip) |
| Method | base-calibrated Min-K%++ — reference.mode=disable_adapter |
| TPR @ 1% FPR | 0.100 (10/100) — 95% CI [0.049, 0.176] — headline valid |
| Repetition-tier curve (same threshold) | 1x 0/34 — 4x 1/33 — 16x 9/33 |
| AUC (secondary) | 0.837 |
| Regurgitation | 0/100 under this prefix/decoding/exact-match protocol |
| Negative-control regurgitation | 0.00 (n=200) |
| Train / audit / wall | load 3 s — data 304 s — canary gen 1,921 s — train 10,685 s — audit 829 s — wall 3 h 35 min (22:19—01:55 IST) |
A one-epoch LoRA (r=8) of TinyLlama-1.1B-Chat on 20,000 real Stanford Alpaca rows, with an honest 0.874% canary budget and model-scored high-perplexity canaries, did leak membership at the pre-declared operating point: 10 of 100 inserted canaries were detectable at 1% FPR (TPR 0.100, 95% CI [0.049, 0.176]). Ranking agreed — AUC 0.837. The same threshold decomposes as 1x 0/34, 4x 1/33, 16x 9/33: the pooled 10% is substantially a duplication/exposure stress signal, not a 10% detection probability for a single-exposure record. The same run was 0/100 under this prefix/decoding/exact-match protocol (not a claim of no extraction risk).
The prior powered run used a uniform_vocab fallback (no model at generation): TPR 0.180 (18/100), CI [0.110, 0.269], AUC 0.776 — archived as examples/alpaca-powered-report-v0.1-uniformvocab.json. A 12-canary first look (TPR 0.500, CI [0.211, 0.789]) lives in examples/alpaca-case-study-report.json and the case-study appendix — not the headline. The already-measured distilgpt2 n=100/200 rows (TPR 0.000 [0, 0.036], risky AUC 0.848) stay in the LoRA benchmark table below.
pip install "memaudit[peft,trl]"
python examples/alpaca_case_study.pyThese numbers were produced by python examples/demo.py on 2026-08-27 (Apple MPS). The model is a randomly-initialized 1-block TinyDemoLM (hidden=64, vocab=256), full fine-tune, seed 0 — a positive-control run so the instrument can show a clear hit and a clean control side.
| Metric | Measured value |
|---|---|
| Method | base-calibrated Min-K%++ (secret-span) |
| Inserted canaries / held-out controls | 16 / 100 |
| Repetition tier | 16x |
| TPR @ 1% FPR | 1.000 (16/16 detected) |
| 95% CI (Clopper-Pearson) | [0.794, 1.000] |
| Headline valid? | yes (n_controls=100) |
| Regurgitation (exact / BLEU>0.75 / NED<=0.1) | 16/16 = 1.000 at 16x |
| Negative-control regurgitation | 0.00 (n=100) |
| Negative-control mean headline score | -15.31 (well below members) |
| Train wall-clock | 7.0 s (last-batch loss 0.126) |
| Audit wall-clock | 20.3 s |
| Seed / schema / tool | 0 / 1.1.0 / 0.1.0 |
Reproduce:
pip install memaudit
memaudit demo --output-dir examples
# from a clone: pip install -e ".[dev]" && python examples/demo.pyA checked-in copy of that report lives at examples/demo-report.json. Re-running the demo overwrites it with whatever this machine measures.
If a tiny model cannot memorize, memaudit refuses a fake TPR@1%FPR rather than inventing one. This run memorized; Clopper-Pearson intervals are printed with the point estimate.
Measured 2026-08-27 on Apple MPS. Scale: pretrained distilgpt2 + LoRA. Canary token budget 0.93%. Scoring used live peft.disable_adapter() (--ref auto) on one model copy.
| Metric | Run A (1 ep, r=8, lr=2e-4) | Run B (3 ep, r=16, lr=5e-4) |
|---|---|---|
| Model | distilgpt2 + LoRA on c_attn |
same |
| Host / members / controls | 10,000 / 16 / 100 | same |
| Repetitions | {1, 4, 16} | same |
| TPR @ 1% FPR | 0.000 (0/16) | 0.000 (0/16) |
| 95% CI | [0.000, 0.206] | [0.000, 0.206] |
| AUC (secondary) | 0.498 | 0.657 |
| Headline valid | yes | yes |
| Regurgitation | 0/16 | 0/16 |
| Negative-control regurgitation | 0.00 | 0.00 |
| Train / audit wall-clock | 142 s / 68 s | 274 s / 71 s |
reference.mode |
disable_adapter |
disable_adapter |
This is the opposite of the overfit demo: at an honest 0.93% canary budget, LoRA did not leak at 1% FPR. Run B's AUC rose (0.50 -> 0.66) so ranking moved, but the pre-declared headline stayed 0. That is a measured result, not a missing test. Reproduce:
pip install "memaudit[peft,dev]"
python benchmarks/run_lora_benchmark.py --n-host 10000 --n 16 --n-controls 100 --epochs 1 --lora-r 8Same pretrained distilgpt2 + LoRA on MPS, 2026-08-27. The script auto-grew the host to 80,000 rows to keep the canary budget at 0.77% (<=1%). Multi-seed stability (--seeds 0,1,2) included. Scale: distilgpt2 + LoRA.
| Metric | Run C (safe: 1 ep, r=8, lr 2e-4) | Run D (deliberately risky: 5 ep, r=16, lr 1e-3) |
|---|---|---|
| Host / members / controls | 80,000 / 100 / 200 | 80,000 / 100 / 200 |
| Canary token budget | 0.77% | 0.77% |
| TPR @ 1% FPR (primary calibration) | 0.000 (0/100) | 0.000 (0/100) |
| 95% CI | [0.000, 0.036] | [0.000, 0.036] |
| AUC (secondary) | 0.586 | 0.848 |
| Regurgitation / control regurgitation | 0/100 / 0.00 (n=200) | 0/100 / 0.00 (n=200) |
| Stability: per-seed TPR (seeds 0,1,2) | 0.000 / 0.010 / 0.010 (mean 0.007) | 0.000 / 0.090 / 0.090 (mean 0.060) |
| Train / audit wall-clock | 753 s / 173 s | 2,757 s / 170 s |
n=100 is the honest upgrade over n=16: zero detections now cap the true TPR at 3.6% with 95% confidence (vs 20.6% at n=16). Run D is why the risky config is labeled risky -- and why multi-seed mode exists: the AUC jumps 0.59 -> 0.85 (the member/control distributions clearly separated), and while the primary threshold calibration still lands at 0 detections, two of three bootstrap calibrations detect 9/100 canaries at 1% FPR. A single-seed run would have reported Run C and Run D as identical headlines; the stability block shows the risky config is sitting on the detection edge. No verbatim regurgitation in either run. See benchmarks/README.md for all rows and reproduce commands.
benchmarks/run_sft_benchmark.py runs the full claimed path on a live trl.SFTTrainer (TRL 0.29.1): prompt/completion dataset, completion_only_loss=True, LoRA r=8 on distilgpt2, inject() + MemorizationAuditCallback end-to-end. Measured 2026-08-27 on Apple MPS, host 10,000 records, canary budget 0.93% — same scale as Run A:
| Metric | SFT live run (1 ep, r=8, lr=2e-4) |
|---|---|
| Trainer | trl.SFTTrainer, completion_only_loss=True |
| Host / members / controls | 10,000 / 16 / 100 |
| Preflight survival scan | 16/16 found (9 token-level, 7 string-level fallback), 0 fully masked, 10,106 processed rows scanned |
| TPR @ 1% FPR | 0.000 (0/16), 95% CI [0.000, 0.206] |
| AUC (secondary) | 0.516 |
| Regurgitation / neg-control regurgitation | 0/16 / 0.00 (n=100) |
| Stability (seeds 0,1,2) | TPR mean/min/max 0.000 / 0.000 / 0.000 |
reference.mode |
disable_adapter |
| Train / audit wall-clock | 192 s / 52 s |
memaudit verify on the written report |
pass |
The value of this run is the integration evidence: TRL's tokenized prompt/completion pipeline kept all 16 canaries trainable (the survival scan found 7 of them via string-level fallback where BPE merged tokens across the prompt/completion boundary -- exactly the case the scan's fallback exists for), the callback audited an SFTTrainer-owned PEFT model via disable_adapter(), and the result matches the HF-Trainer run at the same scale. Reproduce:
pip install "memaudit[peft,trl,dev]"
python benchmarks/run_sft_benchmark.py --output-dir benchmarks/out-sft \
--n-host 10000 --n 16 --n-controls 100 --epochs 1 --seeds 0,1,2
# or as a gated test: MEMAUDIT_RUN_SFT=1 pytest -m integration| In scope | Out of scope |
|---|---|
| Membership inference (canary MIA, TPR @ 1% FPR + CI) | Model inversion / reconstruction |
| Prefix-prompted regurgitation (exact / BLEU / edit distance) | Attribute inference |
| LoRA / PEFT embedding-trainability pre-flight | Shadow-model LiRA, DP certificates |
| Set-level signal on a sample of your real records | Broad red-teaming, PII discovery |
Membership and regurgitation routinely disagree. A loss-only audit is the wrong answer in both directions; v0.1 always reports both.
Default canaries are high-perplexity regular tokens from the existing vocabulary. memaudit never resizes the vocab. The new-token family is gated and unimplemented in v0.1 (frozen-embedding LoRA leaves new rows untrained and the audit would silently measure noise).
Pre-flight blocks silent false confidence: wrong canary placement, fmt vs column mismatch, ShareGPT from/value, labels=-100 on the secret, canaries longer than max_length, empty inclusion coins, missing tokenizer. TPR@1%FPR is refused (not fabricated) when there are fewer than 100 held-out controls.
pip install memaudit # core: transformers, torch, datasets, numpy, scipy
pip install "memaudit[peft]" # LoRA / adapter-toggle scoring
pip install "memaudit[trl]" # SFTTrainer lint (optional)
pip install "memaudit[peft,trl]" # both extras
pip install "memaudit[hub]" # reserved for later model-card push
# from source:
git clone https://github.com/mem-audit/memaudit.git
cd memaudit && pip install -e ".[dev,peft,trl]"Requires Python 3.10+ and transformers>=4.56.2 (works on 5.x; the callback reads processing_class, not the removed tokenizer= kwarg).
pytest
memaudit demo --output-dir examplesfrom memaudit import generate_canaries, inject, MemorizationAuditCallback
canaries = generate_canaries(
tokenizer, n=32, n_controls=100, family="high_ppl",
repetitions=(1, 4, 16), seed=0,
)
train_ds, manifest = inject(train_ds, canaries, fmt="auto", seed=0)
# build SFTTrainer / Trainer on train_ds as usual
trainer.add_callback(
MemorizationAuditCallback(
trainer=trainer, manifest=manifest, real_sample=64, ref="auto",
)
)
trainer.train() # writes <output_dir>/memaudit-report.json
# ref="auto" is the LoRA one-copy path. Full FT: pass ref=<base model> or ref="none".Injection is a pre-train helper. It cannot live in the callback: transformers builds the dataloader before on_train_begin, and TRL tokenizes / loss-masks / packs inside SFTTrainer.__init__ before any hook fires.
The secret is always placed on the trainable side of the record (completion / assistant turn / text body). A prompt- or user-turn canary is labeled -100 under completion_only_loss / assistant_only_loss and would silently zero the audit - inject() refuses that placement.
Post-hoc / after a ZeRO-3 or FSDP run (in-callback scoring is deferred there):
memaudit audit --model ./out --canary-set ./out/memaudit-manifest.json \
--dataset ./train.jsonl --ref auto
# --manifest is an alias for --canary-set; both accept the inject() manifest--ref auto uses disable_adapter() on an unmerged LoRA so one model copy scores both fine-tuned and base. It refuses to silently fall back on a full fine-tune or a merged adapter: pass --ref <base-checkpoint> or explicit --ref none (target-only Min-K%++, labeled as a downgraded headline).
memaudit-report.json is schema 1.3.0 (schema_version; additive on 1.2.0 / 1.1.0 / 1.0.0 -- every earlier field is still there). Headline fields:
| Field | Meaning |
|---|---|
membership.headline_attack |
Pre-declared base-calibrated Min-K%++ (same two forwards also yield masked loss, loss ratio, Min-K%) |
membership.scorer |
Pluggable backend provenance: {name, version} (default min_k_plus_plus; EZ-MIA is a documented future file, not shipped) |
membership.tpr_at_1pct_fpr |
Detection rate on inserted canaries at the profile target FPR (default 1%), thresholded on held-out canary controls. null when headline_valid=false (underpowered controls, or audit_profile=smoke) |
membership.by_repetition |
Same threshold, split by 1x / 4x / 16x / pooled. 1x is a single-exposure probe; pooled is the powered-audit headline. Per-tier detected / tpr are null when the pooled threshold is unidentified (threshold_identified: false), not a fabricated 0.0 |
membership.calibration_stability |
Bootstrap-resample controls; how much the FPR threshold and TPR move (separate from the member-side CI) |
membership.ci_low / ci_high |
Clopper-Pearson 95% interval. With tens of canaries this interval is wide - that is honest |
membership.auc |
Secondary. Average-case; not the headline |
audit_profile |
smoke (refuses TPR@FPR headline) / routine / powered (or custom if inferred) plus target_fpr |
canaries.requested_family / actual_generator |
Requested construction vs what actually drew the tokens (a high_ppl run can be uniform_vocab) |
regurgitation.execution |
Run-level execution state: executed or not_run (reason: skip_generation). When not run, numerics are unmeasured (rate: null) and skipped rows do not enter a denominator |
regurgitation.overall.rate |
Fraction of inserted canaries the model completes from a 25% / 50% prefix (exact, BLEU>0.75, or sliding-window NED<=0.1). null when regurgitation was not run (skip_generation) |
regurgitation.detected |
Protocol-scoped exact-match count: "N/M under this prefix/decoding/exact-match protocol" -- never "no extraction risk". n: 0 / rate: null when not run |
regurgitation.by_tier |
Same rate at repetition 1 / 4 / 16. 1x is MIA-tier only. Empty object when regurgitation was not run |
negative_controls |
Never-inserted canaries. Membership scores always run; regurgitation on controls is skipped when --skip-generation is set (regurgitation_rate: null) |
real_records.execution |
executed when ranking ran; not_run with reason: no_dataset or real_sample_zero when it did not. A missing block on a hand-built report is not_recorded in the annex |
real_records.exact_dup_rate |
Exact-duplicate rate on extractable training texts. null when no extractable texts were found (not a measured 0.0) |
real_records.set_level |
Inferential member-vs-nonmember test only when held_out= is supplied (comparison_population: held_out). Otherwise descriptive_ranking_only or ranking_only (ran, but no comparison population -- not "skipped"). No FPR, not evidence about any individual record |
audit_seconds |
Wall-clock of the audit engine |
recommendations |
Heuristics (dedup -> fewer epochs -> cooler LoRA -> ...). Not a compliance program |
compliance_annex |
EDPB Opinion 28/2024 mapping: attack-coverage table (para 55), threat models (para 58(c)), test scope, release context (para 46), limitations. New in 1.1.0 |
release_context |
User-declared public-api / internal / open-weights (default unspecified). Never inferred |
stability |
Only with --seeds: multi-seed audit-procedure variance (null on single-seed runs) |
provenance |
Canary-manifest SHA-256, dataset fingerprint, model/adapter fingerprint, resolved config, python/torch/transformers versions |
report_sha256 |
Self-hash of the canonicalized report content, stamped at write time (+ <report>.sha256 sidecar) |
phone_home |
Always false |
local_only |
Always true |
Scores are computed on the secret span only. Full-sequence loss collapses detection.
EDPB-mapped annex. Every report carries a compliance_annex implementing the EDPB Opinion 28/2024 para 46 / para 55 / para 58 mapping: an attack-coverage table (membership inference para 55(i) and regurgitation para 55(iii) in scope with methods; attribute inference, exfiltration para 55(ii), model inversion para 55(iv), reconstruction para 55(v) explicitly out of scope), a threat model per attack and per canary family used (attacker access + assumptions, sourced from the published literature), test-scope metadata (n canaries, reps grid, seeds, dataset rows, negative-control results, run date, tool version), the user-declared release context, and a limitations statement quoting para 55: "successful testing which covers widely known, state-of-the-art attacks can only be evidence for the resistance to those attacks." The annex is documented test evidence -- it does not constitute a determination of anonymity or GDPR compliance. Render it as markdown for a DPO:
memaudit report --annex out/memaudit-report.json # markdown to stdout
memaudit report out/memaudit-report.json -o annex.md # or to a fileRelease context (para 46). Declare how the model will be exposed -- it changes which attack surface is "reasonably likely": --release-context public-api|internal|open-weights (API: run_audit(..., release_context=...) or MemorizationAuditCallback(..., release_context=...)). Default unspecified; the annex then says so.
Provenance + verify. Reports are self-hashed at write time: report_sha256 is the SHA-256 of the canonicalized report content (sorted keys, compact separators, minus the hash field), stamped into the JSON and into a <report>.sha256 sidecar. Check integrity later:
memaudit verify out/memaudit-report.json # exit 0 = intact, 1 = mismatchThis proves content integrity, not authorship. Cryptographic signing of the report file (GPG / sigstore) is a release-runbook step outside memaudit; memaudit does not implement key management.
Multi-seed mode. --seeds 0,1,2 (API: run_audit(..., seeds=[0,1,2])) adds a stability block. The model is trained once and canary scoring / greedy generation are deterministic, so what varies per seed is the randomness that actually exists in the audit procedure: bootstrap resampling of held-out control scores (threshold calibration) and real-record sampling. The block is labeled audit-procedure variance, not training variance (re-training across seeds is out of scope) and reports variance: {tpr_mean, tpr_min, tpr_max, tpr_std, per_seed: [...]}. Single-seed stays the default.
generate_canaries() # pure; no Trainer
inject() # raw dataset only
MemorizationAuditCallback
on_train_begin # PEFT pre-flight + survival scan (raises on silent-zero configs)
on_train_end # run_audit, or write a deferred CLI command under ZeRO-3/FSDP
run_audit() # shared engine (callback + CLI)
memaudit audit # post-hoc; --canary-set == --manifest == inject() JSON
memaudit demo # tiny overfit; measured metrics, not paper numbers
Ten landmines encoded in the implementation (source-checked against transformers 5.x / TRL / PEFT):
- No callback-time injection
- Secret never in the prompt / user turn
- No vocab resize
- Secret-span scoring
- No in-callback forwards under ZeRO-3 / FSDP
disable_adapter()skipped whenbias != "none"or merged- Standalone short canary records; warn on
wrappedpacking; skip first packed token model.eval()+inference_mode+ unwrapprocessing_classonly- Two verdicts, always
| Family | Construction |
|---|---|
high_ppl (default) |
Rejection-sample from the base model at high temperature into a PPL band. If no model is passed and a corpus is supplied, falls back to rare-token unigram; if no model and no corpus, falls back to uniform-from-vocab (recorded as actual_generator / metadata.source) |
unigram / bigram |
Least-likely tokens under corpus n-gram counts; uniform-from-vocab if no corpus |
structured |
CANARY-ID:... template + random fill (exposure metric later) |
random |
Uniform existing-vocab draws (also used as control twins) |
new_token |
Gated by the PEFT pre-flight — frozen embeddings cannot train new-token canaries; memaudit does not resize your vocab |
Defaults: 32 insert-eligible + 100 never-inserted controls (the TPR@1% FPR floor; the routine profile shape), 25-64 tokens, repetitions {1,4,16}, Bernoulli(1/2) inclusion coins. Named profiles: smoke (cheap, refuses the TPR@FPR headline), routine (those defaults), powered (100/200/{1,4,16}, calibration stability required). Going below 100 controls emits a warning and the report refuses the TPR@1% FPR headline unless you asked for smoke. The public powered case study stays 100/200/{1,4,16}.
- Small canary counts give wide CIs. Published audits use hundreds to thousands of canaries. v0.1 defaults are a CPU-friendly starting point, not a regulatory sample size.
- Thresholds are calibrated on this run's controls and do not transfer across model families.
- Real-record ranking is exploratory and descriptive: no FPR is attached, and it is not evidence about any individual record. Set-level member-vs-nonmember inference needs a user-supplied genuine held-out population.
- Black-box, final-model audits are structurally loose. A small TPR is not a privacy certificate.
- The README demo overfits on purpose (canaries ~ 99% of tokens). Your production run should stay near the 0.1% token-budget target.
- Multi-seed mode measures audit-procedure variance only (bootstrap threshold calibration + real-record sampling); re-training across seeds is out of scope.
- DPO / GRPO / Hub model-card push / PII flagging are out of scope for v0.1.
- LoRA-aware, not LoRA-only. Full fine-tunes need
--ref <base-checkpoint>or explicit--ref none. memaudit demo --loraneedsmemaudit[peft]and a transformersPreTrainedModel. The checked-in demo is full FT onTinyDemoLM.
| Piece | Buyer stack (LoRA bench) | Wheel install (clean venv) |
|---|---|---|
| Python | 3.12.11 | 3.12.11 |
| torch | 2.7.1 (PyPI, MPS) | 2.13.0 (PyPI, MPS) |
| transformers | 4.56.2 | 5.16.1 |
| peft | 0.20.0 | not installed (optional extra) |
| trl | 0.29.1 | not installed (optional extra) |
| datasets | 3.6.0 | 5.0.1 (pulled by pip install wheel) |
Known-bad combo: transformers 5.16.x + torch 2.6.dev hangs on FSDP imports (CPUOffloadPolicy). The hang is the dev torch, not 5.16 itself: a clean venv with transformers 5.16.1 + torch 2.13.0 imported and ran memaudit doctor here. Do not use --system-site-packages over a conda torch nightly. Recommended LoRA pin: transformers==4.56.2 + torch>=2.5,<2.8 + peft==0.20.0.
memaudit doctor --output-dir examples # env + tiny demo + schema
# or, if a report already exists:
memaudit doctor --skip-demo --report examples/demo-report.json
bash scripts/acceptance.shThe implementation module is memaudit.injection. The public helper remains from memaudit import inject.
Apache-2.0. Local execution is the product; a SaaS re-host does not capture it.