PerceptFence v0.4.0: rendered-screen evaluation
Use case. A developer or support engineer shares their screen with an AI assistant to debug a failing build, a broken deploy, or a customer ticket. The screen shows terminals, .env files, CI logs, admin consoles, and chat toasts, and all of them can hold live secrets and personal data.
What's new
eval/screen/: 480 held-out developer-support screens across 8 screen types. Each is rendered by headless Chrome, either kept lossless or degraded (0.75x downscale + JPEG quality 55), and read back by tesseract OCR. Every defense receives the same OCR text.- Protocol and rule-file hashes were frozen before the test split ran (
eval/screen/PROTOCOL.md, commit502b97a). Three screen types were held out of development entirely. - Four new redaction families (T7–T10) for text as it looks after OCR. The v0.3 engine is still available as
RedactionEngine(screen_families=False), so the two can be compared directly. - 50 tests (up from 42), including false-positive checks for the new families.
Results (test split, from supplement/screen_eval/test_summary.json)
| Defense | OCR-surviving secrets/PII neutralised |
|---|---|
| PerceptFence v0.4 | 889 / 968 (0.918; Wilson 95% 0.899–0.934) |
| Microsoft Presidio | 0.581 |
| gitleaks | 0.179 |
On the three held-out screen types, PerceptFence neutralises 0.974. The cost there is 0.763 retention of task-relevant tokens, which is stated in the paper as the trade-off.
Limits. The screens are synthetic, and the evaluation covers only the deterministic redaction layer. It does not test live capture, a production model, or formal privacy guarantees.
Archive. Zenodo DOI 10.5281/zenodo.23004093 (source zip identical to this tag, plus the paper PDF). Note: this tag's pyproject.toml and CITATION.cff still read 0.3.0; that was corrected on main in 065f6b1. The code and results are unchanged.
Paper. paper/main.tex, Section 5.8. The preprint was submitted to arXiv (cs.CR) on 2026-09-27; this release will be updated with its identifier once it is announced.
Reproduce
Needs Google Chrome, tesseract, gitleaks, and Python 3.12.
python3.12 -m venv .evalvenv && .evalvenv/bin/pip install -r requirements-eval.txt
eval/screen/py.sh -m pytest -q
eval/screen/py.sh eval/screen/render_ocr.py --split test --out /tmp/ocr-test.jsonl
eval/screen/py.sh eval/screen/score.py --split test --ocr /tmp/ocr-test.jsonl --csv /tmp/test.csv --summary /tmp/test.json