Skip to content

v0.4.0 — rendered-screen evaluation (developer-support use case)

Latest

Choose a tag to compare

@neerazz neerazz released this 27 Sep 23:59
· 2 commits to main since this release

PerceptFence v0.4.0: rendered-screen evaluation

Use case. A developer or support engineer shares their screen with an AI assistant to debug a failing build, a broken deploy, or a customer ticket. The screen shows terminals, .env files, CI logs, admin consoles, and chat toasts, and all of them can hold live secrets and personal data.

What's new

  • eval/screen/: 480 held-out developer-support screens across 8 screen types. Each is rendered by headless Chrome, either kept lossless or degraded (0.75x downscale + JPEG quality 55), and read back by tesseract OCR. Every defense receives the same OCR text.
  • Protocol and rule-file hashes were frozen before the test split ran (eval/screen/PROTOCOL.md, commit 502b97a). Three screen types were held out of development entirely.
  • Four new redaction families (T7–T10) for text as it looks after OCR. The v0.3 engine is still available as RedactionEngine(screen_families=False), so the two can be compared directly.
  • 50 tests (up from 42), including false-positive checks for the new families.

Results (test split, from supplement/screen_eval/test_summary.json)

Defense OCR-surviving secrets/PII neutralised
PerceptFence v0.4 889 / 968 (0.918; Wilson 95% 0.899–0.934)
Microsoft Presidio 0.581
gitleaks 0.179

On the three held-out screen types, PerceptFence neutralises 0.974. The cost there is 0.763 retention of task-relevant tokens, which is stated in the paper as the trade-off.

Limits. The screens are synthetic, and the evaluation covers only the deterministic redaction layer. It does not test live capture, a production model, or formal privacy guarantees.

Archive. Zenodo DOI 10.5281/zenodo.23004093 (source zip identical to this tag, plus the paper PDF). Note: this tag's pyproject.toml and CITATION.cff still read 0.3.0; that was corrected on main in 065f6b1. The code and results are unchanged.

Paper. paper/main.tex, Section 5.8. The preprint was submitted to arXiv (cs.CR) on 2026-09-27; this release will be updated with its identifier once it is announced.

Reproduce
Needs Google Chrome, tesseract, gitleaks, and Python 3.12.

python3.12 -m venv .evalvenv && .evalvenv/bin/pip install -r requirements-eval.txt
eval/screen/py.sh -m pytest -q
eval/screen/py.sh eval/screen/render_ocr.py --split test --out /tmp/ocr-test.jsonl
eval/screen/py.sh eval/screen/score.py --split test --ocr /tmp/ocr-test.jsonl --csv /tmp/test.csv --summary /tmp/test.json