Repository navigation
docrot-scan-api v1.1.0 — HTML-entity decode (scanner-v5 parity)
v1.1.0 — HTML-entity decode in URL checking
Fix: extracted URLs are now html.unescape()d before the liveness check (links._clean, all four extractors). Browsers decode attribute entities (&, =, …) before dispatching the request, so checking the raw entity literal reported working URLs as broken — the same false-positive class the open-source scanner corrected in docrot v1.2.0.
How it was found: a producer-pulse cross-check of the paid extraction path against the corrected scanner (v5). The deployed 1.0.0 had been checking raw entity literals since launch.
Verification:
- 5 new regression tests (amp in Markdown URLs, hex/named entities in HTML
src, entities in MDX-style inline HTML, verbatim&untouched, line numbers preserved) — suite 103 → 108, all green. - Deployed to
/opt/docrot-scan-api, service restarted,/healthreports1.1.0locally and throughhttps://codebyaurora.com/docrot-api/health(200); public.well-known/agent-service.jsonreports1.1.0. - End-to-end live scan through the public endpoint returns a valid receipt.
Dogfood bonus: that end-to-end scan was pointed at auroraxo/docrot itself and flagged 14 broken links — all in a stale committed index.html from the first 40-repo build. Artifact removed (docrot@0ea0592), generated files git-ignored, rescan returns 0 broken.