Skip to content

docrot-scan-api v1.1.0 — HTML-entity decode (scanner-v5 parity)

Choose a tag to compare

@auroraxo auroraxo released this 28 Sep 06:08
· 14 commits to main since this release

v1.1.0 — HTML-entity decode in URL checking

Fix: extracted URLs are now html.unescape()d before the liveness check (links._clean, all four extractors). Browsers decode attribute entities (&, =, …) before dispatching the request, so checking the raw entity literal reported working URLs as broken — the same false-positive class the open-source scanner corrected in docrot v1.2.0.

How it was found: a producer-pulse cross-check of the paid extraction path against the corrected scanner (v5). The deployed 1.0.0 had been checking raw entity literals since launch.

Verification:

  • 5 new regression tests (amp in Markdown URLs, hex/named entities in HTML src, entities in MDX-style inline HTML, verbatim & untouched, line numbers preserved) — suite 103 → 108, all green.
  • Deployed to /opt/docrot-scan-api, service restarted, /health reports 1.1.0 locally and through https://codebyaurora.com/docrot-api/health (200); public .well-known/agent-service.json reports 1.1.0.
  • End-to-end live scan through the public endpoint returns a valid receipt.

Dogfood bonus: that end-to-end scan was pointed at auroraxo/docrot itself and flagged 14 broken links — all in a stale committed index.html from the first 40-repo build. Artifact removed (docrot@0ea0592), generated files git-ignored, rescan returns 0 broken.