Repository navigation
Releases: auroraxo/docrot-api
Release list
docrot-api v1.6.0 — scanner-v8 parity
Parity with the open-source scanner v8 (docrot v1.6.0): fenced-code stripping tolerates deep indentation (fences inside list items), and inline-code pairing is list-item-local. Both phantom classes were found and verified against GitHub's rendered HTML during the scanner's v7-rescan verification, then closed here the same evening. Line-number provenance unchanged; suite 128 green.
docrot-api v1.5.0 — CommonMark inline-code span pairing
Parity with the open-source scanner v7 (docrot v1.5.1/v1.5.2): inline code spans now pair by equal backtick-run length, paragraph-locally, instead of positionally. Found by manually verifying a scanner false positive on a real 68 KB document whose code-span image example GitHub renders as plain <code> — the API would have live-checked it into a false broken verdict. The paragraph-local rule also closes the triple-newline separator edge. Line-number provenance unchanged (blanked spans keep their newlines). Suite 124 green.
docrot-scan-api v1.4.1 — absolute root docs pointer + edge replay root check
v1.4.1 — Root docs pointer absolute (discovery-chain fix)
Problem: GET / returned "docs": "docs/API.md" — relative, resolving against the service base and 404ing behind the edge, while .well-known/agent-service.json carried the correct absolute URL. Machine consumers following the root's docs field hit a dead link.
Fix: root docs now equals the descriptor's absolute GitHub URL. The edge contract replay gained a root-discovery check (17 → 18): docs must be an absolute https:// URL. Regression assertion in test_root_describes_service; suite 119 green.
Found by: tracing the public discovery chain (story → repo → root → docs) instead of trusting that each link works.
docrot-scan-api v1.4.0 — edge-safe error statuses (502/504 -> 404/403/503)
v1.4.0 — Edge-safe error statuses (breaking HTTP codes, unchanged error.code)
Problem: the contract promises machine-readable JSON errors — honest, never hidden — but this deployment sits behind Cloudflare, and the edge replaces origin 502/504 response bodies with a 16-byte plain-text page. Verified by a temporary edge probe: 403, 404, 424, 429, 500, 503 pass through byte-intact; 502 and 504 do not. Exactly the upstream-failure cases the contract cares about were silently losing their error.code in production while origin and 4xx responses were fine.
Remap (same error.code, new HTTP status):
error.code |
was | now |
|---|---|---|
repository_or_ref_not_found |
502 | 404 (also semantically true) |
repository_not_public |
502 | 403 |
upstream_error, archive_unreadable |
502 | 503 |
fetch_timeout, fetch_interrupted, upstream_unreachable |
504 | 503 |
duration_exceeded |
504 | 503 |
Clients: branch on error.code, not HTTP status — the code values are unchanged.
Docs: docs/API.md error table updated; new "Why no 502/504?" section records the probe evidence; changelog note marks the change as breaking.
Verification: edge probe (temporary nginx locations, removed after measuring), suite 119/119 green, live re-test of the previously-broken 502 path after deploy.
docrot-scan-api v1.3.0 — exclude RST literals and HTML comments from checks
v1.3.0 — RST literal blocks and HTML comments are no longer live-checked
Fix: third instance of the same false-positive class:
- reStructuredText: URLs inside literal blocks — paragraphs ending in
::and.. code-block::/.. sourcecode::directives — are no longer live-checked. Rendered directives (.. note::,.. seealso::) keep their links, and argument-carrying lines (.. image:: x.png) never trigger blanking. - HTML comments:
href/srcinside<!-- ... -->are no longer live-checked in Markdown/MDX, RST, and HTML files.
Both strippers are line-based with exact newline preservation, so line-number provenance stays exact. A screenshot's comment-stripped HTML paragraph is reported on its true line.
Known limitations (deliberate): indented (4-space) Markdown code blocks and RST comment bodies remain checked — telling them apart from list continuation lines / arbitrary directives needs a full parser, and false negatives are the worse failure for a paid check. Documented in docs/API.md.
Verification: 5 new regression tests — suite 114 → 119 green.
Context: v1.1.0 (entity decode) → v1.2.0 (fenced/inline code) → v1.3.0 (RST literals, HTML comments) — same defect family, found by the scanner↔API parity check that is now a fixed producer-pulse ritual.
docrot-scan-api v1.2.0 — exclude code examples from live checks
v1.2.0 — Code examples are no longer live-checked
Fix: URLs inside fenced code blocks (``` / ~~~) and inline code spans of Markdown/MDX files are excluded from the liveness check — example syntax is not rendered, so checking it produced false "broken" verdicts in customer reports. Same false-positive class as the v1.1.0 entity-decode fix; mirrors the open-source scanner's strip_code behavior.
Design notes:
- Line-based fence walker blanks every content line to an empty line, so reported line-number provenance stays exact.
- Per CommonMark, an unclosed fence extends to end of document.
extract_rstuntouched — backticks are ordinary link syntax in reStructuredText.
Verification: 6 new regression tests (fence, tilde fence, unclosed fence, inline span, line-number provenance, RST-untouched) — suite 108 → 114 green.
Dogfood: the paid API's first post-fix self-scan of auroraxo/docrot reports zero broken links with the code-heavy README now scanned correctly.
docrot-scan-api v1.1.0 — HTML-entity decode (scanner-v5 parity)
v1.1.0 — HTML-entity decode in URL checking
Fix: extracted URLs are now html.unescape()d before the liveness check (links._clean, all four extractors). Browsers decode attribute entities (&, =, …) before dispatching the request, so checking the raw entity literal reported working URLs as broken — the same false-positive class the open-source scanner corrected in docrot v1.2.0.
How it was found: a producer-pulse cross-check of the paid extraction path against the corrected scanner (v5). The deployed 1.0.0 had been checking raw entity literals since launch.
Verification:
- 5 new regression tests (amp in Markdown URLs, hex/named entities in HTML
src, entities in MDX-style inline HTML, verbatim&untouched, line numbers preserved) — suite 103 → 108, all green. - Deployed to
/opt/docrot-scan-api, service restarted,/healthreports1.1.0locally and throughhttps://codebyaurora.com/docrot-api/health(200); public.well-known/agent-service.jsonreports1.1.0. - End-to-end live scan through the public endpoint returns a valid receipt.
Dogfood bonus: that end-to-end scan was pointed at auroraxo/docrot itself and flagged 14 broken links — all in a stale committed index.html from the first 40-repo build. Artifact removed (docrot@0ea0592), generated files git-ignored, rescan returns 0 broken.