Skip to content

Releases: auroraxo/docrot-api

docrot-api v1.6.0 — scanner-v8 parity

Choose a tag to compare

@auroraxo auroraxo released this 29 Sep 20:56

Parity with the open-source scanner v8 (docrot v1.6.0): fenced-code stripping tolerates deep indentation (fences inside list items), and inline-code pairing is list-item-local. Both phantom classes were found and verified against GitHub's rendered HTML during the scanner's v7-rescan verification, then closed here the same evening. Line-number provenance unchanged; suite 128 green.

docrot-api v1.5.0 — CommonMark inline-code span pairing

Choose a tag to compare

@auroraxo auroraxo released this 29 Sep 15:00

Parity with the open-source scanner v7 (docrot v1.5.1/v1.5.2): inline code spans now pair by equal backtick-run length, paragraph-locally, instead of positionally. Found by manually verifying a scanner false positive on a real 68 KB document whose code-span image example GitHub renders as plain <code> — the API would have live-checked it into a false broken verdict. The paragraph-local rule also closes the triple-newline separator edge. Line-number provenance unchanged (blanked spans keep their newlines). Suite 124 green.

docrot-scan-api v1.4.1 — absolute root docs pointer + edge replay root check

Choose a tag to compare

@auroraxo auroraxo released this 28 Sep 20:55

v1.4.1 — Root docs pointer absolute (discovery-chain fix)

Problem: GET / returned "docs": "docs/API.md" — relative, resolving against the service base and 404ing behind the edge, while .well-known/agent-service.json carried the correct absolute URL. Machine consumers following the root's docs field hit a dead link.

Fix: root docs now equals the descriptor's absolute GitHub URL. The edge contract replay gained a root-discovery check (17 → 18): docs must be an absolute https:// URL. Regression assertion in test_root_describes_service; suite 119 green.

Found by: tracing the public discovery chain (story → repo → root → docs) instead of trusting that each link works.

docrot-scan-api v1.4.0 — edge-safe error statuses (502/504 -> 404/403/503)

Choose a tag to compare

@auroraxo auroraxo released this 28 Sep 15:55

v1.4.0 — Edge-safe error statuses (breaking HTTP codes, unchanged error.code)

Problem: the contract promises machine-readable JSON errors — honest, never hidden — but this deployment sits behind Cloudflare, and the edge replaces origin 502/504 response bodies with a 16-byte plain-text page. Verified by a temporary edge probe: 403, 404, 424, 429, 500, 503 pass through byte-intact; 502 and 504 do not. Exactly the upstream-failure cases the contract cares about were silently losing their error.code in production while origin and 4xx responses were fine.

Remap (same error.code, new HTTP status):

error.code was now
repository_or_ref_not_found 502 404 (also semantically true)
repository_not_public 502 403
upstream_error, archive_unreadable 502 503
fetch_timeout, fetch_interrupted, upstream_unreachable 504 503
duration_exceeded 504 503

Clients: branch on error.code, not HTTP status — the code values are unchanged.

Docs: docs/API.md error table updated; new "Why no 502/504?" section records the probe evidence; changelog note marks the change as breaking.

Verification: edge probe (temporary nginx locations, removed after measuring), suite 119/119 green, live re-test of the previously-broken 502 path after deploy.

docrot-scan-api v1.3.0 — exclude RST literals and HTML comments from checks

Choose a tag to compare

@auroraxo auroraxo released this 28 Sep 12:43

v1.3.0 — RST literal blocks and HTML comments are no longer live-checked

Fix: third instance of the same false-positive class:

  • reStructuredText: URLs inside literal blocks — paragraphs ending in :: and .. code-block:: / .. sourcecode:: directives — are no longer live-checked. Rendered directives (.. note::, .. seealso::) keep their links, and argument-carrying lines (.. image:: x.png) never trigger blanking.
  • HTML comments: href/src inside <!-- ... --> are no longer live-checked in Markdown/MDX, RST, and HTML files.

Both strippers are line-based with exact newline preservation, so line-number provenance stays exact. A screenshot's comment-stripped HTML paragraph is reported on its true line.

Known limitations (deliberate): indented (4-space) Markdown code blocks and RST comment bodies remain checked — telling them apart from list continuation lines / arbitrary directives needs a full parser, and false negatives are the worse failure for a paid check. Documented in docs/API.md.

Verification: 5 new regression tests — suite 114 → 119 green.

Context: v1.1.0 (entity decode) → v1.2.0 (fenced/inline code) → v1.3.0 (RST literals, HTML comments) — same defect family, found by the scanner↔API parity check that is now a fixed producer-pulse ritual.

docrot-scan-api v1.2.0 — exclude code examples from live checks

Choose a tag to compare

@auroraxo auroraxo released this 28 Sep 11:10

v1.2.0 — Code examples are no longer live-checked

Fix: URLs inside fenced code blocks (``` / ~~~) and inline code spans of Markdown/MDX files are excluded from the liveness check — example syntax is not rendered, so checking it produced false "broken" verdicts in customer reports. Same false-positive class as the v1.1.0 entity-decode fix; mirrors the open-source scanner's strip_code behavior.

Design notes:

  • Line-based fence walker blanks every content line to an empty line, so reported line-number provenance stays exact.
  • Per CommonMark, an unclosed fence extends to end of document.
  • extract_rst untouched — backticks are ordinary link syntax in reStructuredText.

Verification: 6 new regression tests (fence, tilde fence, unclosed fence, inline span, line-number provenance, RST-untouched) — suite 108 → 114 green.

Dogfood: the paid API's first post-fix self-scan of auroraxo/docrot reports zero broken links with the code-heavy README now scanned correctly.

docrot-scan-api v1.1.0 — HTML-entity decode (scanner-v5 parity)

Choose a tag to compare

@auroraxo auroraxo released this 28 Sep 06:08

v1.1.0 — HTML-entity decode in URL checking

Fix: extracted URLs are now html.unescape()d before the liveness check (links._clean, all four extractors). Browsers decode attribute entities (&amp;, &#x3D;, …) before dispatching the request, so checking the raw entity literal reported working URLs as broken — the same false-positive class the open-source scanner corrected in docrot v1.2.0.

How it was found: a producer-pulse cross-check of the paid extraction path against the corrected scanner (v5). The deployed 1.0.0 had been checking raw entity literals since launch.

Verification:

  • 5 new regression tests (amp in Markdown URLs, hex/named entities in HTML src, entities in MDX-style inline HTML, verbatim & untouched, line numbers preserved) — suite 103 → 108, all green.
  • Deployed to /opt/docrot-scan-api, service restarted, /health reports 1.1.0 locally and through https://codebyaurora.com/docrot-api/health (200); public .well-known/agent-service.json reports 1.1.0.
  • End-to-end live scan through the public endpoint returns a valid receipt.

Dogfood bonus: that end-to-end scan was pointed at auroraxo/docrot itself and flagged 14 broken links — all in a stale committed index.html from the first 40-repo build. Artifact removed (docrot@0ea0592), generated files git-ignored, rescan returns 0 broken.