Skip to content

v0.7.0: the container formats stop needing a server

Choose a tag to compare

@ivanusto ivanusto released this 16 Sep 05:52
· 10 commits to main since this release

SVG, HTML and Markdown are cleaned in the browser now, the engines run off the thread that draws the page, and one upstream defect is deliberately not copied.

The container formats stop needing a server

These three have been in the accept list since the first release, but only as text: Layer A ran over the body and a findings line said the metadata inside needed the Python service. The text-markup half of upstream's container_meta.py is pure regex and string work with no zipfile, no PDF, no subprocess and no network, so it runs in a page unchanged.

  • SVG: <metadata> and <x:xmpmeta> blocks, XML DOCTYPE and ENTITY declarations (#288), comments carrying AI markers, generator-like root attributes.
  • HTML: <meta> tags that name a generator, JSON-LD provenance blocks, data-ai* attributes. A plain CMS generator tag stays, which is the rule upstream #336 and #342 settled: a value names a generator only under a naming key, and everything else is free prose, so a description about "a static site generator" survives.
  • Markdown: AI frontmatter keys, with the nested lines under them. Dropping model: and leaving name: claude-opus behind is the failure this avoids.
  • All three: embedded data:image/… payloads run back through the PNG, JPEG, WebP and ISOBMFF strippers this port already had, then Layer A over the body, which is the order upstream's clean_container uses.

A file that is not valid UTF-8 comes back byte for byte

Upstream decodes these formats with surrogateescape and re-encodes at the end. The page used to read them through File.text(), which substitutes replacement characters, so such a file came back rewritten rather than cleaned. Both of upstream's decode modes are written out by hand in js/container_meta.js, because TextDecoder offers neither, and so are CPython's base64.b64decode, quote_from_bytes and unquote_to_bytes, whose rules differ from atob and encodeURIComponent in ways that change bytes. All of them are tested against CPython directly, over random bytes as well as named cases.

One upstream defect this port does not copy

_iter_script_blocks and _iter_data_uris lowercase the whole document to locate <script and data:image/, then index the original string with the result's offsets. That holds only while lowercasing preserves length, and U+0130 (LATIN CAPITAL LETTER I WITH DOT ABOVE) lowercases to two characters, in Python and in JavaScript alike.

With one such character earlier in the file, upstream's inspect_html reports nothing and its clean_html changes nothing for a document whose JSON-LD says trainedAlgorithmicMedia, and an embedded PNG parses as MIME type ng with a truncated payload. A single character of Turkish or Azerbaijani text is enough.

This port lowercases ASCII only, which cannot change length. Reported as guillaumemeyer/watermarks-remover#354. The divergence is asserted in both directions, so it fails loudly if upstream fixes this and the two converge, and it is recorded in scripts/upstream-sources.json rather than left in a commit message. This is the second such divergence, after the truncated-MP4 one upstream later took as #242.

The engines run in a Web Worker

A 64 MiB image is scanned and rebuilt in one synchronous pass, keyed-Gumbel does four pure-JS SHA-256 compressions per token, and the stylometry marker table walks the whole text once per pattern: 383 ms for a megabyte of ASCII, 1692 ms for the same text carrying zero-width characters, which is V8 taking the two-byte string path rather than anything this code does. For however long those took, nothing repainted and no button answered.

js/engine_ops.js holds the work as plain functions with no DOM in it, and both sides load it, so the worker path and the same-thread path run the same code rather than two copies that can drift. A File crosses by reference and the cleaned bytes come back transferred, so nothing is copied either way.

A page opened straight off the filesystem cannot start a worker, and the README says to open index.html that way, so the same op table runs in place there. Both paths are verified in headless Chromium against the real page, not against the modules.

Smaller things

  • The Inspector no longer refuses a recording the clean tab accepts. It read every input whole and applied the 64 MiB cap to audio and video too, while the clean tab has exempted them since v0.4.0 because of the slice driver. The detector input contract now carries either the bytes or the File, and the five metadata detectors share one pass through the headers.
  • A stale upstream checkout used to abort the whole test run. tests/test_contains_any_parity.py read a marker list at import time, outside its own skipif, so an AttributeError there ended collection for every file: 965 collected, 1 error. The lists are read through getattr now, and a set built from a missing list is reported as a skip rather than dropped in silence.
  • The dropzone said EPUB needed a server in one line and offered it in the next. Both READMEs left av_meta.py out of the line-for-line port list it has been in since v0.4.0. Two dictionary keys were dead in all three locales.
  • The Python re and str emulation moved out of js/stylometry.js into js/pyre.js, unchanged, because the container port needs the same rules; the stylometry parity suite is what proves it is the same code.
  • The nine script tags moved into the head with defer.

Parity anchors

container_meta.py is tracked as six slices rather than as a file: it is 4,000 lines and most of it is ZIP and PDF work with no counterpart here, so whole-file hashing would open a drift issue every time any of that moved. The check now covers twelve sources.

1403 tests against upstream fd44c13, no skips, up from 1035. 259 of the new ones are the SVG and data-URI port, whose fixtures are upstream's own hardening suite, all thirteen cases, compared byte for byte rather than asserted about.