Skip to content

Gather 1.7.1: Corpus newline integrity

Choose a tag to compare

@HarperZ9 HarperZ9 released this 08 Sep 23:14
· 50 commits to main since this release
cb35c6d

Gather 1.7.1

Gather 1.7.1 is a patch release for corpus newline integrity and readable-context provenance.

It resolves the inherited CR/LF and mixed-newline corpus integrity mismatch documented in 1.7.0 by writing new corpus objects as exact UTF-8 bytes and sealing a versioned storage witness into new catalog rows. Readable context now keeps the exact source-text identity separate from the LF-normalized readable view: sha256, verified_sha256, and source_sha256 bind the exact source text, while view_sha256 and view_codec describe the normalized view used for excerpt display and Python character offsets.

This patch also fixes a writer admission edge case found during independent review: adding a receipt now refuses a corrupt preexisting object at the content-addressed path before appending to the catalog. Valid legacy objects are still reused without rewriting when Gather can reconstruct exactly one matching source text from exact UTF-8 bytes or the old Windows LF-to-CRLF expansion.

Validation for this candidate used the full Gather pytest suite, ruff, mypy, source gather status/doctor, no-index wheel installs on Windows and WSL, CLI/MCP context selection, and tamper refusal checks.

Boundaries: legacy reconstruction verifies source text for an existing receipt, not historical raw object-byte integrity. Storage-witness stripping or codec changes are downgrade-detectable only when a consumer pins or independently verifies the prior corpus digest with --expect-digest / expected_corpus_digest. Readable context is acquired source context; it does not prove source truth, claim support, completeness, downstream model use, caller workspace-parent resolution, or absence of sensitive material in selected source text or URLs.