Skip to content

Releases: jlevy/flexdoc

v0.4.1

Choose a tag to compare

@jlevy jlevy released this 19 Sep 01:40
cbe2f5d

What's Changed

Changed

GIL CPython 3.11–3.14; refuse free-threaded 3.14t

requires-python is >=3.11,<3.15. Token diffs refuse a free-threaded build (Py_GIL_DISABLED) before importing cydifflib. Use --python 3.13 or a GIL 3.14 (uv python find 3.14 may resolve to 3.14t).

cydifflib is the optional diff extra

Token diffs still require cydifflib>=1.2.0 (no stdlib fallback). Core flexdoc no longer depends on it. Install flexdoc[diff]. Callers that use token diffs, including chopdiff, should depend on flexdoc[diff].

Full Changelog

v0.4.0...v0.4.1

v0.4.0

Choose a tag to compare

@github-actions github-actions released this 20 Jul 23:47
60a19b0

What's Changed

This release adds the portable reference and annotation layer from the 2026-07 TextRef design cycle plus cross-language logical word metrics, finalized under an independent senior release review (review and breaking-change catalog; PRs #18, #20, #21). Because it changes documented behavior on a pre-1.0 library, it bumps the minor version (per the pre-1.0 rule in docs/publishing.md). If you depend on an unbounded flexdoc>=0.3.0, review the breaking changes below or pin <0.4 before upgrading.

New Features

Native TextRef integration

Strict DocRef/TextRef values identify whole documents, spans, points, and semantic sections with canonical JSON and reversible textref:0.1 URIs. FlexDoc.references() binds a document locator to one source snapshot: it maps every locatable public value (for_target, for_span, for_point, for_section), resolves references with typed outcomes on independent document/source-validation/selector axes (resolution never silently picks a duplicate), retrieves structured line-window context, and renders deterministic annotation views for humans and LLMs. Span quote evidence is configurable per context or per span; compact exact-less spans stay bound to one source hash.

Consumer-owned annotations

TextAnnotation and the one-document AnnotationSet sidecar round-trip strict JSON and safe YAML, and every sidecar entry is validated through the complete TextRef evidence contract at parse time. DocGraph/v0.2 embeds a matching sidecar after verifying its document and source hash against the snapshot, and bounds-checks embedded positions. Committed JSON Schemas ship for both formats in the wheel.

Cross-language logical word metrics

TextUnit.words now measures normalized word-equivalent volume across natural language, CJK text, source code, URLs, and punctuation-dense content. TextUnit.raw_words and raw_word_count() keep literal whitespace counting, and logical_word_count() exposes the dependency-free primitive. See the logical-word definition and validation.

Breaking Changes

Made cleanly with no compatibility aliases, given the pre-1.0 status. Full catalog with rationales and migrations: review doc §4.

  • FlexDoc.graph() requires a document locator — doc.graph() becomes doc.graph(document="path/or/id.md"); build_doc_graph() and the debug helpers doc_graph_yaml()/dump_views() take the same argument. Every serialized graph is now self-identifying.
  • DocGraph wire contract v0.1 → v0.2 — schema is "DocGraph/v0.2"; source.document is required; the unqualified source.sha256 field is replaced by algorithm-qualified source.source_hash (shared with TextRef); models are strict (unknown fields rejected, frozen, strict types) and validate node references, span/text consistency, and annotation bounds.
  • TextUnit.words is logical — it matches raw counts for ordinary non-wide prose but differs for wide/fullwidth scripts, URLs, code, and symbolic content; use TextUnit.raw_words for the previous whitespace-split behavior. Ripples through size(), size_summary(), section_size_tree(), and debug reports.
  • Token estimates scale logical words — estimate_tokens(text, tokens_per_logical_word=1.6) replaces the chars_per_token parameter, and TOKENS_PER_LOGICAL_WORD replaces CHARS_PER_TOKEN; estimates change numerically even at defaults.
  • format_read_time() takes logical word counts — same signature; the default rate corresponds to roughly 450 CJK characters per minute under the default wide-character weight.
  • Pinned outputs change — regenerate any golden DocGraph/report snapshots (schema string, source block, and words fields).

Full Changelog

v0.3.0...v0.4.0

v0.3.0

Choose a tag to compare

@github-actions github-actions released this 11 Jul 02:00
74dff0c

What's Changed

This release stabilizes the pre-1.0 API boundary from the 2026-07 pre-promotion design review (senior engineering review, PRs #9 and #10). It fixes correctness bugs in source-coordinate handling and anchoring, and settles the remaining pre-1.0 API decisions so later releases can add annotation and synthetic-layer mechanisms without reopening the foundation. Because it changes documented behavior on a pre-1.0 library, it bumps the minor version (per the pre-1.0 rule in docs/publishing.md).

Bug Fixes

CRLF input no longer corrupts the structural views

marko computes block positions against LF-only text, so \r in the input desynchronized every structural span (blocks(), sections(), base_blocks(), links(), prose_text(), the node table) from source_text, silently garbling content. from_text now normalizes \r\n and lone \r to \n and retains the normalized string as source_text, so all layers share one offset space. Callers anchoring offsets to an external CRLF original must normalize it the same way first.

Markdown constructs inside frontmatter can no longer swallow the body

The shared parse previously included the frontmatter region, so e.g. a YAML block scalar containing a code fence opened a fenced block spanning the rest of the document, leaving blocks() empty. The frontmatter region is now blanked out of the shared parse (offsets preserved); frontmatter remains a non-content region.

resolve() no longer guesses on ambiguous quotes

A SpanRef quote that occurs multiple times with no disambiguating prefix/suffix (or a tied context score) now resolves to None instead of silently anchoring to the first occurrence. A context-free position hint also cannot select a duplicate, and a zero-width quote (exact="") resolves to None.

collect(overlaps=...) treats empty intervals as empty

A degenerate [x, x) region now overlaps nothing, matching half-open interval semantics; point queries use (x, x + 1).

Render helpers harden their HTML output

render_node_attrs attribute-escapes node.id, and wrap_with_node_attrs validates the tag name (raising ValueError), matching the flexdoc.html tag helpers.

Breaking Changes

Made cleanly with no compatibility aliases, given the preview (pre-1.0) status:

  • SpanRef owns its public resolution API — call ref.resolve(source_text) / ref.resolve_and_update(source_text); the generic resolve / resolve_and_update names are no longer promoted from flexdoc.docs.
  • flexdoc.docs now promotes the document model only — word-token/search and diff/mapping names moved to their owning modules (flexdoc.docs.wordtoks, search_tokens, token_diffs, token_mapping).
  • Recursive collect() includes inline descendants by default — inline is now tri-state; pass inline=False for the previous block-only recursive result.
  • Paragraph.heading_level / Paragraph.heading_title are properties — remove the () from calls.
  • NAVIGABLE_LINK_FORMS replaces TRUE_LINK_FORMS — no alias retained.
  • TextUnit is a StrEnum — TextUnit.words == "words" now holds; only str()/equality behavior changes.
  • Cached structural views are mutation-safe — Block is frozen, Block.children and TableInfo.alignments are tuples, and sections() returns isolated copies; direct Block construction must pass child tuples.

Other Changes

  • graph() accepts any Set for include/detail, so plain set literals type-check.
  • Frontmatter delimiters tolerate trailing horizontal whitespace.
  • The dependency lock is refreshed under the 14-day cool-off (cutoff 2026-06-26) with no per-package exceptions or audit ignores; CI runs pip-audit clean.
  • The OS-independent classifier is backed by a macOS CI job (Python 3.13 full lint/test gate on macos-latest).
  • Local release preparation is tag-aware, preventing a tagless clone from producing a 0.0.1.devN artifact.

Full Changelog

v0.2.0...v0.3.0

v0.2.0

Choose a tag to compare

@jlevy jlevy released this 14 Jun 01:43
8c13171

What's Changed

This release lands the document-metrics use case (#6) for pprose: two correctness fixes plus a completed, typed inline/heading/link/prose surface, verified end-to-end against a 150-document corpus. Because it carries breaking signature changes on a pre-1.0 library, it bumps the minor version (per the pre-1.0 rule in docs/publishing.md).

Bug Fixes

node_table() / collect() / graph() no longer raise on valid Markdown

Inline elements were discovered over the whole source and parented by start offset, so backtick pairing across a block boundary (an empty fence next to inline backticks) produced an inline span that escaped its parent block and raised a layer-nesting error. Inline discovery is now scoped per leaf content block, so an inline node can never straddle a block boundary.

sections() / toc() recover every heading blocks() finds

Headings were re-derived from the blank-line paragraph view, which dropped tight headings and headings preceded by a non-blank line (e.g. an HTML-comment marker), and section content was bucketed from that same view, so a heading glued to its body lost the body. Sections now derive entirely from the structural block tree, so tight and marker-preceded headings own exactly their content. Section spans also nest correctly when a blank-line paragraph straddles a later heading.

Inline links near unbalanced backticks no longer misclassify as bare_url

links() scanned link atomic-spans over the whole document, so an unbalanced backtick run in one block could pair across a later block and swallow an intervening [text](url), dropping it to the bare-url fallback. The scan is now bounded per leaf block.

Reference-definition nodes attach to their block

A link_ref_def span included the line's trailing newline and escaped the containing paragraph; spans are now trimmed like every other structural block.

New Features

  • Typed heading metadata — Block.heading_info (HeadingInfo with parser-authoritative level and title) and Block.heading_level; HeadingInfo is exported from flexdoc.docs.
  • Typed link forms — LinkForm (inline / reference / autolink / bare_url / image / reference_definition) and Link.link_form. FlexDoc.links(link_forms=…) selects forms (default: navigable links only), FlexDoc.images() is a convenience accessor, and reference definitions ([id]: url) are surfaced as NodeKind.link_ref_def nodes.
  • FlexDoc.prose_text() — prose-only text for editorial linting and prose metrics. Drops inline code, footnote refs, and inline-HTML tags (keeping the wrapped text); reduces links/images to their text/alt; strips heading/blockquote/list markers and reference-definition lines. With include_tables=True it also flattens table cells. Slices come from verbatim source, so line wrapping and editorial spacing (e.g. a spaced em-dash) are preserved exactly.
  • FlexDoc.block_at_offset() — the innermost structural Block containing an offset (the structural counterpart of paragraph_at_offset).

Plus test-suite hardening: adversarial corpus docs, cross-projection invariants (toc-count == heading-block count, inline span ⊆ parent, link-form accounting), and a dogfood test over every Markdown file in the repo.

Breaking Changes

Made cleanly with no compatibility aliases, given the preview (pre-1.0) status:

  • Link gains a required link_form: LinkForm field — direct Link(...) construction must pass it.
  • block_links() returns all link-like constructs (navigable links, images, and reference definitions), each tagged with a link_form; previously it returned navigable links only. FlexDoc.links() still filters to navigable links by default, so its default result is unchanged.
  • collect() returns inline-kind nodes without recursive=True — an inline-kind request (e.g. collect(kinds={NodeKind.link})) now widens the candidate set instead of returning [].
  • New NodeKind.link_ref_def member in the cross-language DocGraph contract.

Full Changelog

v0.1.0...v0.2.0

flexdoc 0.1.0

Choose a tag to compare

@jlevy jlevy released this 12 Jun 21:02
c94c977

First release of flexdoc: a source-grounded, layered document model for Markdown and text, extracted from chopdiff as a standalone package.

What's Changed

Features

  • The document model: parse a document into a FlexDoc — the blank-line editing view (paragraphs and sentences with exact source offsets) plus derived structural views: the recursive block tree, the flat base-block partition, sections and TOC, links, and the cross-layer node table. One immutable source string and one Unicode-code-point offset space anchor everything.
  • One query primitive: collect() gathers nodes by kind, predicate, subtree, or cross-layer offset containment (within/overlaps), in document order.
  • Serialization: doc.graph() emits a DocGraph — a language-neutral JSON contract (DocGraph/v0.1) with composable include layers and detail payloads, plus toc/blocks/links/paragraphs/sentences view indexes for UIs.
  • Durable references: SpanRef anchors annotations and edits to quoted text (offsets as recomputable hints) and exports Chrome-style text fragments.
  • Sizes at every grain: bytes, lines, words, sentences, paragraphs, and approximate LLM token estimates via TextUnit, rolled up per document or per section.
  • A deliberate root API: from flexdoc import FlexDoc, DocGraph, Detail, SpanRef, BlockType, NodeKind, Layer, TextUnit, pinned by contract tests; render helpers for source-linked HTML are public in flexdoc.docs.

Documentation

  • The design of record ships with the package: docs/flexdoc-spec.md, a standalone spec with first-principles terminology and per-layer error-handling coverage.

For users migrating from chopdiff's chopdiff.docs modules: the migration is one pass — chopdiff.docs.TextDoc becomes flexdoc.FlexDoc, chopdiff.docs.* becomes flexdoc.docs.*, with method renames listed in CHANGELOG.md.

Full commit history: https://github.com/jlevy/flexdoc/commits/v0.1.0