Releases: jlevy/flexdoc
Release list
v0.4.1
What's Changed
Changed
GIL CPython 3.11–3.14; refuse free-threaded 3.14t
requires-python is >=3.11,<3.15. Token diffs refuse a free-threaded build (Py_GIL_DISABLED) before importing cydifflib. Use --python 3.13 or a GIL 3.14 (uv python find 3.14 may resolve to 3.14t).
cydifflib is the optional diff extra
Token diffs still require cydifflib>=1.2.0 (no stdlib fallback). Core flexdoc no longer depends on it. Install flexdoc[diff]. Callers that use token diffs, including chopdiff, should depend on flexdoc[diff].
Full Changelog
v0.4.0
What's Changed
This release adds the portable reference and annotation layer from the 2026-07 TextRef design cycle plus cross-language logical word metrics, finalized under an independent senior release review (review and breaking-change catalog; PRs #18, #20, #21). Because it changes documented behavior on a pre-1.0 library, it bumps the minor version (per the pre-1.0 rule in docs/publishing.md). If you depend on an unbounded flexdoc>=0.3.0, review the breaking changes below or pin <0.4 before upgrading.
New Features
Native TextRef integration
Strict DocRef/TextRef values identify whole documents, spans, points, and semantic sections with canonical JSON and reversible textref:0.1 URIs. FlexDoc.references() binds a document locator to one source snapshot: it maps every locatable public value (for_target, for_span, for_point, for_section), resolves references with typed outcomes on independent document/source-validation/selector axes (resolution never silently picks a duplicate), retrieves structured line-window context, and renders deterministic annotation views for humans and LLMs. Span quote evidence is configurable per context or per span; compact exact-less spans stay bound to one source hash.
Consumer-owned annotations
TextAnnotation and the one-document AnnotationSet sidecar round-trip strict JSON and safe YAML, and every sidecar entry is validated through the complete TextRef evidence contract at parse time. DocGraph/v0.2 embeds a matching sidecar after verifying its document and source hash against the snapshot, and bounds-checks embedded positions. Committed JSON Schemas ship for both formats in the wheel.
Cross-language logical word metrics
TextUnit.words now measures normalized word-equivalent volume across natural language, CJK text, source code, URLs, and punctuation-dense content. TextUnit.raw_words and raw_word_count() keep literal whitespace counting, and logical_word_count() exposes the dependency-free primitive. See the logical-word definition and validation.
Breaking Changes
Made cleanly with no compatibility aliases, given the pre-1.0 status. Full catalog with rationales and migrations: review doc §4.
FlexDoc.graph()requires a document locator —doc.graph()becomesdoc.graph(document="path/or/id.md");build_doc_graph()and the debug helpersdoc_graph_yaml()/dump_views()take the same argument. Every serialized graph is now self-identifying.- DocGraph wire contract v0.1 → v0.2 —
schemais"DocGraph/v0.2";source.documentis required; the unqualifiedsource.sha256field is replaced by algorithm-qualifiedsource.source_hash(shared with TextRef); models are strict (unknown fields rejected, frozen, strict types) and validate node references, span/text consistency, and annotation bounds. TextUnit.wordsis logical — it matches raw counts for ordinary non-wide prose but differs for wide/fullwidth scripts, URLs, code, and symbolic content; useTextUnit.raw_wordsfor the previous whitespace-split behavior. Ripples throughsize(),size_summary(),section_size_tree(), and debug reports.- Token estimates scale logical words —
estimate_tokens(text, tokens_per_logical_word=1.6)replaces thechars_per_tokenparameter, andTOKENS_PER_LOGICAL_WORDreplacesCHARS_PER_TOKEN; estimates change numerically even at defaults. format_read_time()takes logical word counts — same signature; the default rate corresponds to roughly 450 CJK characters per minute under the default wide-character weight.- Pinned outputs change — regenerate any golden DocGraph/report snapshots (schema string,
sourceblock, andwordsfields).
Full Changelog
v0.3.0
What's Changed
This release stabilizes the pre-1.0 API boundary from the 2026-07 pre-promotion design review (senior engineering review, PRs #9 and #10). It fixes correctness bugs in source-coordinate handling and anchoring, and settles the remaining pre-1.0 API decisions so later releases can add annotation and synthetic-layer mechanisms without reopening the foundation. Because it changes documented behavior on a pre-1.0 library, it bumps the minor version (per the pre-1.0 rule in docs/publishing.md).
Bug Fixes
CRLF input no longer corrupts the structural views
marko computes block positions against LF-only text, so \r in the input desynchronized every structural span (blocks(), sections(), base_blocks(), links(), prose_text(), the node table) from source_text, silently garbling content. from_text now normalizes \r\n and lone \r to \n and retains the normalized string as source_text, so all layers share one offset space. Callers anchoring offsets to an external CRLF original must normalize it the same way first.
Markdown constructs inside frontmatter can no longer swallow the body
The shared parse previously included the frontmatter region, so e.g. a YAML block scalar containing a code fence opened a fenced block spanning the rest of the document, leaving blocks() empty. The frontmatter region is now blanked out of the shared parse (offsets preserved); frontmatter remains a non-content region.
resolve() no longer guesses on ambiguous quotes
A SpanRef quote that occurs multiple times with no disambiguating prefix/suffix (or a tied context score) now resolves to None instead of silently anchoring to the first occurrence. A context-free position hint also cannot select a duplicate, and a zero-width quote (exact="") resolves to None.
collect(overlaps=...) treats empty intervals as empty
A degenerate [x, x) region now overlaps nothing, matching half-open interval semantics; point queries use (x, x + 1).
Render helpers harden their HTML output
render_node_attrs attribute-escapes node.id, and wrap_with_node_attrs validates the tag name (raising ValueError), matching the flexdoc.html tag helpers.
Breaking Changes
Made cleanly with no compatibility aliases, given the preview (pre-1.0) status:
SpanRefowns its public resolution API — callref.resolve(source_text)/ref.resolve_and_update(source_text); the genericresolve/resolve_and_updatenames are no longer promoted fromflexdoc.docs.flexdoc.docsnow promotes the document model only — word-token/search and diff/mapping names moved to their owning modules (flexdoc.docs.wordtoks,search_tokens,token_diffs,token_mapping).- Recursive
collect()includes inline descendants by default —inlineis now tri-state; passinline=Falsefor the previous block-only recursive result. Paragraph.heading_level/Paragraph.heading_titleare properties — remove the()from calls.NAVIGABLE_LINK_FORMSreplacesTRUE_LINK_FORMS— no alias retained.TextUnitis aStrEnum—TextUnit.words == "words"now holds; onlystr()/equality behavior changes.- Cached structural views are mutation-safe —
Blockis frozen,Block.childrenandTableInfo.alignmentsare tuples, andsections()returns isolated copies; directBlockconstruction must pass child tuples.
Other Changes
graph()accepts anySetforinclude/detail, so plainsetliterals type-check.- Frontmatter delimiters tolerate trailing horizontal whitespace.
- The dependency lock is refreshed under the 14-day cool-off (cutoff 2026-06-26) with no per-package exceptions or audit ignores; CI runs
pip-auditclean. - The OS-independent classifier is backed by a macOS CI job (Python 3.13 full lint/test gate on
macos-latest). - Local release preparation is tag-aware, preventing a tagless clone from producing a
0.0.1.devNartifact.
Full Changelog
v0.2.0
What's Changed
This release lands the document-metrics use case (#6) for pprose: two correctness fixes plus a completed, typed inline/heading/link/prose surface, verified end-to-end against a 150-document corpus. Because it carries breaking signature changes on a pre-1.0 library, it bumps the minor version (per the pre-1.0 rule in docs/publishing.md).
Bug Fixes
node_table() / collect() / graph() no longer raise on valid Markdown
Inline elements were discovered over the whole source and parented by start offset, so backtick pairing across a block boundary (an empty fence next to inline backticks) produced an inline span that escaped its parent block and raised a layer-nesting error. Inline discovery is now scoped per leaf content block, so an inline node can never straddle a block boundary.
sections() / toc() recover every heading blocks() finds
Headings were re-derived from the blank-line paragraph view, which dropped tight headings and headings preceded by a non-blank line (e.g. an HTML-comment marker), and section content was bucketed from that same view, so a heading glued to its body lost the body. Sections now derive entirely from the structural block tree, so tight and marker-preceded headings own exactly their content. Section spans also nest correctly when a blank-line paragraph straddles a later heading.
Inline links near unbalanced backticks no longer misclassify as bare_url
links() scanned link atomic-spans over the whole document, so an unbalanced backtick run in one block could pair across a later block and swallow an intervening [text](url), dropping it to the bare-url fallback. The scan is now bounded per leaf block.
Reference-definition nodes attach to their block
A link_ref_def span included the line's trailing newline and escaped the containing paragraph; spans are now trimmed like every other structural block.
New Features
- Typed heading metadata —
Block.heading_info(HeadingInfowith parser-authoritativelevelandtitle) andBlock.heading_level;HeadingInfois exported fromflexdoc.docs. - Typed link forms —
LinkForm(inline/reference/autolink/bare_url/image/reference_definition) andLink.link_form.FlexDoc.links(link_forms=…)selects forms (default: navigable links only),FlexDoc.images()is a convenience accessor, and reference definitions ([id]: url) are surfaced asNodeKind.link_ref_defnodes. FlexDoc.prose_text()— prose-only text for editorial linting and prose metrics. Drops inline code, footnote refs, and inline-HTML tags (keeping the wrapped text); reduces links/images to their text/alt; strips heading/blockquote/list markers and reference-definition lines. Withinclude_tables=Trueit also flattens table cells. Slices come from verbatim source, so line wrapping and editorial spacing (e.g. a spaced em-dash) are preserved exactly.FlexDoc.block_at_offset()— the innermost structuralBlockcontaining an offset (the structural counterpart ofparagraph_at_offset).
Plus test-suite hardening: adversarial corpus docs, cross-projection invariants (toc-count == heading-block count, inline span ⊆ parent, link-form accounting), and a dogfood test over every Markdown file in the repo.
Breaking Changes
Made cleanly with no compatibility aliases, given the preview (pre-1.0) status:
Linkgains a requiredlink_form: LinkFormfield — directLink(...)construction must pass it.block_links()returns all link-like constructs (navigable links, images, and reference definitions), each tagged with alink_form; previously it returned navigable links only.FlexDoc.links()still filters to navigable links by default, so its default result is unchanged.collect()returns inline-kind nodes withoutrecursive=True— an inline-kind request (e.g.collect(kinds={NodeKind.link})) now widens the candidate set instead of returning[].- New
NodeKind.link_ref_defmember in the cross-languageDocGraphcontract.
Full Changelog
flexdoc 0.1.0
First release of flexdoc: a source-grounded, layered document model for Markdown and text, extracted from chopdiff as a standalone package.
What's Changed
Features
- The document model: parse a document into a
FlexDoc— the blank-line editing view (paragraphs and sentences with exact source offsets) plus derived structural views: the recursive block tree, the flat base-block partition, sections and TOC, links, and the cross-layer node table. One immutable source string and one Unicode-code-point offset space anchor everything. - One query primitive:
collect()gathers nodes by kind, predicate, subtree, or cross-layer offset containment (within/overlaps), in document order. - Serialization:
doc.graph()emits aDocGraph— a language-neutral JSON contract (DocGraph/v0.1) with composableincludelayers anddetailpayloads, plustoc/blocks/links/paragraphs/sentencesview indexes for UIs. - Durable references:
SpanRefanchors annotations and edits to quoted text (offsets as recomputable hints) and exports Chrome-style text fragments. - Sizes at every grain: bytes, lines, words, sentences, paragraphs, and approximate LLM token estimates via
TextUnit, rolled up per document or per section. - A deliberate root API:
from flexdoc import FlexDoc, DocGraph, Detail, SpanRef, BlockType, NodeKind, Layer, TextUnit, pinned by contract tests; render helpers for source-linked HTML are public inflexdoc.docs.
Documentation
- The design of record ships with the package: docs/flexdoc-spec.md, a standalone spec with first-principles terminology and per-layer error-handling coverage.
For users migrating from chopdiff's chopdiff.docs modules: the migration is one pass — chopdiff.docs.TextDoc becomes flexdoc.FlexDoc, chopdiff.docs.* becomes flexdoc.docs.*, with method renames listed in CHANGELOG.md.
Full commit history: https://github.com/jlevy/flexdoc/commits/v0.1.0