Skip to content

Releases: ForLegalAI/legaldown-validator

v0.3.0

Choose a tag to compare

@dvejsada dvejsada released this 29 Sep 20:35
9864ad8

The parser now reads CommonMark/GFM blocks as cmark-gfm does, the model keeps what it reads, diagnostics carry lines, and the release adds template assembly and a public API for renderers. Minor release before 1.0: the model and helpers change; read Breaking changes before upgrading. Spec target: LegalDown v0.2.

Breaking changes

  • List items are ListItems holding blocks (#64, #71), not strings; nested lists are blocks inside items. Use item_text, render_item, list_items.
  • New block kinds: heading (in items/quotes, with Block.level; #81), source (written as-is; #82), html (raw HTML block; #49).
  • Paragraph text holds \n between lines (#87); a line ending in \ or two spaces is a hard break.
  • document_to_dict changed shape: items are {"blocks": [...]}, blocks gain align/level. Migrate stored dicts, or read them back with document_from_dict.
  • Tables follow GFM (#77): empty cells/rows kept, Block.align per column.
  • Quote content is read as blocks (#72): code and HTML in quotes hold no directives.
  • Section numbers stay unique after a skipped heading level (#50).
  • lex() no longer blanks fenced code (#87): lex block_fragments(block) output.
  • Diagnostic gains line and file (#85): compare fields, not whole objects.
  • FrontmatterError for unreadable frontmatter (#73), with .line; scalar/list frontmatter is now body with frontmatter-absent, not an error.
  • Serializer writes incomplete frontmatter entries and party custom fields (#89), paragraphs as their lines, and new list/table formatting: expect diffs. empty_document() has no blank representatives.
  • CLI: file:line: text output sorted by line; JSON has line.

New

  • Assembly (§15.7, #55): assemble(), template_questions(), needed_questions(), legaldown assemble; answer rules answer-*.
  • Line numbers in diagnostics (§16.9, #85).
  • Public API for renderers (#90): ValidationResult.is_template, ValidationResult.placed_markers (PlacedMarker), is_template, is_drafting_note, lex, block_fragments/list_fragments (Fragment, ListFragment), condition helpers in legaldown.validator. The README now defines the public API and its versioning.
  • Rules: raw-html (#80), value-curly-quote, frontmatter-absent; 84 of 113 corpus rules.
  • Directives split across lines are directive-malformed (#87).

Fixes

CommonMark/GFM fidelity across lists, fences, HTML, tables and blank lines (#49–#77, #86); round trips of nested lists, include_signatures, open fences and frontmatter (#61–#69, #89); a parser hang on an indented table header (#87).

Full notes: see the pull requests linked above.

v0.2.0

Choose a tag to compare

@dvejsada dvejsada released this 23 Sep 17:41
2fef419

Targets specification v0.2 at conformance Level 1 — Core (§17.2). Rule ids are unchanged by the specification's renumbering (§15–18 became §16–19), so existing --ignore lists keep working.

What's new

  • Templates (§15). Questions and their answers, placeholders tied to questions, {{choose:}}, drafting notes (> [!DRAFTING]), insertion boundaries, and conditions on sections, list items, paragraphs, includes and attachments ({#id when=vat}, when: forum:courts).
  • Alternatives and reference safety (§15.4). Declarations that can never appear together may share an identifier. Every {{ref:}}, {{term:}} and {{attach:}} must resolve under every combination of answers in which it is present. A unit whose condition can never hold is reported.
  • Final documents (§15.9). legaldown validate --final (or validate_document(final=True)) reports blanks and template constructs left in a document meant for signature.
  • Automatic identifiers follow §5.3 exactly. NFKD normalization, the transliteration table, and collapsed hyphens: Smluvní pokuta → smluvni-pokuta, Haftungsausschluß → haftungsausschluss. Identifiers that lose letters without an ASCII form, such as Cyrillic or CJK, are reported.
  • Parsing. Setext headings, fenced code blocks kept whole, the preamble before the first heading (§4.4), and directives lexed by the §11.2 grammar.
  • New rules include condition-invalid, condition-never-true, condition-reference-unsafe, question-unused, choose-invalid, drafting-note-def, drafting-note-unrecognized, insertion-boundary, placeholder-unfilled, template-construct-present, anchor-lossy-slug, def-lossy-slug and legaldown-version-newer.

Breaking changes

Documents:

  • Automatic identifiers change for headings and terms with accented or special Latin letters, runs of hyphens, or long numeric text. A {{ref:}} written against the old identifier is now ref-broken. An automatic definition id with no usable text falls back to section (was term).
  • {{ref: a.b}} no longer resolves a section's dotted path, which §5.4 forbids.

Python API:

  • Section.identifier holds only an explicit {#id}. Generated identifiers are no longer written into the document; read them from ValidationResult.sections. Section gains condition.
  • The serializer writes a heading marker only where the source had one, so generated identifiers stay generated.
  • DefinitionRef.section_identifier is replaced by section_index and fragment_index; DefinitionRef gains lossy_id.
  • ensure_unique_identifier and DEF_ANCHOR_RE are removed. slugify_identifier no longer takes fallback; generate_identifier also returns whether letters were lost.
  • Block.format is removed. Metadata.supersedes is str | Amends (it was a string holding a Python repr). Metadata gains legaldown, questions and not_line_editable, and Attachment gains when.

Conformance

Verified against the specification's fixtures corpus: 79 of the corpus's 113 rules are implemented, and every one the corpus exercises at Core level passes. Per §17.5, the rules not implemented are named in CONFORMANCE.md, with the known limits: include fragments, attachment files and amended originals are not read; nested lists are flattened (#16).

Install

pip install --upgrade legaldown-validator

Python 3.11+. Only dependency: PyYAML.

v0.1.0

Choose a tag to compare

@dvejsada dvejsada released this 19 Aug 14:13
0978602

First release of the LegalDown reference implementation: a parser, document model, serializer, and validator for the LegalDown plain-text legal document format.

Targets specification v0.1 at conformance Level 1 — Core (§16.2): parse and validate a single document.

What's in it

  • Parser and document model — YAML frontmatter plus a heading/block document tree, with {#identifier} anchors on headings, list items, and paragraphs.
  • Validator — the §15 rule set at Core level. Every diagnostic carries the specification's stable rule id (§15.1) and severity, so tooling can suppress one check, escalate another, or gate a build on exactly the rules it cares about.
  • Serializer — Document back to LegalDown source.
  • Definitions (§7) — quoted term followed by {{def:}}, across the eight accepted quotation mark pairs, with auto-derived identifiers.
  • Command line — legaldown validate with plain-text and JSON output (§15.9), --ignore by rule id, --warnings-as-errors, --strict, recursive directory walking, and exit codes suitable for gating a build.
  • SPEC_VERSION and CONFORMANCE_LEVEL exported so consumers can assert what they validate against.

Frontmatter is read as exactly the shape §3 defines, and the parser never rewrites its input to make a document valid — what you authored is what the validator judges.

Conformance

Verified against the specification's own fixtures corpus: 57 of the corpus's 95 rules are implemented, and every one the corpus exercises at Core level passes. Per §16.5, the rules that are not implemented are named in CONFORMANCE.md rather than left to be discovered — they are the ones needing filesystem access: includes, attachment file contents, and bilingual document sets.

Install

pip install legaldown-validator

Python 3.11+. The distribution is legaldown-validator; the import name is legaldown. Only dependency: PyYAML. Ships a py.typed marker.