Skip to content

Releases: SylphxAI/anymd

anymd 8.5.1

Choose a tag to compare

@github-actions github-actions released this 02 Oct 23:49
f1d369b

Patch Changes

  • anymd now re-checks the OCR engine program (anymd-ocr-vlm) every time it starts to use it, not only when it is installed, so a changed or replaced file is caught before it runs. A brief "text file busy" error while the engine is being refreshed no longer fails the run; anymd retries it.

  • When OCR runs on a page or image and finds no text, the result now says OCR found no text instead of looking like a page that was never read.

  • anymd pro buy [--no-browser] [--json] opens the Pro page; in-terminal purchase turns on when the checkout service is live. anymd pro activate and anymd pro status are unchanged, and existing tokens and token files keep working (same key, ANYMD_PRO_TOKEN, and <config dir>/anymd/pro-token). The licence code now comes from sylphx-mcp-kit 0.6, so anymd pro status also shows where the token was read and warns before a licence with an expiry runs out.

anymd 8.5.0

Choose a tag to compare

@github-actions github-actions released this 02 Oct 22:24
d80a930

Minor Changes

  • The default anymd binary is smaller: about 25 MB instead of 35 MB on Linux x64, with no change to anything it converts. The local VLM OCR engine moved out into its own small program, anymd-ocr-vlm (about 10 MB). anymd setup ocr downloads it in the same one-time step as the model weights; anymd checks a signature from a key built into anymd itself (bound to this exact version), downloads over HTTPS only, and checks the program's own version before installing it. --ocr vlm needs the same setup as before, and its output is unchanged. After an upgrade, anymd refreshes the 10 MB engine once by itself for anyone who already ran anymd setup ocr; the 2 GB model weights are never downloaded without asking. If you ask for VLM OCR before running anymd setup ocr, anymd says VLM OCR needs a one-time setup: run anymd setup ocr``; over MCP that is a normal reply rather than an error. Tesseract OCR and every other feature need no setup. Builds from source with --features ocr-vlm keep the engine inside the binary.

  • read now marks scanned PDF pages it could not read. A page with no text layer that images cover gets <!-- page N: scanned image, no text layer; enable OCR to read it: anymd setup ocr / ocr: true --> after its page marker, the front matter gains scanned_pages: [N, …], and a result made only of scans opens with a one-line hint, so agents can tell an empty page from an unread scan. Pages with text, and pages OCR reads, are unchanged.

anymd 8.4.0

Choose a tag to compare

@github-actions github-actions released this 02 Oct 06:53
97a5983

The anymd core stays free and open source (MIT), forever; nothing that was free before is now paid. 8.4.0 adds anymd Pro (US$29 once): video evidence and cite-check. Pro funds development.

Minor Changes

  • anymd Pro licence check. anymd pro status and anymd pro activate <token> manage an offline-verified (Ed25519) licence, read from ANYMD_PRO_TOKEN or <config dir>/anymd/pro-token. Pro unlocks only the new operations below; everything that was free before 8.4.0 stays free, MIT and ungated. Without a licence, the Pro operations return a short message with the link instead of doing any work. See anymd Pro.

  • (anymd Pro) Add bounded local video timelines and decoded frames to inspect, with matching read and outline projections. Cuts are heuristic FFmpeg scene-score 0.4 detections, not semantic scenes or confidence values. Chapters and timed subtitle/ASR cues retain their playback clock and source hash. Optional OCR and captions are sampled single-frame observations; OCR reuses the existing request permit, and captions need a user-configured local-command adapter, with no built-in captioning. Malformed subtitle files become a subtitles gap instead of failing the request. Frames, manifests and successful captions share the existing generated-image cache budget. Ordinary reads remain unchanged without timeline.

  • (anymd Pro) Add inspect operation cite_check for deterministic PDF quote and location checks. Exact, case-sensitive matching is the default; optional whitespace_v1 only collapses and trims whitespace. Results distinguish supported quotes, complete non-matches and insufficient evidence, keep native/OCR geometry provenance, and check source SHA-256 when requested. This checks extracted text at a location, not semantic truth or OCR accuracy.

anymd 8.3.0

Choose a tag to compare

@github-actions github-actions released this 01 Oct 11:08
d4b5074

Minor Changes

  • Add an opt-in local doc-VLM OCR route using PaddleOCR-VL-1.6 and PP-DocLayoutV3 on Candle, including Metal on Macs. anymd setup ocr installs SHA-256-pinned weights; ocr accepts auto, vlm and tesseract alongside the existing MCP booleans. Automatic OCR never downloads models. The worker has a hard page deadline, a per-region token cap, decode-time repetition stopping, shared OCR admission, and a fail-fast aggregate PDF deadline. Tables become Markdown and formulas become LaTeX. Experimental CPU q8/q4 decoder quantization is available through ANYMD_OCR_QUANTIZATION; Linux arm64 checks actual FP16 hardware and keeps the plain CLI on older machines.

  • Local transcripts now use bundled transcribe-cpp and one Qwen3-ASR-1.7B Q8 model for every language; the Whisper engine is removed. Pinned weights download only when explicitly requested and are SHA-256 verified. Long audio uses bounded 20-second chunks with source-relative segment timestamps. Optional Qwen3-ForcedAligner via standalone CrispASR supplies validated word timestamps where available. download_asr_model / --download-asr-model replaces the old spelling, which remains a Qwen-only compatibility alias. Japanese trails whisper-turbo on the 200-utterance FLEURS sample (5.93 vs 4.80 raw CER; 5.56 vs 4.59 with symmetric numeral/kana normalization), an accepted one-model trade-off. ASR benchmark tables publish both raw and explicitly named normalized scoring; they do not claim to reproduce Qwen's official scores.

  • New outline MCP tool and anymd outline <file> CLI command return a local, deterministic heading tree for PDF, DOCX, PPTX, EPUB, HTML and Markdown, as JSON or tree text. Nodes carry stable ids, title paths, page/slide/chapter ranges, Markdown byte ranges and child counts. read gains node (--node on the CLI), with the existing page selections, token budgets and cursors. Literal and ranked search hits carry node ids and title paths. Reads without a node keep their output unchanged.

  • The platform wheel includes a thin Python API: from anymd import convert returns a Document with Markdown text and source metadata. Optional langchain and llamaindex extras provide AnyMDLoader and AnyMDReader, using the same native converter without downloading another binary.

anymd 8.2.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 22:00
716534f

Minor Changes

  • pip install anymd and uvx anymd install the same native binary from PyPI (wheels for Linux x64/arm64, macOS x64/arm64 and Windows x64), and docker run ghcr.io/sylphxai/anymd runs it from a multi-arch image.
  • Every release carries SHA256SUMS, a CycloneDX SBOM and GitHub build-provenance attestations for its binaries and MCP Bundles; the install guide shows how to verify a download.
  • CI tests every change on Linux, macOS and Windows before it merges.
  • The benchmarks page adds OmniDocBench v1.6 for scans and page images.
  • The CLI prints one GitHub star line to stderr after the fifth successful interactive run, once ever. It is silent in MCP mode, in CI, and when stderr is not a terminal; ANYMD_NO_STAR_HINT=1 turns it off.
  • cargo install anymd works. Each release now also publishes the Rust crates to crates.io: anymd, anymd-core, anymd-formats, anymd-pdf, and two forks of upstream crates that carry our fixes, anymd-pdf-extract (from pdf-extract) and anymd-adobe-cmap-parser (from adobe-cmap-parser). The forks keep the upstream MIT licence and credit. npm stays the primary install.
  • sylphx-mcp-kit now comes from crates.io (0.2.3) instead of a git tag.
  • The Claude Desktop MCP Bundles (.mcpb) on each release now carry the anymd icon and a long description, and the install docs point at them.
  • read and the CLI gain images (--images refs|none). By default (refs) the raster images embedded in PDFs, DOCX, PPTX and EPUB files, which an agent cannot open from inside the container, are saved once each to <anymd cache>/images/<sha256>.<ext> and marked in the Markdown where they sit: ![caption](absolute path) and an <!-- image: WxH, page N --> comment. A PDF figure takes the caption line (Figure 3, Table 2, 圖1) directly below or above it, otherwise the alt text, otherwise image. Images under 48 x 48 px, PDF images under 2% of the page, and pictures repeated on three or more pages, slides or chapters (logos, headers) are skipped. Images over 50 megapixels are refused, and nothing is written beside the source. inspect structure also lists a PDF's exported images (embeddedImages). The image cache is pruned at start, at most daily: files unused for 30 days go, and the folder stays under 2 GiB. Vector-drawn figures and charts are not exported yet.
  • Word tracked changes and comments come out as CriticMarkup: insertions {++…++}, deletions {--…--}, a deletion next to an insertion as {~~old~>new~~}, and commented text as {==text==}{>>Author (date): comment<<}. Each tracked change is followed by its author and date, {++new++}{>>Author (date)<<}, the way CriticMarkup tracks several authors. Authors and dates are exactly as Word stores them. Inserted and deleted paragraph breaks, table rows and cells, text boxes, and footnotes are covered, and the output nests cleanly so simple CriticMarkup parsers read it. Contributed by @arthrod (#814).
  • read and the CLI gain revisions (--revisions markup|accept|reject) for Word files. markup (default) writes tracked changes and comments as CriticMarkup; accept and reject give the text as Word shows it after Accept All or Reject All, without markup or comments. A document with no tracked changes or comments converts exactly as before under all three.

AgentDocBench corpus v1

Choose a tag to compare

@shtse8 shtse8 released this 01 Oct 00:39
5051332

Byte-for-byte mirror of the existing 38-document AgentDocBench corpus at source commit 5051332. Every document was checked against the existing corpus.json SHA-256 and byte count before upload. No benchmark inputs, IDs, ground truth, reference text, scores or scoring changed.

This is a corpus-only release, not an anymd software release. Original source URLs, license evidence and attribution are in corpus.json and CORPUS-NOTICE. Each document retains its original public-domain or permissive license; Microsoft test files include the MIT notice, CC BY/CC BY-SA files retain attribution, and Project Gutenberg EPUBs retain embedded notices (public domain in the USA).

SHA256SUMS covers all 38 documents plus corpus.json and CORPUS-NOTICE. SHA256SUMS itself has SHA-256 67cbec8ef5e5d1ffe570c33fe883481d38dacda01a4d12a97c87135bc1e217fe. These assets are fixed: any future mirror release uses a new corpus tag.

anymd 8.1.0

Choose a tag to compare

@github-actions github-actions released this 26 Sep 02:31
3dbe810

Minor Changes

  • PDF tables are rebuilt from the page's drawn lines and aligned whitespace. A missing line between two cells makes a merged cell; wrapped cell text stays in its cell; stacked header lines become one header row, and a short header over several columns is repeated over each of them; rows banded between two lines split into one row per line; charts and boxes around whole passages are no longer mistaken for tables. Columns closer than a word gap apart are still split.
  • Scanned pages are read at 300 dpi, several at a time, and the words tesseract finds are laid out like a text page: paragraphs join across line breaks, hyphenated words rejoin, and tables come out as tables. A paragraph-final comma, a common misreading of a typewritten full stop, becomes a full stop.
  • Reading order handles three-column pages, and a table or figure inside one column of a two-column page no longer pulls in the other column.

Patch Changes

  • Text a reader cannot see is left out: invisible text (rendering mode 3) and text painted in the colour of the box behind it.
  • Text shown with the ' and " operators is no longer dropped, and pages whose text is turned (a landscape table on a portrait page) are laid out in their own direction.
  • Numbered paragraphs whose first line wraps are no longer read as headings.
  • A spreadsheet sheet with a name of its own (not "Sheet1") gets its name as a heading above its table.
  • The GitHub release carries an MCP Bundle (.mcpb) for one-click install in Claude Desktop.
  • crates/anymd-pdf is split into modules (extraction, rows, reading order, blocks, tables, OCR layout, rendering), each under 1,500 lines.

anymd 8.0.0

Choose a tag to compare

@github-actions github-actions released this 25 Sep 21:09
c32b5b7

Major Changes

  • The npm packages follow the mcp-kit layout. @sylphx/anymd is now only a small launcher (bin/anymd.js) that runs the binary from the matching @sylphx/anymd-<platform> package; the platform packages ship the binary at their root. The @sylphx/citra and @sylphx/pdf-reader-mcp aliases run the same launcher. MCP client configs (npx -y @sylphx/anymd) need no change.
  • Removed the @sylphx/anymd/sdk and @sylphx/anymd/pure-rust exports, the Anymd/Citra SDK classes, and the CITRA_RUST_BIN variable. Use the anymd CLI, or connect an MCP client to anymd mcp. The package no longer ships examples/ or corpus/.
  • The binary override is ANYMD_BIN=/path/to/anymd; ANYMD_RUST_BIN still works.

Minor Changes

  • anymd setup adds anymd to the MCP clients on this machine (Claude Code, Codex, Cursor, VS Code, Claude Desktop, Windsurf, Gemini CLI). --dry-run previews the changes, --remove undoes them, and running it again changes nothing. Try npx -y @sylphx/anymd setup.
  • anymd version prints anymd X.Y.Z. The binary, the npm packages and the MCP Registry entry now carry one version (the Rust crates were at 3.1.1).

Patch Changes

  • Releases publish through the shared mcp-kit release workflow: 5 native builds, npm trusted publishing, an npx smoke test, the GitHub release and the MCP Registry entry, whenever main carries a version that is not on npm yet. The Changesets flow and the release admission gate (release/admission.json, the capability matrix and its review specs) are retired.

v7.1.1

Choose a tag to compare

@sylphx-release-bot sylphx-release-bot released this 25 Sep 14:45
ed0eccd

What's Changed

  • fix(playground): show the anymd release version, not the crate version by @shtse8 in #760
  • ci(release): publish to npm through trusted publishing (OIDC), not a long-lived token by @shtse8 in #761
  • fix(http): only the exact health route skips the API key by @shtse8 in #762
  • chore: remove committed junk, source-text tests, and the duplicate native CI build by @shtse8 in #764
  • docs(copy): one source for descriptions and benchmark numbers by @shtse8 in #763
  • refactor(core): split document_twin into focused modules by @shtse8 in #765
  • ci: plain-language check on added lines by @shtse8 in #766
  • ci(publish): npm trusted publishing only; drop the token fallback by @shtse8 in #768
  • feat(bench): AgentDocBench, an open document-to-Markdown benchmark for agents by @shtse8 in #758
  • refactor: crates renamed to anymd-*; pdf-reader-cli folded into tests; inline test modules moved out by @shtse8 in #769
  • chore(release): version packages by @sylphx-release-bot[bot] in #767
  • chore(release): admit the 7.1.1 candidate for publish by @shtse8 in #770

Full Changelog: v7.1.0...v7.1.1

v7.1.0

Choose a tag to compare

@sylphx-release-bot sylphx-release-bot released this 25 Sep 11:17
c47b358

What's Changed

  • docs: link repomap from the README by @shtse8 in #753
  • refactor: PDF layout engine moves to its own crate (anymd-pdf, wasm-ready) by @shtse8 in #749
  • chore: retire the v3.0.14 differential harness and TypeScript oracle by @shtse8 in #752
  • feat(transcript): whisper model cache, verified opt-in download, OS install hints by @shtse8 in #750
  • fix: monospace text keeps its line breaks (receipts, code, terminal output) by @shtse8 in #751
  • docs: social preview image + real terminal demo GIF above the fold by @shtse8 in #754
  • fix(pdf): structured JSON text keeps word spaces on TeX PDFs by @shtse8 in #756
  • feat(docs): browser playground: any file → Markdown via WebAssembly by @shtse8 in #755
  • chore(release): version packages by @sylphx-release-bot[bot] in #757
  • chore(release): admit the 7.1.0 candidate for publish by @shtse8 in #759

Full Changelog: v7.0.0...v7.1.0