Repository navigation
Releases: SylphxAI/anymd
Release list
anymd 8.5.1
Patch Changes
-
anymd now re-checks the OCR engine program (
anymd-ocr-vlm) every time it starts to use it, not only when it is installed, so a changed or replaced file is caught before it runs. A brief "text file busy" error while the engine is being refreshed no longer fails the run; anymd retries it. -
When OCR runs on a page or image and finds no text, the result now says
OCR found no textinstead of looking like a page that was never read. -
anymd pro buy [--no-browser] [--json]opens the Pro page; in-terminal purchase turns on when the checkout service is live.anymd pro activateandanymd pro statusare unchanged, and existing tokens and token files keep working (same key,ANYMD_PRO_TOKEN, and<config dir>/anymd/pro-token). The licence code now comes fromsylphx-mcp-kit0.6, soanymd pro statusalso shows where the token was read and warns before a licence with an expiry runs out.
anymd 8.5.0
Minor Changes
-
The default
anymdbinary is smaller: about 25 MB instead of 35 MB on Linux x64, with no change to anything it converts. The local VLM OCR engine moved out into its own small program,anymd-ocr-vlm(about 10 MB).anymd setup ocrdownloads it in the same one-time step as the model weights; anymd checks a signature from a key built into anymd itself (bound to this exact version), downloads over HTTPS only, and checks the program's own version before installing it.--ocr vlmneeds the same setup as before, and its output is unchanged. After an upgrade, anymd refreshes the 10 MB engine once by itself for anyone who already rananymd setup ocr; the 2 GB model weights are never downloaded without asking. If you ask for VLM OCR before runninganymd setup ocr, anymd saysVLM OCR needs a one-time setup: runanymd setup ocr``; over MCP that is a normal reply rather than an error. Tesseract OCR and every other feature need no setup. Builds from source with--features ocr-vlmkeep the engine inside the binary. -
readnow marks scanned PDF pages it could not read. A page with no text layer that images cover gets<!-- page N: scanned image, no text layer; enable OCR to read it: anymd setup ocr / ocr: true -->after its page marker, the front matter gainsscanned_pages: [N, …], and a result made only of scans opens with a one-line hint, so agents can tell an empty page from an unread scan. Pages with text, and pages OCR reads, are unchanged.
anymd 8.4.0
The anymd core stays free and open source (MIT), forever; nothing that was free before is now paid. 8.4.0 adds anymd Pro (US$29 once): video evidence and cite-check. Pro funds development.
Minor Changes
-
anymd Pro licence check.
anymd pro statusandanymd pro activate <token>manage an offline-verified (Ed25519) licence, read fromANYMD_PRO_TOKENor<config dir>/anymd/pro-token. Pro unlocks only the new operations below; everything that was free before 8.4.0 stays free, MIT and ungated. Without a licence, the Pro operations return a short message with the link instead of doing any work. See anymd Pro. -
(anymd Pro) Add bounded local video timelines and decoded frames to
inspect, with matchingreadandoutlineprojections. Cuts are heuristic FFmpeg scene-score 0.4 detections, not semantic scenes or confidence values. Chapters and timed subtitle/ASR cues retain their playback clock and source hash. Optional OCR and captions are sampled single-frame observations; OCR reuses the existing request permit, and captions need a user-configured local-command adapter, with no built-in captioning. Malformed subtitle files become a subtitles gap instead of failing the request. Frames, manifests and successful captions share the existing generated-image cache budget. Ordinary reads remain unchanged withouttimeline. -
(anymd Pro) Add
inspectoperationcite_checkfor deterministic PDF quote and location checks. Exact, case-sensitive matching is the default; optionalwhitespace_v1only collapses and trims whitespace. Results distinguish supported quotes, complete non-matches and insufficient evidence, keep native/OCR geometry provenance, and check source SHA-256 when requested. This checks extracted text at a location, not semantic truth or OCR accuracy.
anymd 8.3.0
Minor Changes
-
Add an opt-in local doc-VLM OCR route using PaddleOCR-VL-1.6 and PP-DocLayoutV3 on Candle, including Metal on Macs.
anymd setup ocrinstalls SHA-256-pinned weights;ocraccepts auto, vlm and tesseract alongside the existing MCP booleans. Automatic OCR never downloads models. The worker has a hard page deadline, a per-region token cap, decode-time repetition stopping, shared OCR admission, and a fail-fast aggregate PDF deadline. Tables become Markdown and formulas become LaTeX. Experimental CPU q8/q4 decoder quantization is available throughANYMD_OCR_QUANTIZATION; Linux arm64 checks actual FP16 hardware and keeps the plain CLI on older machines. -
Local transcripts now use bundled transcribe-cpp and one Qwen3-ASR-1.7B Q8 model for every language; the Whisper engine is removed. Pinned weights download only when explicitly requested and are SHA-256 verified. Long audio uses bounded 20-second chunks with source-relative segment timestamps. Optional Qwen3-ForcedAligner via standalone CrispASR supplies validated word timestamps where available.
download_asr_model/--download-asr-modelreplaces the old spelling, which remains a Qwen-only compatibility alias. Japanese trails whisper-turbo on the 200-utterance FLEURS sample (5.93 vs 4.80 raw CER; 5.56 vs 4.59 with symmetric numeral/kana normalization), an accepted one-model trade-off. ASR benchmark tables publish both raw and explicitly named normalized scoring; they do not claim to reproduce Qwen's official scores. -
New
outlineMCP tool andanymd outline <file>CLI command return a local, deterministic heading tree for PDF, DOCX, PPTX, EPUB, HTML and Markdown, as JSON or tree text. Nodes carry stable ids, title paths, page/slide/chapter ranges, Markdown byte ranges and child counts.readgainsnode(--nodeon the CLI), with the existing page selections, token budgets and cursors. Literal and rankedsearchhits carry node ids and title paths. Reads without a node keep their output unchanged. -
The platform wheel includes a thin Python API:
from anymd import convertreturns aDocumentwith Markdown text and source metadata. Optionallangchainandllamaindexextras provideAnyMDLoaderandAnyMDReader, using the same native converter without downloading another binary.
anymd 8.2.0
Minor Changes
pip install anymdanduvx anymdinstall the same native binary from PyPI (wheels for Linux x64/arm64, macOS x64/arm64 and Windows x64), anddocker run ghcr.io/sylphxai/anymdruns it from a multi-arch image.- Every release carries
SHA256SUMS, a CycloneDX SBOM and GitHub build-provenance attestations for its binaries and MCP Bundles; the install guide shows how to verify a download. - CI tests every change on Linux, macOS and Windows before it merges.
- The benchmarks page adds OmniDocBench v1.6 for scans and page images.
- The CLI prints one GitHub star line to stderr after the fifth successful interactive run, once ever. It is silent in MCP mode, in CI, and when stderr is not a terminal;
ANYMD_NO_STAR_HINT=1turns it off. cargo install anymdworks. Each release now also publishes the Rust crates to crates.io:anymd,anymd-core,anymd-formats,anymd-pdf, and two forks of upstream crates that carry our fixes,anymd-pdf-extract(frompdf-extract) andanymd-adobe-cmap-parser(fromadobe-cmap-parser). The forks keep the upstream MIT licence and credit. npm stays the primary install.sylphx-mcp-kitnow comes from crates.io (0.2.3) instead of a git tag.- The Claude Desktop MCP Bundles (
.mcpb) on each release now carry the anymd icon and a long description, and the install docs point at them. readand the CLI gainimages(--images refs|none). By default (refs) the raster images embedded in PDFs, DOCX, PPTX and EPUB files, which an agent cannot open from inside the container, are saved once each to<anymd cache>/images/<sha256>.<ext>and marked in the Markdown where they sit:and an<!-- image: WxH, page N -->comment. A PDF figure takes the caption line (Figure 3,Table 2,圖1) directly below or above it, otherwise the alt text, otherwiseimage. Images under 48 x 48 px, PDF images under 2% of the page, and pictures repeated on three or more pages, slides or chapters (logos, headers) are skipped. Images over 50 megapixels are refused, and nothing is written beside the source.inspectstructurealso lists a PDF's exported images (embeddedImages). The image cache is pruned at start, at most daily: files unused for 30 days go, and the folder stays under 2 GiB. Vector-drawn figures and charts are not exported yet.- Word tracked changes and comments come out as CriticMarkup: insertions
{++…++}, deletions{--…--}, a deletion next to an insertion as{~~old~>new~~}, and commented text as{==text==}{>>Author (date): comment<<}. Each tracked change is followed by its author and date,{++new++}{>>Author (date)<<}, the way CriticMarkup tracks several authors. Authors and dates are exactly as Word stores them. Inserted and deleted paragraph breaks, table rows and cells, text boxes, and footnotes are covered, and the output nests cleanly so simple CriticMarkup parsers read it. Contributed by @arthrod (#814). readand the CLI gainrevisions(--revisions markup|accept|reject) for Word files.markup(default) writes tracked changes and comments as CriticMarkup;acceptandrejectgive the text as Word shows it after Accept All or Reject All, without markup or comments. A document with no tracked changes or comments converts exactly as before under all three.
AgentDocBench corpus v1
Byte-for-byte mirror of the existing 38-document AgentDocBench corpus at source commit 5051332. Every document was checked against the existing corpus.json SHA-256 and byte count before upload. No benchmark inputs, IDs, ground truth, reference text, scores or scoring changed.
This is a corpus-only release, not an anymd software release. Original source URLs, license evidence and attribution are in corpus.json and CORPUS-NOTICE. Each document retains its original public-domain or permissive license; Microsoft test files include the MIT notice, CC BY/CC BY-SA files retain attribution, and Project Gutenberg EPUBs retain embedded notices (public domain in the USA).
SHA256SUMS covers all 38 documents plus corpus.json and CORPUS-NOTICE. SHA256SUMS itself has SHA-256 67cbec8ef5e5d1ffe570c33fe883481d38dacda01a4d12a97c87135bc1e217fe. These assets are fixed: any future mirror release uses a new corpus tag.
anymd 8.1.0
Minor Changes
- PDF tables are rebuilt from the page's drawn lines and aligned whitespace. A missing line between two cells makes a merged cell; wrapped cell text stays in its cell; stacked header lines become one header row, and a short header over several columns is repeated over each of them; rows banded between two lines split into one row per line; charts and boxes around whole passages are no longer mistaken for tables. Columns closer than a word gap apart are still split.
- Scanned pages are read at 300 dpi, several at a time, and the words tesseract finds are laid out like a text page: paragraphs join across line breaks, hyphenated words rejoin, and tables come out as tables. A paragraph-final comma, a common misreading of a typewritten full stop, becomes a full stop.
- Reading order handles three-column pages, and a table or figure inside one column of a two-column page no longer pulls in the other column.
Patch Changes
- Text a reader cannot see is left out: invisible text (rendering mode 3) and text painted in the colour of the box behind it.
- Text shown with the
'and"operators is no longer dropped, and pages whose text is turned (a landscape table on a portrait page) are laid out in their own direction. - Numbered paragraphs whose first line wraps are no longer read as headings.
- A spreadsheet sheet with a name of its own (not "Sheet1") gets its name as a heading above its table.
- The GitHub release carries an MCP Bundle (
.mcpb) for one-click install in Claude Desktop. crates/anymd-pdfis split into modules (extraction, rows, reading order, blocks, tables, OCR layout, rendering), each under 1,500 lines.
anymd 8.0.0
Major Changes
- The npm packages follow the mcp-kit layout.
@sylphx/anymdis now only a small launcher (bin/anymd.js) that runs the binary from the matching@sylphx/anymd-<platform>package; the platform packages ship the binary at their root. The@sylphx/citraand@sylphx/pdf-reader-mcpaliases run the same launcher. MCP client configs (npx -y @sylphx/anymd) need no change. - Removed the
@sylphx/anymd/sdkand@sylphx/anymd/pure-rustexports, theAnymd/CitraSDK classes, and theCITRA_RUST_BINvariable. Use theanymdCLI, or connect an MCP client toanymd mcp. The package no longer shipsexamples/orcorpus/. - The binary override is
ANYMD_BIN=/path/to/anymd;ANYMD_RUST_BINstill works.
Minor Changes
anymd setupadds anymd to the MCP clients on this machine (Claude Code, Codex, Cursor, VS Code, Claude Desktop, Windsurf, Gemini CLI).--dry-runpreviews the changes,--removeundoes them, and running it again changes nothing. Trynpx -y @sylphx/anymd setup.anymd versionprintsanymd X.Y.Z. The binary, the npm packages and the MCP Registry entry now carry one version (the Rust crates were at 3.1.1).
Patch Changes
- Releases publish through the shared mcp-kit release workflow: 5 native builds, npm trusted publishing, an
npxsmoke test, the GitHub release and the MCP Registry entry, whenevermaincarries a version that is not on npm yet. The Changesets flow and the release admission gate (release/admission.json, the capability matrix and its review specs) are retired.
v7.1.1
What's Changed
- fix(playground): show the anymd release version, not the crate version by @shtse8 in #760
- ci(release): publish to npm through trusted publishing (OIDC), not a long-lived token by @shtse8 in #761
- fix(http): only the exact health route skips the API key by @shtse8 in #762
- chore: remove committed junk, source-text tests, and the duplicate native CI build by @shtse8 in #764
- docs(copy): one source for descriptions and benchmark numbers by @shtse8 in #763
- refactor(core): split document_twin into focused modules by @shtse8 in #765
- ci: plain-language check on added lines by @shtse8 in #766
- ci(publish): npm trusted publishing only; drop the token fallback by @shtse8 in #768
- feat(bench): AgentDocBench, an open document-to-Markdown benchmark for agents by @shtse8 in #758
- refactor: crates renamed to anymd-*; pdf-reader-cli folded into tests; inline test modules moved out by @shtse8 in #769
- chore(release): version packages by @sylphx-release-bot[bot] in #767
- chore(release): admit the 7.1.1 candidate for publish by @shtse8 in #770
Full Changelog: v7.1.0...v7.1.1
v7.1.0
What's Changed
- docs: link repomap from the README by @shtse8 in #753
- refactor: PDF layout engine moves to its own crate (anymd-pdf, wasm-ready) by @shtse8 in #749
- chore: retire the v3.0.14 differential harness and TypeScript oracle by @shtse8 in #752
- feat(transcript): whisper model cache, verified opt-in download, OS install hints by @shtse8 in #750
- fix: monospace text keeps its line breaks (receipts, code, terminal output) by @shtse8 in #751
- docs: social preview image + real terminal demo GIF above the fold by @shtse8 in #754
- fix(pdf): structured JSON text keeps word spaces on TeX PDFs by @shtse8 in #756
- feat(docs): browser playground: any file → Markdown via WebAssembly by @shtse8 in #755
- chore(release): version packages by @sylphx-release-bot[bot] in #757
- chore(release): admit the 7.1.0 candidate for publish by @shtse8 in #759
Full Changelog: v7.0.0...v7.1.0