Skip to content

v1.8.0

Latest

Choose a tag to compare

@stwhwing stwhwing released this 21 Sep 09:27

markitdown-skill v1.8.0

Security hardening + provenance for the web→Markdown pipeline (public build, no telemetry).

Security

  • Response / decompressed size caps (P0): raw responses > 32 MiB rejected; gzip/deflate/br decoded with a streaming, bounded reader that aborts past 64 MiB (defeats decompression bombs).
  • Redirect-hop limit (P0): 3xx chains longer than 10 hops refused.
  • Reject embedded credentials (P0): user:pass@host URLs refused outright.
  • Prompt-injection boundary (P1): url_to_markdown.py --sanitize strips <script>/<style>, neutralises javascript:/data: URIs, and wraps output in --- EXTERNAL CONTENT --- markers (content stays data, not instructions).
  • New references/SECURITY.md documents the full threat model (SSRF, resource caps, prompt-injection boundary, browser sandbox, data flow).

Provenance & reliability

  • Per-conversion manifest (P1): --manifest <file> records source, output, UTC time, sha256, byte size, and a heuristic quality score for url_to_markdown.py and batch_convert.py — useful for knowledge-base ingestion trails.
  • Atomic output writes (P1): output files are written via temp file + os.replace (no half-written files on crash).
  • Empty / scanned-source hint (P1): low-quality or image-only results print an Azure Document Intelligence / local-OCR upgrade suggestion.
  • Honest positioning (P2): SKILL.md now states plainly when to use url_to_markdown.py vs a bare markitdown <url>, and that the public build ships no reporting component.
  • uvx install fallback (P2): uvx --with 'markitdown[all]' markitdown ... for environments without pip write access.

Published to GitHub + skillhub.cn + ClawHub at the same version.