markitdown-skill v1.8.0
Security hardening + provenance for the web→Markdown pipeline (public build, no telemetry).
Security
- Response / decompressed size caps (P0): raw responses > 32 MiB rejected; gzip/deflate/br decoded with a streaming, bounded reader that aborts past 64 MiB (defeats decompression bombs).
- Redirect-hop limit (P0): 3xx chains longer than 10 hops refused.
- Reject embedded credentials (P0):
user:pass@hostURLs refused outright. - Prompt-injection boundary (P1):
url_to_markdown.py --sanitizestrips<script>/<style>, neutralisesjavascript:/data:URIs, and wraps output in--- EXTERNAL CONTENT ---markers (content stays data, not instructions). - New
references/SECURITY.mddocuments the full threat model (SSRF, resource caps, prompt-injection boundary, browser sandbox, data flow).
Provenance & reliability
- Per-conversion manifest (P1):
--manifest <file>records source, output, UTC time, sha256, byte size, and a heuristic quality score forurl_to_markdown.pyandbatch_convert.py— useful for knowledge-base ingestion trails. - Atomic output writes (P1): output files are written via temp file +
os.replace(no half-written files on crash). - Empty / scanned-source hint (P1): low-quality or image-only results print an Azure Document Intelligence / local-OCR upgrade suggestion.
- Honest positioning (P2): SKILL.md now states plainly when to use
url_to_markdown.pyvs a baremarkitdown <url>, and that the public build ships no reporting component. uvxinstall fallback (P2):uvx --with 'markitdown[all]' markitdown ...for environments without pip write access.
Published to GitHub + skillhub.cn + ClawHub at the same version.