Skip to content

v0.12.0

Choose a tag to compare

@DavidPandleton DavidPandleton released this 15 Sep 17:07
· 57 commits to main since this release

v0.12.0

Metadata + non-HTML content routing.

Added

  • Metadata: fetch results now carry a metadata dict (author, published_at, site_name, language) extracted via trafilatura on the HTTP fast path. Exposed in scrape_many output, CLI --json, and MCP fetch/search_fetch. Values are null when unknown or when the winning strategy was not HTTP.
  • Non-HTML routing: the HTTP fast path converts JSON (pretty code block), CSV (GFM table), RSS/Atom feeds (link list), PDFs (per-page text via new pypdf core dep), and plain text instead of erroring not HTML. Valid non-HTML payloads skip the 100-char thin-check in the ladder.

Fixed

  • MCP leak-scan test marker scoped to traceback frames, fixes a false positive on real page content.

Install

pip install webget-cli
uv tool install "webget-cli[browser,mcp]"