Skip to content

Releases: DavidPandleton/webget

v0.16.1

Choose a tag to compare

@DavidPandleton DavidPandleton released this 25 Sep 12:27

Fix HttpOnly cookie loading and epoch-0 cookie expiry on the HTTP fast path. See CHANGELOG.md.

v0.16.0 - resumable crawler, structured extraction, search-provider seam

Choose a tag to compare

@DavidPandleton DavidPandleton released this 24 Sep 16:04

[0.16.0] - 2026-09-24

Added

  • SQLite-backed resumable crawler with URL deduplication, same-domain policy, depth/page budgets, durable page persistence, and stale lease recovery.
  • webget crawl URL DB CLI command and MCP crawl tool with bounded validation and JSON output.
  • Dependency-light structured extraction primitives for JSON-LD and HTML tables, with explicit success, incomplete, and error states.
  • Search-provider seam plus optional HTTP SearxngSearchProvider; SearXNG is configured externally through WEBGET_SEARXNG_URL and is not bundled.

Fixed

  • RSS/Atom feed titles and descriptions no longer break generated Markdown when feed text contains newlines.
  • Feed descriptions are truncated at word boundaries instead of cutting words in half.
  • Packaging metadata now uses the modern SPDX license form without setuptools deprecation warnings.
  • CI routes MCP-dependent crawl tests to the MCP job instead of collecting them in the no-MCP offline job.

Verification

  • Full offline suite, Ruff, wheel/sdist build, clean Python 3.11 install, MCP stdio smoke, and GitHub Actions CI pass.

See CHANGELOG.md for the full entry.

v0.15.0 - browser discovery + webget doctor

Choose a tag to compare

@DavidPandleton DavidPandleton released this 24 Sep 16:04

[0.15.0] - 2026-09-21

Added

  • Browser discovery for the crawl4ai pass (webget/browser.py): look for a Chromium-family browser already installed before falling back to Playwright's bundled download. Resolution order: WEBGET_BROWSER_CDP, WEBGET_BROWSER_CHANNEL, WEBGET_BROWSER_PATH, detected browser, then bundled download. No install-time probing.
  • webget doctor prints the resolved browser, every browser found, which are usable, crawl4ai/Playwright-cache state, and active env overrides. --json for machine-readable output.

See CHANGELOG.md for the full entry.

v0.14.0 - learned engine health, time-budgeted failover, artifact CI gate

Choose a tag to compare

@DavidPandleton DavidPandleton released this 18 Sep 20:45

Added

  • Learned engine health for failover ordering (webget/health.py). Every search observation (which engine answered, how fast) folds into a per-install ledger at ~/.local/state/webget/engine_health.json (override: WEBGET_ENGINE_HEALTH). When the requested engine fails, alternates are tried best-first by learned score instead of registry order. The ledger is advisory - it reorders candidates, never removes them - and it decays: entries older than 7 days score as unseen, because the blocking that made an engine dead is exactly what changes. Latency is penalized via EWMA (cap 8s, weight 0.25), so a reliable-but-glacial engine ranks below a fast one. There is no baked-in ranking in the source: the same install learns a different order on different networks. Diagnostics: health.health_snapshot().
  • Time-budgeted failover (FAILOVER_BUDGET_S = 15.0, env WEBGET_FAILOVER_BUDGET_S), replacing the fixed 3-attempt cap. A count is a bad proxy for "enough": 3 x 20s timeouts is a minute of dead air, while 8 fast engines finish in 3s and a live engine fifth in line never got reached. At least MIN_FAILOVER_ATTEMPTS = 2 alternates are always tried regardless of the clock, and the budget starts after the requested engine's own attempt, so one slow first call cannot silently disable the safety net. Provenance gains budget_exhausted and untried so a budget-limited failure is distinguishable from "every engine is dead"; budget exhaustion is announced on stderr.
  • Provenance gains health_ranked: True when the failover order came from the learned ledger rather than registry order.

Fixed

  • search_with_provenance crashed with UnboundLocalError on every successful first-engine call (the _health_ranked variable was defined after _prov could read it), and the failure was swallowed by failover, so every search silently burned a second engine call. Found by driving the installed artifact, not by the suite.
  • The CLI's user-facing help did not document --engine or most options: the module docstring carrying them was overwritten by a shorter __doc__ = (...) assignment further down the file, so the option documentation written for 0.13.0 was invisible for the entire release line. The dead docstring is gone; the live one now documents all options. webget s --help previously ran a search for the word "help" instead of showing usage - --help is now honoured after subcommands.
  • MCP tests pinned one MCP SDK spelling: CallToolResult.isError (mcp 1.x / fastmcp 3.x) vs is_error (mcp 2.x / fastmcp 4.x). A tool_failed() helper in conftest reads whichever exists, and the suite now passes on both fastmcp 3.4.7 and 4.0.5. mcp extra floor raised to fastmcp>=4 for installs (the code path never used the old name; only the tests did).
  • Test isolation: the search fixtures wrote real engine-health observations into the user's ~/.local/state and read whatever the previous run left there, making outcomes depend on disk state outside the repo (the same suite passed and failed on an unchanged tree). The ledger path is now redirected per-test. Multi-engine labels like brave,duckduckgo are recorded against each named engine instead of being stored as one dead ledger key.

Changed

  • CI reshaped: test (offline unit, 3 Python versions, plus a socket-blocked run proving no test touches the network), mcp-test (offline MCP surface), network (non-gating, runs the live_network-marked tests and the stratified engine benchmark), artifact (installs the built wheel in a clean venv and runs scripts/artifact_smoke.py against it - the check that catches the bugs above, which are invisible to source-tree tests), and build (sdist/wheel + twine check). Tests that genuinely reach the public internet carry a live_network marker (verified empirically by blocking sockets except localhost; exactly three tests failed without a network and they are the three marked).
  • New tooling: scripts/bench_engines_stratified.py (102 queries across 8 strata - docs/code/news/commerce/science/local plus Tranco head/tail - with per-stratum success rates, top-1 accuracy where ground truth exists, and cross-engine agreement; the old 8-query benchmark could not support the reliability claims made from it), scripts/classify_tests.py (empirical network-dependence classification), scripts/find_live_tests.py (MCP live-test detection), and scripts/artifact_smoke.py (installed-artifact smoke test; non-gating online checks are separated from gating offline ones).

v0.12.1

Choose a tag to compare

@DavidPandleton DavidPandleton released this 15 Sep 18:31

v0.12.1

Token-waste fix: inline base64 image payloads in extracted markdown.

Fixed

  • Oversized inline base64 image payloads (data:image/...;base64, with a payload over 200 chars) are replaced with stripped in HTTP fast-path markdown. The markdownify fallback used to keep them verbatim, so a single hero image could burn hundreds of KB of tokens as unreadable noise. Alt text and the mime prefix are preserved; short payloads are left untouched.

Install

pip install webget-cli
uv tool install "webget-cli[browser,mcp]"

v0.12.0

Choose a tag to compare

@DavidPandleton DavidPandleton released this 15 Sep 17:07

v0.12.0

Metadata + non-HTML content routing.

Added

  • Metadata: fetch results now carry a metadata dict (author, published_at, site_name, language) extracted via trafilatura on the HTTP fast path. Exposed in scrape_many output, CLI --json, and MCP fetch/search_fetch. Values are null when unknown or when the winning strategy was not HTTP.
  • Non-HTML routing: the HTTP fast path converts JSON (pretty code block), CSV (GFM table), RSS/Atom feeds (link list), PDFs (per-page text via new pypdf core dep), and plain text instead of erroring not HTML. Valid non-HTML payloads skip the 100-char thin-check in the ladder.

Fixed

  • MCP leak-scan test marker scoped to traceback frames, fixes a false positive on real page content.

Install

pip install webget-cli
uv tool install "webget-cli[browser,mcp]"

v0.11.0 - DoH DNS fallback, retry pass, and map discovery

Choose a tag to compare

@DavidPandleton DavidPandleton released this 03 Sep 08:28

Added

  • SSRF guard: Dual-resolver DNS fallback via DoH (Cloudflare 1.1.1.1 / Google 8.8.8.8) when system getaddrinfo fails.
  • Ladder & CLI: Opt-in transient retry pass (--retry / -r) for HTTP fast path timeout recovery.
  • URL Discovery: 'webget map ' command and MCP 'map' tool for sitemap/robots.txt discovery.

v0.10.0 - package refactor + MCP test fix

Choose a tag to compare

@DavidPandleton DavidPandleton released this 01 Sep 06:26

What's changed

  • Internal: package refactor. Split webget_cli.py (1762 lines) into a webget/ package with focused sub-modules (cache, ssrf, profile, http, firecrawl, \搜索, ladder, cli). webget_cli.pyis kept as a ~50-line backwards-compatible shim soimport webget_cli as webgetand the full test fixture keep working unchanged. Entry point migrated towebget.cli:main`. No behavior change.
  • MCP tests: rename .is_error -> .isError for fastmcp 3.4.7 (camelCase CallToolResult field). 19 mechanical replacements across 5 test files.

Verification

  • Test suite: 344 passed, 0 failed (was 21 pre-existing MCP failures)
  • \ruff check .: all checks passed

— Rouge

v0.9.0 - bulk-crawl hardening

Choose a tag to compare

@DavidPandleton DavidPandleton released this 31 Aug 19:41
  • Absolute wall-clock per_url_timeout (slow-drip proof)
  • DNS-bounded concurrent SSRF pre-check (bulk groups 90s to 20-36s)
  • Cache eviction cap 5000 (WEBGET_CACHE_MAX) + lazy scan
  • Strategy memory 14-day TTL + legacy format migration
  • Error diagnostics: empty-string exceptions fall back to type name
  • reasons: [] always present (consistent shape)
  • SSRF route guard handles BrowserContext (profile path)
  • MCP login auto-headless on servers without display
  • Test suite: 344 passed, 0 failed (was 3 pre-existing failures) - Rouge

webget 0.8.0 - MCP authenticated sessions

Choose a tag to compare

@DavidPandleton DavidPandleton released this 08 Aug 12:50

Added

  • MCP server: list_profiles tool (non-sensitive session metadata) and profile parameter on fetch / search_fetch for authenticated scraping with locally stored sessions (webget login). Cookie values are never exposed; invalid or unknown profile names return clean errors instead of a silent anonymous fallback.
  • WEBGET_PROFILE_DIR env override for the profile root (tests/ops).

Backward compatible: profile omitted behaves exactly like 0.7.3. Full suite + browser integration + MCP E2E (cookie-gated auth) verified.