Repository navigation
Releases: DavidPandleton/webget
Releases · DavidPandleton/webget
Release list
v0.16.1
v0.16.0 - resumable crawler, structured extraction, search-provider seam
[0.16.0] - 2026-09-24
Added
- SQLite-backed resumable crawler with URL deduplication, same-domain policy, depth/page budgets, durable page persistence, and stale lease recovery.
webget crawl URL DBCLI command and MCPcrawltool with bounded validation and JSON output.- Dependency-light structured extraction primitives for JSON-LD and HTML tables, with explicit success, incomplete, and error states.
- Search-provider seam plus optional HTTP
SearxngSearchProvider; SearXNG is configured externally throughWEBGET_SEARXNG_URLand is not bundled.
Fixed
- RSS/Atom feed titles and descriptions no longer break generated Markdown when feed text contains newlines.
- Feed descriptions are truncated at word boundaries instead of cutting words in half.
- Packaging metadata now uses the modern SPDX license form without setuptools deprecation warnings.
- CI routes MCP-dependent crawl tests to the MCP job instead of collecting them in the no-MCP offline job.
Verification
- Full offline suite, Ruff, wheel/sdist build, clean Python 3.11 install, MCP stdio smoke, and GitHub Actions CI pass.
See CHANGELOG.md for the full entry.
v0.15.0 - browser discovery + webget doctor
[0.15.0] - 2026-09-21
Added
- Browser discovery for the crawl4ai pass (
webget/browser.py): look for a Chromium-family browser already installed before falling back to Playwright's bundled download. Resolution order:WEBGET_BROWSER_CDP,WEBGET_BROWSER_CHANNEL,WEBGET_BROWSER_PATH, detected browser, then bundled download. No install-time probing. webget doctorprints the resolved browser, every browser found, which are usable, crawl4ai/Playwright-cache state, and active env overrides.--jsonfor machine-readable output.
See CHANGELOG.md for the full entry.
v0.14.0 - learned engine health, time-budgeted failover, artifact CI gate
Added
- Learned engine health for failover ordering (
webget/health.py). Every search observation (which engine answered, how fast) folds into a per-install ledger at~/.local/state/webget/engine_health.json(override:WEBGET_ENGINE_HEALTH). When the requested engine fails, alternates are tried best-first by learned score instead of registry order. The ledger is advisory - it reorders candidates, never removes them - and it decays: entries older than 7 days score as unseen, because the blocking that made an engine dead is exactly what changes. Latency is penalized via EWMA (cap 8s, weight 0.25), so a reliable-but-glacial engine ranks below a fast one. There is no baked-in ranking in the source: the same install learns a different order on different networks. Diagnostics:health.health_snapshot(). - Time-budgeted failover (
FAILOVER_BUDGET_S = 15.0, envWEBGET_FAILOVER_BUDGET_S), replacing the fixed 3-attempt cap. A count is a bad proxy for "enough": 3 x 20s timeouts is a minute of dead air, while 8 fast engines finish in 3s and a live engine fifth in line never got reached. At leastMIN_FAILOVER_ATTEMPTS = 2alternates are always tried regardless of the clock, and the budget starts after the requested engine's own attempt, so one slow first call cannot silently disable the safety net. Provenance gainsbudget_exhaustedanduntriedso a budget-limited failure is distinguishable from "every engine is dead"; budget exhaustion is announced on stderr. - Provenance gains
health_ranked: True when the failover order came from the learned ledger rather than registry order.
Fixed
search_with_provenancecrashed with UnboundLocalError on every successful first-engine call (the_health_rankedvariable was defined after_provcould read it), and the failure was swallowed by failover, so every search silently burned a second engine call. Found by driving the installed artifact, not by the suite.- The CLI's user-facing help did not document
--engineor most options: the module docstring carrying them was overwritten by a shorter__doc__ = (...)assignment further down the file, so the option documentation written for 0.13.0 was invisible for the entire release line. The dead docstring is gone; the live one now documents all options.webget s --helppreviously ran a search for the word "help" instead of showing usage ---helpis now honoured after subcommands. - MCP tests pinned one MCP SDK spelling:
CallToolResult.isError(mcp 1.x / fastmcp 3.x) vsis_error(mcp 2.x / fastmcp 4.x). Atool_failed()helper in conftest reads whichever exists, and the suite now passes on both fastmcp 3.4.7 and 4.0.5.mcpextra floor raised tofastmcp>=4for installs (the code path never used the old name; only the tests did). - Test isolation: the search fixtures wrote real engine-health observations into the user's
~/.local/stateand read whatever the previous run left there, making outcomes depend on disk state outside the repo (the same suite passed and failed on an unchanged tree). The ledger path is now redirected per-test. Multi-engine labels likebrave,duckduckgoare recorded against each named engine instead of being stored as one dead ledger key.
Changed
- CI reshaped:
test(offline unit, 3 Python versions, plus a socket-blocked run proving no test touches the network),mcp-test(offline MCP surface),network(non-gating, runs thelive_network-marked tests and the stratified engine benchmark),artifact(installs the built wheel in a clean venv and runsscripts/artifact_smoke.pyagainst it - the check that catches the bugs above, which are invisible to source-tree tests), andbuild(sdist/wheel + twine check). Tests that genuinely reach the public internet carry alive_networkmarker (verified empirically by blocking sockets except localhost; exactly three tests failed without a network and they are the three marked). - New tooling:
scripts/bench_engines_stratified.py(102 queries across 8 strata - docs/code/news/commerce/science/local plus Tranco head/tail - with per-stratum success rates, top-1 accuracy where ground truth exists, and cross-engine agreement; the old 8-query benchmark could not support the reliability claims made from it),scripts/classify_tests.py(empirical network-dependence classification),scripts/find_live_tests.py(MCP live-test detection), andscripts/artifact_smoke.py(installed-artifact smoke test; non-gating online checks are separated from gating offline ones).
v0.12.1
v0.12.1
Token-waste fix: inline base64 image payloads in extracted markdown.
Fixed
- Oversized inline base64 image payloads (
data:image/...;base64,with a payload over 200 chars) are replaced withstrippedin HTTP fast-path markdown. The markdownify fallback used to keep them verbatim, so a single hero image could burn hundreds of KB of tokens as unreadable noise. Alt text and the mime prefix are preserved; short payloads are left untouched.
Install
pip install webget-cli
uv tool install "webget-cli[browser,mcp]"
v0.12.0
v0.12.0
Metadata + non-HTML content routing.
Added
- Metadata: fetch results now carry a
metadatadict (author,published_at,site_name,language) extracted via trafilatura on the HTTP fast path. Exposed inscrape_manyoutput, CLI--json, and MCPfetch/search_fetch. Values arenullwhen unknown or when the winning strategy was not HTTP. - Non-HTML routing: the HTTP fast path converts JSON (pretty code block), CSV (GFM table), RSS/Atom feeds (link list), PDFs (per-page text via new
pypdfcore dep), and plain text instead of erroringnot HTML. Valid non-HTML payloads skip the 100-char thin-check in the ladder.
Fixed
- MCP leak-scan test marker scoped to traceback frames, fixes a false positive on real page content.
Install
pip install webget-cli
uv tool install "webget-cli[browser,mcp]"
v0.11.0 - DoH DNS fallback, retry pass, and map discovery
Added
- SSRF guard: Dual-resolver DNS fallback via DoH (Cloudflare 1.1.1.1 / Google 8.8.8.8) when system getaddrinfo fails.
- Ladder & CLI: Opt-in transient retry pass (--retry / -r) for HTTP fast path timeout recovery.
- URL Discovery: 'webget map ' command and MCP 'map' tool for sitemap/robots.txt discovery.
v0.10.0 - package refactor + MCP test fix
What's changed
- Internal: package refactor. Split
webget_cli.py(1762 lines) into awebget/package with focused sub-modules (cache,ssrf,profile,http,firecrawl, \搜索,ladder,cli).webget_cli.pyis kept as a ~50-line backwards-compatible shim soimport webget_cli as webgetand the full test fixture keep working unchanged. Entry point migrated towebget.cli:main`. No behavior change. - MCP tests: rename
.is_error->.isErrorfor fastmcp 3.4.7 (camelCase CallToolResult field). 19 mechanical replacements across 5 test files.
Verification
- Test suite: 344 passed, 0 failed (was 21 pre-existing MCP failures)
- \ruff check .: all checks passed
— Rouge
v0.9.0 - bulk-crawl hardening
- Absolute wall-clock per_url_timeout (slow-drip proof)
- DNS-bounded concurrent SSRF pre-check (bulk groups 90s to 20-36s)
- Cache eviction cap 5000 (WEBGET_CACHE_MAX) + lazy scan
- Strategy memory 14-day TTL + legacy format migration
- Error diagnostics: empty-string exceptions fall back to type name
- reasons: [] always present (consistent shape)
- SSRF route guard handles BrowserContext (profile path)
- MCP login auto-headless on servers without display
- Test suite: 344 passed, 0 failed (was 3 pre-existing failures) - Rouge
webget 0.8.0 - MCP authenticated sessions
Added
- MCP server:
list_profilestool (non-sensitive session metadata) andprofileparameter onfetch/search_fetchfor authenticated scraping with locally stored sessions (webget login). Cookie values are never exposed; invalid or unknown profile names return clean errors instead of a silent anonymous fallback. WEBGET_PROFILE_DIRenv override for the profile root (tests/ops).
Backward compatible: profile omitted behaves exactly like 0.7.3. Full suite + browser integration + MCP E2E (cookie-gated auth) verified.