Releases: anishfyi/curl_reap
Release list
v1.1.0
curl_reap 1.1.0 fixes the reap command line tool and adds HTTP/2, a cookie jar, streaming downloads, and versioned impersonation targets.
What changed
- Fixed: the
reapcommand failed at startup in 1.0.0 with aSyntaxError. It works again, and CI now runsreap --helpon every Python version. - Added: optional HTTP/2 transport with
Session(http2=True)andreap --http2(install theh2extra). It falls back to HTTP/1.1 when HTTP/2 is not available. Response bodies over HTTP/2 are not content-decoded yet. - Added:
impersonate=as an alias ofprofile=, with 38 versioned Chrome, Edge, Firefox, and Safari targets. - Added:
Session.cookies, a cookie jar that stores cookies from responses and redirects and replays them. - Added: streaming downloads with
Session.download(),reap.download(), andAsyncSession.download(). - Added:
reap.render()andreap.render_if_empty()through Playwright (install thejsextra). - Tests: 67 pass on Python 3.9 to 3.13.
Install
PyPI still serves 0.2.2, which predates the 1.0 rewrite. Until 1.1.0 is uploaded there, install it from this release:
pip install "curl_reap @ https://github.com/anishfyi/curl_reap/releases/download/v1.1.0/curl_reap-1.1.0-py3-none-any.whl"
Or from the tag:
pip install "git+https://github.com/anishfyi/curl_reap@v1.1.0"
Documentation
Full changelog: https://velofy.co/curl_reap/changelog/
v1.0.0: the great rewrite
The breaking, from-scratch release
Transport, rewritten
- curl_cffi is gone. No cffi bindings, no binary wheels, no runtime downloads: a hand-rolled HTTP/1.1 engine on the Python standard library (
socket+ssl) with byte-level control of request-line formatting and header order/casing, keep-alive connection pooling, chunked transfer decoding, and gzip/deflate decoding. - TLS/header profiles (
curl_reap.tls.Profile): curated cipher orderings, ALPN, and browser-family default headers forchrome,firefox,safari. Bring your own Profile too. - Breaking:
impersonate=is removed. Useprofile=everywhere (reap.get(url, profile="chrome")). - Retries with exponential backoff + jitter, proxy rotation, and profile rotation across retries.
Everything else you know and love
Self-healing selectors, structured extraction (jsonld/meta/links/images/tables/markdown), concurrent crawl engine with AutoThrottle, robots.txt Crawl-delay, sitemap discovery, disk cache with 304 revalidation, async facade, pipelines, CLI.
Packaging
Dependencies are now just lxml and cssselect. Python 3.9+.
Docs
The website is rebuilt mobile-first: hamburger nav under 768px, dark mode, zero external requests: https://anishfyi.github.io/curl_reap/
Honesty section
Standard-library TLS cannot produce an exact browser ClientHello. Profiles are best effort, not JA3 parity; hard-blocked enterprise anti-bot edges may still tell it apart. curl_reap does not solve CAPTCHAs or defeat anti-bot services.
Full changelog: ad5cd5b...d108f7f
v0.2.0 — resilient transport, structured extraction, continuous crawl engine, CLI
Major upgrade turning curl_reap into a full-stack scraper. The v0.1 API is unchanged and still works.
Transport
RetryPolicy: exponential backoff + jitter, honorsRetry-After, auto-retries 429/5xx- Fingerprint rotation (
rotate="random"|"sequence") and proxy rotation (proxy=[...]) DiskCache: content-addressed cache so repeat GETs are instant (resp.from_cache)AsyncSession+aget/apostforasyncio.gatherfan-out- Richer
Response:.status_codealias,.cookies,.urljoin(),.follow(),.raise_for_status()
Parsing — one-call structured extraction
jsonld(), meta_tags(), links(), images(), tables(), markdown(), plus re_first() and a find_by_text that returns the deepest match.
Crawl engine
- Continuous priority-queue scheduler (no more batch-wave stalls)
- Per-domain AutoThrottle that backs off on 429/503
- Opt-in
robots.txt,allowed_domains/max_depthconfinement, errbacks, richer stats/logging
Spider & pipelines
Request(priority=, errback=, dont_filter=),Response.follow, newSitemapSpiderSqlitePipeline(auto columns + upsert), thread-safeJsonLinesPipeline
CLI
New reap command: reap get|meta|links|crawl (installed via entry point).
Tests
24 passing, including new extraction, engine (offline), and cache/pipeline suites.