Skip to content

Releases: anishfyi/curl_reap

v1.1.0

Choose a tag to compare

@anishfyi anishfyi released this 29 Sep 20:15
7685bf2

curl_reap 1.1.0 fixes the reap command line tool and adds HTTP/2, a cookie jar, streaming downloads, and versioned impersonation targets.

What changed

  • Fixed: the reap command failed at startup in 1.0.0 with a SyntaxError. It works again, and CI now runs reap --help on every Python version.
  • Added: optional HTTP/2 transport with Session(http2=True) and reap --http2 (install the h2 extra). It falls back to HTTP/1.1 when HTTP/2 is not available. Response bodies over HTTP/2 are not content-decoded yet.
  • Added: impersonate= as an alias of profile=, with 38 versioned Chrome, Edge, Firefox, and Safari targets.
  • Added: Session.cookies, a cookie jar that stores cookies from responses and redirects and replays them.
  • Added: streaming downloads with Session.download(), reap.download(), and AsyncSession.download().
  • Added: reap.render() and reap.render_if_empty() through Playwright (install the js extra).
  • Tests: 67 pass on Python 3.9 to 3.13.

Install

PyPI still serves 0.2.2, which predates the 1.0 rewrite. Until 1.1.0 is uploaded there, install it from this release:

pip install "curl_reap @ https://github.com/anishfyi/curl_reap/releases/download/v1.1.0/curl_reap-1.1.0-py3-none-any.whl"

Or from the tag:

pip install "git+https://github.com/anishfyi/curl_reap@v1.1.0"

Documentation

https://velofy.co/curl_reap/

Full changelog: https://velofy.co/curl_reap/changelog/

v1.0.0: the great rewrite

Choose a tag to compare

@anishfyi anishfyi released this 22 Aug 00:26
d108f7f

The breaking, from-scratch release

Transport, rewritten

  • curl_cffi is gone. No cffi bindings, no binary wheels, no runtime downloads: a hand-rolled HTTP/1.1 engine on the Python standard library (socket + ssl) with byte-level control of request-line formatting and header order/casing, keep-alive connection pooling, chunked transfer decoding, and gzip/deflate decoding.
  • TLS/header profiles (curl_reap.tls.Profile): curated cipher orderings, ALPN, and browser-family default headers for chrome, firefox, safari. Bring your own Profile too.
  • Breaking: impersonate= is removed. Use profile= everywhere (reap.get(url, profile="chrome")).
  • Retries with exponential backoff + jitter, proxy rotation, and profile rotation across retries.

Everything else you know and love

Self-healing selectors, structured extraction (jsonld/meta/links/images/tables/markdown), concurrent crawl engine with AutoThrottle, robots.txt Crawl-delay, sitemap discovery, disk cache with 304 revalidation, async facade, pipelines, CLI.

Packaging

Dependencies are now just lxml and cssselect. Python 3.9+.

Docs

The website is rebuilt mobile-first: hamburger nav under 768px, dark mode, zero external requests: https://anishfyi.github.io/curl_reap/

Honesty section

Standard-library TLS cannot produce an exact browser ClientHello. Profiles are best effort, not JA3 parity; hard-blocked enterprise anti-bot edges may still tell it apart. curl_reap does not solve CAPTCHAs or defeat anti-bot services.

Full changelog: ad5cd5b...d108f7f

v0.2.0 — resilient transport, structured extraction, continuous crawl engine, CLI

Choose a tag to compare

@anishfyi anishfyi released this 08 Jul 23:39

Major upgrade turning curl_reap into a full-stack scraper. The v0.1 API is unchanged and still works.

Transport

  • RetryPolicy: exponential backoff + jitter, honors Retry-After, auto-retries 429/5xx
  • Fingerprint rotation (rotate="random"|"sequence") and proxy rotation (proxy=[...])
  • DiskCache: content-addressed cache so repeat GETs are instant (resp.from_cache)
  • AsyncSession + aget/apost for asyncio.gather fan-out
  • Richer Response: .status_code alias, .cookies, .urljoin(), .follow(), .raise_for_status()

Parsing — one-call structured extraction

jsonld(), meta_tags(), links(), images(), tables(), markdown(), plus re_first() and a find_by_text that returns the deepest match.

Crawl engine

  • Continuous priority-queue scheduler (no more batch-wave stalls)
  • Per-domain AutoThrottle that backs off on 429/503
  • Opt-in robots.txt, allowed_domains / max_depth confinement, errbacks, richer stats/logging

Spider & pipelines

  • Request(priority=, errback=, dont_filter=), Response.follow, new SitemapSpider
  • SqlitePipeline (auto columns + upsert), thread-safe JsonLinesPipeline

CLI

New reap command: reap get|meta|links|crawl (installed via entry point).

Tests

24 passing, including new extraction, engine (offline), and cache/pipeline suites.