v1.0.0
A complete rewrite. There are no compatibility shims: almost every import changes.
docs/migration.md is the mapping, and
docs/model.md is why.
The change
The library is now organised around a model of what it is up against, rather than
around a request pipeline with anti-detection features bolted on. A mitigation engine
folds many detectors into one score, and admission is close to a conjunction — so the
weakest layer bounds the outcome, and effort spent on any other layer buys nothing.
Detectors that read an artifact the client emits are reproducible; detectors that
read a property it must possess are not.
Every behaviour below follows from those two statements.
Breaking
ScraperEngineis gone, andScraperis no longer arequests.Session
subclass. The transport is a two-method seam, so the escalation ladder can move
between transports and every tier is testable without a network.- The in-process Cloudflare solvers are gone (v1, v2, v3, Turnstile), along with
theexejsdependency. They cannot keep up with the challenge format, and the layer
they targeted is only reachable by a real browser. A challenged site now needs
ScraperConfig.browser; without one it raises and says so, instead of attempting a
solve that usually failed. - TLS cipher rotation is gone. Reordering the cipher list per request does not
produce a browser fingerprint, it produces an unstable one — and an unstable TLS
fingerprint invalidates any clearance bound to it. The feature was breaking the layer
above it. - Header randomisation is gone. Header order is read, not just header values. An
impersonation profile emits a complete, correctly ordered set;
scraper.identity.OVERRIDABLEnow caps what may be written over it. - The User-Agent is taken from the transport, not imposed on it. The generated-UA
machinery is gone. A profile supplies the User-Agent until a real browser earns a
clearance, at which point the browser is the source of truth and its exact string is
reproduced — because that is what the clearance is bound to. - Impersonation is a core dependency, not the
impersonateextra. An ordinary
Python client fails layers 2–5 in the first round trip, so a build without it is not
a degraded scraper but one that cannot reach a protected page. default_config(),StealthConfig,BrowserConfig,ProxyConfig,ProxyUrl,
TorProxyUrl,apply_browser_clearance()and thescraper.enginepackage are
removed.SharedLimiterbecomesSharedState.AbortedExceptionbecomesAborted,
and theCloudflareExceptionhierarchy becomesBlocked/Impassable/
Exhausted, each carrying the layer it is attributed to — orNone, when the
failure is ours rather than the site's. Code that dereferencesexc.layeror
exc.layer_infomust handle that; a type checker will point at every such place.- New extras:
browser(nodriver) andbotauth(cryptography).impersonateis gone. ScraperConfig.impersonatedefaults to"", which resolves tofirefox— or to
chromewhen a browser solver is configured, because the bundled solver drives
Chrome and a clearance is bound to a User-Agent and a TLS fingerprint together.
Read it through the newScraperConfig.profile(); the field is no longer the answer
on its own.
Measured against 0.2.6
The rewrite was A/B'd against 0.2.6 on 150 hosts from lightnovel-crawler's source
index, three arms per host, curl_cffi pinned to the same version in both, one
classifier for every arm, arm order shuffled per host. livetest/compare.py runs it.
The first run said the rewrite had made per-request retrieval slightly worse: 0.2.6
took 56 of 150 hosts with impersonation off, 55 with it on, and 1.0 took 51 — losing
five hosts head to head and winning one. Everything in this section is what closed
that gap, and each item is here because it was measured, not reasoned about:
- A first request now carries a
Refererand matchingSec-Fetch-Site. No
browser does this — a typed address has no referrer, which is exactly why the
rewrite sent none. Over 85 hosts that refuse an impersonated client it recovered
three and cost zero, turning a 403-with-a-challenge into a full page. The coherent
header set matters more than the header: a synthesised referrer alongside
Sec-Fetch-Site: noneis a contradiction, and one host only yielded once both
agreed. - A JavaScript-only redirect is followed instead of being mistaken for content.
A family of bot checks answers with a few hundred bytes of
window.location.replace('…?token=…')— a200, no challenge marker, no Cloudflare
header. Both releases handed that back as a successful retrieval, so a scraper
reported success and collected an empty document. It was 19% of the corpus. No
browser is needed because the destination is emitted in the HTML; running the
script would produce the URL already sitting there in plain text. New
Action.FOLLOW,Diagnosis.locationanddiagnosis.js_redirect(). - The default impersonation profile is Firefox. Over a random 150-host sample:
firefox 85, safari 84, edge 82, chrome 81 — and against chrome, firefox won four
hosts and lost none. Chrome being the most common browser is a reason to expect it
to be unremarkable, not evidence that it is the least remarkable. - Pool sessions are released when a scraper closes.
ExitPool.release()existed,
was never called, and only dropped the local lease. Every lease minted a fresh key
that the pool then held untilSESSION_TTL, so a process building several scrapers
in a row walked the pool out of capacity — and the symptom was the misleading part:
the next lease could not connect, a transport failure through a proxy is evidence
about the exit, and the model reported a reputation block on a destination that
never saw the request. Needs tor-pool 0.2.1, which puts
DELETE /api/sessions/{key}on theproxyscope.
After those, on the same corpus with the archive tier off — so nothing is borrowed
that 0.2.6 cannot do — 1.0 takes 82 of 150 against 0.2.6's best of 76, and 30.8% of the
hard hosts against 25.6%. Head to head it wins nine and loses none (sign test
p = 0.004). Hosts serving an unrecognised stub fell from 28 to 6. With the archive tier
on, total reach was 73% against 39%.
Added
scraper.layers— the model as code: nineteen layers, what each reads
(Trait), what this library does about it (Stance), the bound (weakest) and the
arithmetic that shows why fixing the wrong layer gains nothing (marginal_gain).
Layers 2–5 are declared as one barrier, andexpand()keeps any reach set closed
over the group.scraper.diagnosis— a response becomes a binding layer plus an action, as a pure
function over primitives. Three readings that a status-code table gets wrong: a200
carrying a challenge is a failure, a429is a pacing problem rather than a spent
address, and a403with error 1010 is about the automation channel rather than the
address. A407is reported as our own proxy credential, not as the site needing a
login.scraper.planner— chooses the cheapest capability whose reach covers the binding
layer. Three rules that contradict the conventional table: a possessed property is
never rotated away from; rotation requires somewhere better to go, so a pool of
published ranges stops with an explanation instead of cycling; and escalation only
goes to a tier that actually reaches the layer. Repeated failure at an
already-covered emitted layer is re-attributed to the per-zone composite, because
recurrence is the only evidence available from outside.scraper.identity— the emitted signals as one indivisible thing.
Clearance.usable_by()refuses to replay a clearance under a different identity,
which makes the classic rotating-proxy failure structurally impossible rather than
merely documented.scraper.exits— addresses described by kind, andExitKind.reachdeciding what
layer 1 can be told. Leased per origin and held; rotation happens on evidence, never
on a timer. tor-pool support is retained, and a failure report now carries the kind
derived from the binding layer.scraper.pacing— inter-request gaps drawn from a gamma distribution rather than set
to a constant, occasional reading pauses, homepage warm-up, and a real referrer chain
with fetch metadata. Throttles widen a learned per-origin interval that persists.scraper.memory— per-origin state that survives the process: the binding layer, the
working tier, a clearance and the identity it belongs to, the learned interval,
observed JSON endpoints, and recorded decoy URLs. On by default, because the layer it
exists for cannot be satisfied by a process that forgets.scraper.state.SharedState— shares the address, identity, history, pacing, referrer
chain and decoy list between scrapers pointed at one site. Two scrapers with separate
state do not look like one visitor going faster; they look like two who contradict
each other.scraper.tiers—archive(Wayback, serving the original URL so links resolve
against the real site),direct(the baseline),clearance(solve once, reuse many,
delegating every request todirectso the solve and the fetch cannot diverge), and
managed(a provider callable; none bundled, since a wrapper that guesses a vendor
format wrong fails in a way that looks like the site blocking you).scraper.browser— a two-methodBrowserSolverprotocol, anodriveradapter, and
CallableSolverfor anything else. Headed and WebRTC-disabled by default, both
deliberately: a headless build reports a software renderer, and a STUN request reports
the host's real address past the proxy without any request failing. One browser
profile directory per address.scraper.links—safe_linksenumerates only anchors a person could click, and
TopicGuardnotices content that stopped being about the site. This is the only
defence against the one layer that returns no error, so the guard runs on the way out
rather than on demand.scraper.botauth— RFC 9421 Ed25519 request signing with theweb-bot-authtag,
plus the key directory document to publish. The one layer with no bypass, and for a
crawler willing to identify itself, the cheapest tier in the stack.Scraper.explain(url)andScraper.knows(url)— what the library concluded is
binding, which tier settled, how fast it has learned it can go, and what it has
available.livetest/— a live verification harness that exercises every path against real
Cloudflare deployments, using every host in lightnovel-crawler's source index as
the corpus. Separate fromtests/, which stays offline. See
livetest/README.md; the current run is
livetest/report.html.
Fixed before release, found by live traffic
Most of what was fixed before release was found this way, and none of it was visible to
a stubbed transport. These are not regressions from 0.2.x — they are defects in the new
code, and two of them made a whole feature silently useless while every unit test
passed. Each has a regression test whose docstring says it was found live.
- Cloudflare's injected JavaScript-Detections script was read as a challenge.
The script is served from/cdn-cgi/challenge-platform/scripts/jsd/…on ordinary
successful pages, and that path was a challenge marker — so content pages were
diagnosed as interstitials. Measured across two live populations before changing
it: the bare prefix appeared on 9 of 10 normally-served pages, the challenge-only
/h/orchestrate sub-path on none. 18 of 22 hosts reported as challenged were
serving content fine. The same marker also stopped the browser solve loop from
ever detecting "cleared", so every solve burned its full timeout. - The archive tier could never find a capture. A negative Wayback CDX
limit
is documented as "the last N rows" but returns an empty body once a filter is
applied. The query is now bounded server-side and the newest rows taken from the
tail. Separately, an unbounded query timed out on popular URLs, and a
rate-limited index was reported as "nothing archived" — all three produced the
same misleading message, so a lookup failure, an empty index and an age limit now
say which they are, and the index is retried once. - Real navigation was being dropped as decoy content. An anchor containing only
an icon-font element counted as "nothing rendered", and a URL took the verdict of
whichever anchor appeared first — so a card's empty overlay anchor rejected a page
its own text anchor linked to. On one real page 11 of 11 rejections were wrong. - A stop could advise configuring a capability that was already configured. A
browser solver that ran and produced no clearance yielded "Configure a browser
solver". The message now says the tier ran and failed, and quotes what it
reported. - Rotating with nowhere to go spent the rotation budget on one address. Found
against a host that bans this machine's ASN outright. The pool now reports whether
an alternative exists, and with none the stop is immediate. - A proxy refusing our own credential was diagnosed as the site's IP reputation.
tor-pool 0.2 enforces authentication, and a rejected SOCKS5 handshake never becomes
an HTTP response — so it reacheddiagnose_transport, which blamed the exit for
every proxied transport error. Three wrong things followed: the address was rotated
though nothing was wrong with it, the pool was told a healthy exit wasblocked,
and layer 1 was written to the origin's persisted profile — so a missing token left
behind a permanent verdict that the site refuses this address. A proxy that
refuses us is nowREFUSEwith no layer, matching how HTTP 407 was already handled.
The distinction is drawn on curl's wording rather than the exception class, because
an unreachable destination reported through a SOCKS5 reply raises the same
ProxyErrorand that one really is evidence about the exit. - A failure with nothing to attribute was reported as layer 15.
Blocked
required a layer, so a layer-less stop borrowedLayer.WORKERS— and "L15 Operator
edge code" is indistinguishable from a Cloudflare Worker refusing the request.
Blocked.layeris nowOptionaland renders as "no detection layer". - The
browserextra failed with a raw dependency error on unsupported Pythons.
nodriver raisesTypeErrorbefore 3.10 andSyntaxErroron 3.14; both now
produce a message naming the version floor, and the extra is marked so it does not
install where it cannot load.
Fixed
- One share-button link took out a page's whole crawl frontier.
extract_host
readurlparse(...).port, which raises rather than returningNonewhen the
netloc's:is followed by something that is not a number — so an ordinary
whatsapp:send?text=…anchor abortedsafe_linksfor the entire page. The port is
now optional and the host survives without it.