Skip to content

v1.0.0

Choose a tag to compare

@github-actions github-actions released this 29 Jul 19:03
· 48 commits to main since this release

A complete rewrite. There are no compatibility shims: almost every import changes.
docs/migration.md is the mapping, and
docs/model.md is why.

The change

The library is now organised around a model of what it is up against, rather than
around a request pipeline with anti-detection features bolted on. A mitigation engine
folds many detectors into one score, and admission is close to a conjunction — so the
weakest layer bounds the outcome, and effort spent on any other layer buys nothing.
Detectors that read an artifact the client emits are reproducible; detectors that
read a property it must possess are not.

Every behaviour below follows from those two statements.

Breaking

  • ScraperEngine is gone, and Scraper is no longer a requests.Session
    subclass. The transport is a two-method seam, so the escalation ladder can move
    between transports and every tier is testable without a network.
  • The in-process Cloudflare solvers are gone (v1, v2, v3, Turnstile), along with
    the exejs dependency. They cannot keep up with the challenge format, and the layer
    they targeted is only reachable by a real browser. A challenged site now needs
    ScraperConfig.browser; without one it raises and says so, instead of attempting a
    solve that usually failed.
  • TLS cipher rotation is gone. Reordering the cipher list per request does not
    produce a browser fingerprint, it produces an unstable one — and an unstable TLS
    fingerprint invalidates any clearance bound to it. The feature was breaking the layer
    above it.
  • Header randomisation is gone. Header order is read, not just header values. An
    impersonation profile emits a complete, correctly ordered set;
    scraper.identity.OVERRIDABLE now caps what may be written over it.
  • The User-Agent is taken from the transport, not imposed on it. The generated-UA
    machinery is gone. A profile supplies the User-Agent until a real browser earns a
    clearance, at which point the browser is the source of truth and its exact string is
    reproduced — because that is what the clearance is bound to.
  • Impersonation is a core dependency, not the impersonate extra. An ordinary
    Python client fails layers 2–5 in the first round trip, so a build without it is not
    a degraded scraper but one that cannot reach a protected page.
  • default_config(), StealthConfig, BrowserConfig, ProxyConfig, ProxyUrl,
    TorProxyUrl, apply_browser_clearance() and the scraper.engine package are
    removed. SharedLimiter becomes SharedState. AbortedException becomes Aborted,
    and the CloudflareException hierarchy becomes Blocked / Impassable /
    Exhausted, each carrying the layer it is attributed to — or None, when the
    failure is ours rather than the site's. Code that dereferences exc.layer or
    exc.layer_info must handle that; a type checker will point at every such place.
  • New extras: browser (nodriver) and botauth (cryptography). impersonate is gone.
  • ScraperConfig.impersonate defaults to "", which resolves to firefox — or to
    chrome when a browser solver is configured, because the bundled solver drives
    Chrome and a clearance is bound to a User-Agent and a TLS fingerprint together.
    Read it through the new ScraperConfig.profile(); the field is no longer the answer
    on its own.

Measured against 0.2.6

The rewrite was A/B'd against 0.2.6 on 150 hosts from lightnovel-crawler's source
index, three arms per host, curl_cffi pinned to the same version in both, one
classifier for every arm, arm order shuffled per host. livetest/compare.py runs it.

The first run said the rewrite had made per-request retrieval slightly worse: 0.2.6
took 56 of 150 hosts with impersonation off, 55 with it on, and 1.0 took 51 — losing
five hosts head to head and winning one. Everything in this section is what closed
that gap, and each item is here because it was measured, not reasoned about:

  • A first request now carries a Referer and matching Sec-Fetch-Site. No
    browser does this — a typed address has no referrer, which is exactly why the
    rewrite sent none. Over 85 hosts that refuse an impersonated client it recovered
    three and cost zero, turning a 403-with-a-challenge into a full page. The coherent
    header set matters more than the header: a synthesised referrer alongside
    Sec-Fetch-Site: none is a contradiction, and one host only yielded once both
    agreed.
  • A JavaScript-only redirect is followed instead of being mistaken for content.
    A family of bot checks answers with a few hundred bytes of
    window.location.replace('…?token=…') — a 200, no challenge marker, no Cloudflare
    header. Both releases handed that back as a successful retrieval, so a scraper
    reported success and collected an empty document. It was 19% of the corpus. No
    browser is needed because the destination is emitted in the HTML; running the
    script would produce the URL already sitting there in plain text. New
    Action.FOLLOW, Diagnosis.location and diagnosis.js_redirect().
  • The default impersonation profile is Firefox. Over a random 150-host sample:
    firefox 85, safari 84, edge 82, chrome 81 — and against chrome, firefox won four
    hosts and lost none. Chrome being the most common browser is a reason to expect it
    to be unremarkable, not evidence that it is the least remarkable.
  • Pool sessions are released when a scraper closes. ExitPool.release() existed,
    was never called, and only dropped the local lease. Every lease minted a fresh key
    that the pool then held until SESSION_TTL, so a process building several scrapers
    in a row walked the pool out of capacity — and the symptom was the misleading part:
    the next lease could not connect, a transport failure through a proxy is evidence
    about the exit, and the model reported a reputation block on a destination that
    never saw the request. Needs tor-pool 0.2.1, which puts
    DELETE /api/sessions/{key} on the proxy scope.

After those, on the same corpus with the archive tier off — so nothing is borrowed
that 0.2.6 cannot do — 1.0 takes 82 of 150 against 0.2.6's best of 76, and 30.8% of the
hard hosts against 25.6%. Head to head it wins nine and loses none (sign test
p = 0.004). Hosts serving an unrecognised stub fell from 28 to 6. With the archive tier
on, total reach was 73% against 39%.

Added

  • scraper.layers — the model as code: nineteen layers, what each reads
    (Trait), what this library does about it (Stance), the bound (weakest) and the
    arithmetic that shows why fixing the wrong layer gains nothing (marginal_gain).
    Layers 2–5 are declared as one barrier, and expand() keeps any reach set closed
    over the group.
  • scraper.diagnosis — a response becomes a binding layer plus an action, as a pure
    function over primitives. Three readings that a status-code table gets wrong: a 200
    carrying a challenge is a failure, a 429 is a pacing problem rather than a spent
    address, and a 403 with error 1010 is about the automation channel rather than the
    address. A 407 is reported as our own proxy credential, not as the site needing a
    login.
  • scraper.planner — chooses the cheapest capability whose reach covers the binding
    layer. Three rules that contradict the conventional table: a possessed property is
    never rotated away from; rotation requires somewhere better to go, so a pool of
    published ranges stops with an explanation instead of cycling; and escalation only
    goes to a tier that actually reaches the layer. Repeated failure at an
    already-covered emitted layer is re-attributed to the per-zone composite, because
    recurrence is the only evidence available from outside.
  • scraper.identity — the emitted signals as one indivisible thing.
    Clearance.usable_by() refuses to replay a clearance under a different identity,
    which makes the classic rotating-proxy failure structurally impossible rather than
    merely documented.
  • scraper.exits — addresses described by kind, and ExitKind.reach deciding what
    layer 1 can be told. Leased per origin and held; rotation happens on evidence, never
    on a timer. tor-pool support is retained, and a failure report now carries the kind
    derived from the binding layer.
  • scraper.pacing — inter-request gaps drawn from a gamma distribution rather than set
    to a constant, occasional reading pauses, homepage warm-up, and a real referrer chain
    with fetch metadata. Throttles widen a learned per-origin interval that persists.
  • scraper.memory — per-origin state that survives the process: the binding layer, the
    working tier, a clearance and the identity it belongs to, the learned interval,
    observed JSON endpoints, and recorded decoy URLs. On by default, because the layer it
    exists for cannot be satisfied by a process that forgets.
  • scraper.state.SharedState — shares the address, identity, history, pacing, referrer
    chain and decoy list between scrapers pointed at one site. Two scrapers with separate
    state do not look like one visitor going faster; they look like two who contradict
    each other.
  • scraper.tiers — archive (Wayback, serving the original URL so links resolve
    against the real site), direct (the baseline), clearance (solve once, reuse many,
    delegating every request to direct so the solve and the fetch cannot diverge), and
    managed (a provider callable; none bundled, since a wrapper that guesses a vendor
    format wrong fails in a way that looks like the site blocking you).
  • scraper.browser — a two-method BrowserSolver protocol, a nodriver adapter, and
    CallableSolver for anything else. Headed and WebRTC-disabled by default, both
    deliberately: a headless build reports a software renderer, and a STUN request reports
    the host's real address past the proxy without any request failing. One browser
    profile directory per address.
  • scraper.links — safe_links enumerates only anchors a person could click, and
    TopicGuard notices content that stopped being about the site. This is the only
    defence against the one layer that returns no error, so the guard runs on the way out
    rather than on demand.
  • scraper.botauth — RFC 9421 Ed25519 request signing with the web-bot-auth tag,
    plus the key directory document to publish. The one layer with no bypass, and for a
    crawler willing to identify itself, the cheapest tier in the stack.
  • Scraper.explain(url) and Scraper.knows(url) — what the library concluded is
    binding, which tier settled, how fast it has learned it can go, and what it has
    available.
  • livetest/ — a live verification harness that exercises every path against real
    Cloudflare deployments, using every host in lightnovel-crawler's source index as
    the corpus. Separate from tests/, which stays offline. See
    livetest/README.md; the current run is
    livetest/report.html.

Fixed before release, found by live traffic

Most of what was fixed before release was found this way, and none of it was visible to
a stubbed transport. These are not regressions from 0.2.x — they are defects in the new
code, and two of them made a whole feature silently useless while every unit test
passed. Each has a regression test whose docstring says it was found live.

  • Cloudflare's injected JavaScript-Detections script was read as a challenge.
    The script is served from /cdn-cgi/challenge-platform/scripts/jsd/… on ordinary
    successful pages, and that path was a challenge marker — so content pages were
    diagnosed as interstitials. Measured across two live populations before changing
    it: the bare prefix appeared on 9 of 10 normally-served pages, the challenge-only
    /h/ orchestrate sub-path on none. 18 of 22 hosts reported as challenged were
    serving content fine. The same marker also stopped the browser solve loop from
    ever detecting "cleared", so every solve burned its full timeout.
  • The archive tier could never find a capture. A negative Wayback CDX limit
    is documented as "the last N rows" but returns an empty body once a filter is
    applied. The query is now bounded server-side and the newest rows taken from the
    tail. Separately, an unbounded query timed out on popular URLs, and a
    rate-limited index was reported as "nothing archived" — all three produced the
    same misleading message, so a lookup failure, an empty index and an age limit now
    say which they are, and the index is retried once.
  • Real navigation was being dropped as decoy content. An anchor containing only
    an icon-font element counted as "nothing rendered", and a URL took the verdict of
    whichever anchor appeared first — so a card's empty overlay anchor rejected a page
    its own text anchor linked to. On one real page 11 of 11 rejections were wrong.
  • A stop could advise configuring a capability that was already configured. A
    browser solver that ran and produced no clearance yielded "Configure a browser
    solver". The message now says the tier ran and failed, and quotes what it
    reported.
  • Rotating with nowhere to go spent the rotation budget on one address. Found
    against a host that bans this machine's ASN outright. The pool now reports whether
    an alternative exists, and with none the stop is immediate.
  • A proxy refusing our own credential was diagnosed as the site's IP reputation.
    tor-pool 0.2 enforces authentication, and a rejected SOCKS5 handshake never becomes
    an HTTP response — so it reached diagnose_transport, which blamed the exit for
    every proxied transport error. Three wrong things followed: the address was rotated
    though nothing was wrong with it, the pool was told a healthy exit was blocked,
    and layer 1 was written to the origin's persisted profile — so a missing token left
    behind a permanent verdict that the site refuses this address. A proxy that
    refuses us is now REFUSE with no layer, matching how HTTP 407 was already handled.
    The distinction is drawn on curl's wording rather than the exception class, because
    an unreachable destination reported through a SOCKS5 reply raises the same
    ProxyError and that one really is evidence about the exit.
  • A failure with nothing to attribute was reported as layer 15. Blocked
    required a layer, so a layer-less stop borrowed Layer.WORKERS — and "L15 Operator
    edge code" is indistinguishable from a Cloudflare Worker refusing the request.
    Blocked.layer is now Optional and renders as "no detection layer".
  • The browser extra failed with a raw dependency error on unsupported Pythons.
    nodriver raises TypeError before 3.10 and SyntaxError on 3.14; both now
    produce a message naming the version floor, and the extra is marked so it does not
    install where it cannot load.

Fixed

  • One share-button link took out a page's whole crawl frontier. extract_host
    read urlparse(...).port, which raises rather than returning None when the
    netloc's : is followed by something that is not a number — so an ordinary
    whatsapp:send?text=… anchor aborted safe_links for the entire page. The port is
    now optional and the host survives without it.