Skip to content

v1.1.0

Choose a tag to compare

@github-actions github-actions released this 30 Jul 09:16
· 38 commits to main since this release

Added

  • retry_backoff and max_retry_wait. Base seconds before a retry, doubled per
    attempt and capped. Only consulted when the response named no delay — a Retry-After
    header always wins.

Changed

  • The browser extra is marked for 3.10 to 3.13. nodriver cannot be imported outside
    that range: below 3.10 its module body evaluates a PEP 604 union, and from 3.14 its
    generated cdp/network.py fails to tokenize on a stray non-UTF-8 byte. Only the first
    was handled, and only in the exception clause — so on 3.14, which is what a modern
    Docker image and release build use, the extra installed and then raised a bare
    SyntaxError from inside a dependency. NoDriverSolver now raises MissingDependency
    naming the supported range for either failure, and the marker keeps the extra out of an
    environment that cannot use it.

  • A concurrency gate is keyed per address and origin. With no proxy configured every
    origin shared the literal exit id direct, so max_sessions_per_exit — clamped to the
    low single digits — gated the whole process rather than one address. A consumer
    crawling several sites at once was serialised across all of them, which is what forced
    building one SharedState per domain. The id is now direct#<origin>.

    Identity tokens and stored clearances key on the exit id too, so this narrows what they
    match — the safe direction, since a clearance issued for one site was never usable on
    another. It does mean every clearance already in origins.json stops matching once, on
    upgrade
    . Self-healing: the next solve replaces it. A first run after upgrading will
    look colder than it is.

    The browser profile directory is deliberately not per origin. profile_dir_for keys
    on the address, because a Chrome profile is tens of megabytes and a consumer with a few
    hundred sources would otherwise keep one for each.

  • A transport failure through a proxy no longer claims layer 1. diagnose_transport
    attributed IP_REPUTATION to any connection error through an exit. The address is still
    blamed and still rotated, but with no layer, because the site never answered — there is
    nothing to conclude about it. Attributing reputation wrote a permanent verdict onto the
    origin's profile that the destination refuses us, from evidence that only says one
    address failed to carry a request. Exhausted.layer is None for this case now, and
    the pool is told transport rather than a reputation kind.

Fixed

  • Rotation is reachable with a pool of published ranges. The planner asked
    IP_REPUTATION in exit_reach before every rotation. ExitKind.TOR.reach is empty and
    honestly so — Tor exit lists are published — so with a TorPoolSpec configured
    Move.ROTATE was never emitted and ExitPool.rotate and ExitPool.report were
    unreachable from fetch entirely: a dead pool instance could never be replaced. The
    check now applies only when the site actually attributed reputation, which is the
    question it answers. Whether rotating can produce a different address at all is
    ExitPool.rotatable, and a TorPoolSpec counts as several addresses where a single
    plain proxy does not.

    ExitKind.TOR.reach is deliberately unchanged. Widening it would make the planner
    recommend rotation as a cure for reputation blocks it cannot cure.

  • An unconfigured pool reports no reach. best_kind falls back to DIRECT, whose
    reach includes layer 1 — correct when there is a burnt proxy to move off, meaningless
    with exits=[]. Reported anyway, it told the planner a remedy was available that it had
    no way to perform. ExitKind.DIRECT.reach itself is unchanged: moving off a datacenter
    proxy onto direct genuinely can clear layer 1.

  • An exit that names a kind must name an address. ExitSpec(kind=ExitKind.MOBILE)
    with no url was accepted, reported mobile reach, and printed exits: mobile in
    explain() while every packet left from the local address. It raises ValueError now.
    ExitSpec(kind=ExitKind.DIRECT) is still valid — that is what a fallback-to-direct
    entry looks like.

  • Retries back off. Action.RETRY waited retry_after or 0.0, and 408, 502, 504 and
    the 52x family never parse a Retry-After — so a retry on any of those was sent
    back-to-back, a tight loop aimed at a site already struggling.

  • get_image raises MissingDependency when Pillow is absent, naming the image
    extra, instead of a bare ModuleNotFoundError from the middle of the call. Every cover
    and inline image goes through it.

  • A tier closes only a transport it owns. DirectTier.close() closed whatever
    transport it held, including one handed in through ScraperConfig.transport. Two
    scrapers sharing an injected transport broke each other on the first close().

  • Retired addresses no longer leak their gates and clearances. ExitPool._slots held
    a semaphore per exit id and every rotation minted a new id, so a long-running process
    accumulated one per rotation forever; the gate of a proxied lease is now dropped with the
    lease, whose session key means the id can never be asked for again. ClearanceTier._held
    was keyed by origin and never evicted, holding cookies long past their expiry; expired
    entries are dropped on each solve, with a cap as a backstop.

    Browser profile directories were the third and largest of these, at tens of megabytes
    each. A proxied exit id carries a session key, so every rotation that reached a solve
    left another one behind and nothing ever removed it. profile_dir_for now prunes the
    least recently used beyond MAX_PROFILES, skipping anything touched in the last few
    minutes — two scrapers can share a data dir, and each solver only serialises against
    itself, so the directory being removed must not be one another process has a browser in.
    Keying them coarsely was not the alternative: for a pool endpoint the URL is constant
    while the exit IP is not, so one shared profile would hand a fresh session the
    accumulated history of a burnt exit.

  • Success is recorded only for a response that succeeded. Any status without a
    matching diagnosis is reported as ACCEPT — correctly, since nothing about it says a
    layer is blocking — but the accept path then wrote a success unconditionally. A site
    answering 439 to everything set profile.tier, incremented successes and zeroed
    consecutive_failures, teaching the store that whatever tier had just been tried
    works. The bar is now that the site responded: a 2xx or a 3xx records the tier, and a
    4xx or 5xx records nothing.

    402, 405, 410 and 423 are counted against the origin as failures, because those
    are a site refusing this visitor rather than answering about a path. They are recorded
    with no layer — none of them identifies one, and naming one would retire a healthy
    exit over what may be a URL mistake. A 404 still moves the ledger in neither
    direction and still surfaces as a plain HTTPError.

  • A throttle is counted once. The handler for BACKOFF and ACCUMULATE records the
    failure itself, with the widened interval, and the retrieval loop recorded it a second
    time. With the default promote_after=3 that meant a third failure — and an escalation
    to a tier the caller may not have configured — on the second 429.

  • A failure with nothing to attribute no longer erases the binding layer.
    record_failure(url, None) assigned None over whatever an earlier, attributed
    failure had learned, so a transport error or an unmatched status discarded the single
    most valuable thing the store holds and sent the next run back to guessing.