Repository navigation
v1.1.0
Added
retry_backoffandmax_retry_wait. Base seconds before a retry, doubled per
attempt and capped. Only consulted when the response named no delay — aRetry-After
header always wins.
Changed
-
The
browserextra is marked for 3.10 to 3.13. nodriver cannot be imported outside
that range: below 3.10 its module body evaluates a PEP 604 union, and from 3.14 its
generatedcdp/network.pyfails to tokenize on a stray non-UTF-8 byte. Only the first
was handled, and only in the exception clause — so on 3.14, which is what a modern
Docker image and release build use, the extra installed and then raised a bare
SyntaxErrorfrom inside a dependency.NoDriverSolvernow raisesMissingDependency
naming the supported range for either failure, and the marker keeps the extra out of an
environment that cannot use it. -
A concurrency gate is keyed per address and origin. With no proxy configured every
origin shared the literal exit iddirect, somax_sessions_per_exit— clamped to the
low single digits — gated the whole process rather than one address. A consumer
crawling several sites at once was serialised across all of them, which is what forced
building oneSharedStateper domain. The id is nowdirect#<origin>.Identity tokens and stored clearances key on the exit id too, so this narrows what they
match — the safe direction, since a clearance issued for one site was never usable on
another. It does mean every clearance already inorigins.jsonstops matching once, on
upgrade. Self-healing: the next solve replaces it. A first run after upgrading will
look colder than it is.The browser profile directory is deliberately not per origin.
profile_dir_forkeys
on the address, because a Chrome profile is tens of megabytes and a consumer with a few
hundred sources would otherwise keep one for each. -
A transport failure through a proxy no longer claims layer 1.
diagnose_transport
attributedIP_REPUTATIONto any connection error through an exit. The address is still
blamed and still rotated, but with no layer, because the site never answered — there is
nothing to conclude about it. Attributing reputation wrote a permanent verdict onto the
origin's profile that the destination refuses us, from evidence that only says one
address failed to carry a request.Exhausted.layerisNonefor this case now, and
the pool is toldtransportrather than a reputation kind.
Fixed
-
Rotation is reachable with a pool of published ranges. The planner asked
IP_REPUTATION in exit_reachbefore every rotation.ExitKind.TOR.reachis empty and
honestly so — Tor exit lists are published — so with aTorPoolSpecconfigured
Move.ROTATEwas never emitted andExitPool.rotateandExitPool.reportwere
unreachable fromfetchentirely: a dead pool instance could never be replaced. The
check now applies only when the site actually attributed reputation, which is the
question it answers. Whether rotating can produce a different address at all is
ExitPool.rotatable, and aTorPoolSpeccounts as several addresses where a single
plain proxy does not.ExitKind.TOR.reachis deliberately unchanged. Widening it would make the planner
recommend rotation as a cure for reputation blocks it cannot cure. -
An unconfigured pool reports no reach.
best_kindfalls back toDIRECT, whose
reach includes layer 1 — correct when there is a burnt proxy to move off, meaningless
withexits=[]. Reported anyway, it told the planner a remedy was available that it had
no way to perform.ExitKind.DIRECT.reachitself is unchanged: moving off a datacenter
proxy onto direct genuinely can clear layer 1. -
An exit that names a kind must name an address.
ExitSpec(kind=ExitKind.MOBILE)
with nourlwas accepted, reported mobile reach, and printedexits: mobilein
explain()while every packet left from the local address. It raisesValueErrornow.
ExitSpec(kind=ExitKind.DIRECT)is still valid — that is what a fallback-to-direct
entry looks like. -
Retries back off.
Action.RETRYwaitedretry_after or 0.0, and 408, 502, 504 and
the 52x family never parse aRetry-After— so a retry on any of those was sent
back-to-back, a tight loop aimed at a site already struggling. -
get_imageraisesMissingDependencywhen Pillow is absent, naming theimage
extra, instead of a bareModuleNotFoundErrorfrom the middle of the call. Every cover
and inline image goes through it. -
A tier closes only a transport it owns.
DirectTier.close()closed whatever
transport it held, including one handed in throughScraperConfig.transport. Two
scrapers sharing an injected transport broke each other on the firstclose(). -
Retired addresses no longer leak their gates and clearances.
ExitPool._slotsheld
a semaphore per exit id and every rotation minted a new id, so a long-running process
accumulated one per rotation forever; the gate of a proxied lease is now dropped with the
lease, whose session key means the id can never be asked for again.ClearanceTier._held
was keyed by origin and never evicted, holding cookies long past their expiry; expired
entries are dropped on each solve, with a cap as a backstop.Browser profile directories were the third and largest of these, at tens of megabytes
each. A proxied exit id carries a session key, so every rotation that reached a solve
left another one behind and nothing ever removed it.profile_dir_fornow prunes the
least recently used beyondMAX_PROFILES, skipping anything touched in the last few
minutes — two scrapers can share a data dir, and each solver only serialises against
itself, so the directory being removed must not be one another process has a browser in.
Keying them coarsely was not the alternative: for a pool endpoint the URL is constant
while the exit IP is not, so one shared profile would hand a fresh session the
accumulated history of a burnt exit. -
Success is recorded only for a response that succeeded. Any status without a
matching diagnosis is reported asACCEPT— correctly, since nothing about it says a
layer is blocking — but the accept path then wrote a success unconditionally. A site
answering 439 to everything setprofile.tier, incrementedsuccessesand zeroed
consecutive_failures, teaching the store that whatever tier had just been tried
works. The bar is now that the site responded: a 2xx or a 3xx records the tier, and a
4xx or 5xx records nothing.402,405,410and423are counted against the origin as failures, because those
are a site refusing this visitor rather than answering about a path. They are recorded
with no layer — none of them identifies one, and naming one would retire a healthy
exit over what may be a URL mistake. A404still moves the ledger in neither
direction and still surfaces as a plainHTTPError. -
A throttle is counted once. The handler for
BACKOFFandACCUMULATErecords the
failure itself, with the widened interval, and the retrieval loop recorded it a second
time. With the defaultpromote_after=3that meant a third failure — and an escalation
to a tier the caller may not have configured — on the second 429. -
A failure with nothing to attribute no longer erases the binding layer.
record_failure(url, None)assignedNoneover whatever an earlier, attributed
failure had learned, so a transport error or an unmatched status discarded the single
most valuable thing the store holds and sent the next run back to guessing.