Repository navigation
v1.0.1
Fixed
-
A response is decoded the way it declares itself.
PageSoup.createread the body
of aResponseas UTF-8 witherrors="ignore"and never looked at what the page said
its charset was, so every multi-byte character on a non-UTF-8 site was silently
dropped — a GBK page returned an empty title rather than a wrong one. The charset now
comes from the first source that declares a usable one: an explicitencoding=
argument, then the response'sContent-Typeheader, then a<meta charset>near the
top of the markup, then UTF-8. A charset the server names but Python cannot load falls
through to the next candidate instead of raising.Deliberately not
response.encoding: requests fills that with ISO-8859-1 for any
text/*response that declared no charset, so preferring it would mojibake exactly the
pages this fixes.Scraper._peekdoes read it, which is why diagnosis and the parsed
soup could disagree about the same bytes. -
parsersurvives aResponse.PageSoup.createrecursed into its own bytes
branch without passingparseron, soScraperConfig.parserandScraper(parser=…)
had no effect onget_soup,post_soupormake_soup(response)— the only paths a
caller uses — and everything was parsed with lxml. -
A challenge is no longer written to a download target.
stream_toblanked the body
before diagnosis ran, so a challenge interstitial — which arrives with a 200 and a body
— was streamed to the caller's path and accepted.get_fileproduced a file that was
really a Cloudflare page, indistinguishable from the asset once the response was gone.
The opening bytes are now held back and diagnosed before the file is created, and the
abort signal is honoured while they are buffered as well as while the rest is written.