Skip to content

v1.0.1

Choose a tag to compare

@github-actions github-actions released this 30 Jul 06:06
· 43 commits to main since this release

Fixed

  • A response is decoded the way it declares itself. PageSoup.create read the body
    of a Response as UTF-8 with errors="ignore" and never looked at what the page said
    its charset was, so every multi-byte character on a non-UTF-8 site was silently
    dropped — a GBK page returned an empty title rather than a wrong one. The charset now
    comes from the first source that declares a usable one: an explicit encoding=
    argument, then the response's Content-Type header, then a <meta charset> near the
    top of the markup, then UTF-8. A charset the server names but Python cannot load falls
    through to the next candidate instead of raising.

    Deliberately not response.encoding: requests fills that with ISO-8859-1 for any
    text/* response that declared no charset, so preferring it would mojibake exactly the
    pages this fixes. Scraper._peek does read it, which is why diagnosis and the parsed
    soup could disagree about the same bytes.

  • parser survives a Response. PageSoup.create recursed into its own bytes
    branch without passing parser on, so ScraperConfig.parser and Scraper(parser=…)
    had no effect on get_soup, post_soup or make_soup(response) — the only paths a
    caller uses — and everything was parsed with lxml.

  • A challenge is no longer written to a download target. stream_to blanked the body
    before diagnosis ran, so a challenge interstitial — which arrives with a 200 and a body
    — was streamed to the caller's path and accepted. get_file produced a file that was
    really a Cloudflare page, indistinguishable from the asset once the response was gone.
    The opening bytes are now held back and diagnosed before the file is created, and the
    abort signal is honoured while they are buffered as well as while the rest is written.