Skip to content

v1.0.0

Latest

Choose a tag to compare

@rushter rushter released this 03 Oct 15:01
· 1 commit to master since this release

This release contains breaking changes. It focuses on speed improvements and corner cases.

The main change is that the Modest backend is no longer available.
It is outdated and unmaintained since 2021, contains bugs, and does not follow modern HTML5 standards.

The second is that parse_fragment() is gone. It re-implemented fragment parsing by guessing
whether the input was a document or a fragment; LexborHTMLParser(html, is_fragment=True) already
does this properly, so the guessing layer has been removed.

The rest of the changes fix corner cases where the bugs were happening rarely,
usually when heavily modifying the tree.

  • Breaking: Remove the Modest backend. selectolax.parser is now a stub that raises ImportError
    on import. Use the lexbor backend (from selectolax.lexbor import LexborHTMLParser) instead.
  • Breaking: Remove parse_fragment().
  • Breaking: Fix css_matches and any_css_matches missing matches outside the first top-level node of an HTML
    fragment. They now search the same scope as css.
  • Breaking: Fix tags and strip_tags doing nothing on an HTML fragment. tags returned an empty list and
    strip_tags removed nothing, because the fragment was not reachable from the document node.
  • Breaking: Fix the internal <html> wrapper of an HTML fragment being reported as a match.
  • Breaking: Fix scripts_contain and script_srcs_contain missing matches outside the first top-level
    node of an HTML fragment. They now search the same scope as css.
  • Breaking: Fix scripts_contain and script_srcs_contain answering from another scope's cached
    result. The cache is keyed by the node the search is rooted at rather than the node it was called on.
  • Breaking: Fix select() missing matches outside the first top-level node of an HTML fragment.
    It now searches the same scope as css.
  • Breaking: Fix traverse() covering only the first top-level node of an HTML fragment, and
    yielding nothing when the fragment starts with a text node.
  • Breaking: Fix attribute_longer_than and any_attribute_longer_than returning inconsistent results
  • Breaking: attributes and attrs now report an attribute's qualified name instead of its local name.
    An element carrying both href and xlink:href used to collapse them into a single href key.
  • Breaking: Fix iter() skipping the remaining children when a node is removed during iteration
  • Breaking: Fix an HTML fragment whose first node is text dropping the rest of the fragment.
    LexborHTMLParser('a<span>s</span>', is_fragment=True).text() returned 'a' instead of 'as'.
  • Fix text() silently returning truncated text instead of raising when a fragment fails to be collected.
  • Fix scripts_contain and script_srcs_contain sometimes returning wrong results due to HTML mutations.
  • Fix a single undecodable byte in an untrusted document making html, inner_html, html_pretty,
    attributes, attrs, id, tag, text_lexbor and text_content raise UnicodeDecodeError.
    Bytes Lexbor passes through verbatim are now substituted with U+FFFD, the way text() already did.
  • Fix text_lexbor sometimes holding temporary memory longer than needed
  • Improve memory consumption, potential stack overflow and slow extraction in merge_text_nodes(). It is now up to 10 times faster.
  • Improve speed of __eq__
  • Improve speed of attrs.items() and attrs.values().
  • Fix the inner_html setter attaching element children to non-element nodes
  • Fix head and body dangling after setting inner_html on the <html> element
  • Fix the inner_html setter freeing the replaced children
  • Fix root going stale on an HTML fragment.
  • Fix a segfault when walking the tree from the root of an HTML fragment that had been detached from the
    fragment, for example by unwrap().
  • Avoid segfaults when hitting OOM
  • Improve performance of the text method. Text fragments are now concatenated as raw bytes, up to 5x faster.
  • Fix skip_empty being ignored by text(deep=True)
  • Fix text() raising UnicodeDecodeError on undecodable bytes when deep=False. It now substitutes
    U+FFFD, like the deep path always did
  • Fix handling of the id method on text nodes
  • Fix attrs[key] = value raising AttributeError instead of TypeError when value is not a string
  • Prevent segfaults when instantiating LexborNode or LexborAttributes directly.
  • Fix unwrap() corrupting the tree in some cases.
  • Fix head and body going stale once <head>/<body> is removed from the document.
  • Fix attrs reading freed memory when it outlives the node it was obtained from.
  • Fix memory leak in attrs[key] = None; it leaked the value buffer header on every call.
  • Fix possible memory leak in clone()
  • Add encoding=True to LexborHTMLParser, which detects the encoding of bytes input and
    transcodes it to UTF-8 before parsing.