Skip to content

v2.0.3: Fix HTML parsing to preserve script tags and UTF-8 encoding

Choose a tag to compare

@rolandschuetz rolandschuetz released this 20 Jan 19:46
BUGFIX: Fix HTML parsing to preserve script tags and UTF-8 encoding

When using DOMDocument::loadHTML, script tags at the beginning of the
content were being moved to the head and subsequently lost when
extracting the body content. Additionally, UTF-8 characters were not
always parsed correctly.

This change wraps the input HTML in a body tag and adds an XML encoding
declaration before loading it into DOMDocument. This forces DOMDocument
to keep script tags in the body and correctly parse UTF-8 characters.

Additionally, moved JavaScript detection and removal logic to the DTO
and added comprehensive unit tests.