You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
BUGFIX: Fix HTML parsing to preserve script tags and UTF-8 encoding
When using DOMDocument::loadHTML, script tags at the beginning of the
content were being moved to the head and subsequently lost when
extracting the body content. Additionally, UTF-8 characters were not
always parsed correctly.
This change wraps the input HTML in a body tag and adds an XML encoding
declaration before loading it into DOMDocument. This forces DOMDocument
to keep script tags in the body and correctly parse UTF-8 characters.
Additionally, moved JavaScript detection and removal logic to the DTO
and added comprehensive unit tests.