How to parse CDATA text in HTML #2561
Replies: 1 comment
|
Hi there, CDATA sections are XML syntax, so if the source format uses them inside custom namespaced elements, I’d parse the document as XML: Document doc = Jsoup.parse(input, Parser.xmlParser());
String text = doc.expectFirst("ns|foo").wholeText();The HTML syntax only recognizes CDATA inside SVG or MathML foreign content. A tag such as I wouldn’t recommend extracting it from If the surrounding document genuinely requires HTML parsing, the producer should instead emit ordinary text with Otherwise, preprocessing the source before HTML parsing would be the remaining workaround. |
Uh oh!
There was an error while loading. Please reload this page.
In 1.23.1 the CDATA tokenization was aligned with the HTML spec, so CDATA syntax in HTML content is now parsed as a bogus comment rather than a
CDataNode.Is there a recommended migration path for a case wher I would want CDATA as text in HTML still?
Context
I process markup from a third-party system that mixes HTML with custom namespaced elements. Some of these elements carry text bodies wrapped in CDATA, e.g.:
xml
I've been reading the body with
element.wholeText(). Under 1.23.1 this now returns an empty string, because the CDATA is tokenized as a bogus comment under the HTML parser and no longer appears as aCDataNode.Should I manually pares the
Commentnodes, switch to XML parsing or is there another option?All reactions