Repository navigation
What's Changed
Lexidown 0.2.2 introduces an opt-in llm preset for converting scraped HTML into cleaner Markdown for LLM context.
- Add LLM content extraction plugin with preprocessing hook by @EvickaStudio in #7.
- Reduce common page noise using HTML structure and accessibility attributes, without website-specific rules.
- Preserve headings, lists, fenced code blocks, and declared TeX math, with readable fallbacks for complex tables.
- Omit embedded SVG and image data while retaining useful text descriptions.
- Add
includeLinkUrlsandincludeImageUrlsoptions to keep labels while omitting destination URLs. - Support custom DOM cleanup through an optional
preprocess(root)callback.
Usage
The preset is included in the standard Lexidown installation. This example uses the optional requests package to fetch a page:
import requests
from lexidown import TurndownService
from lexidown.plugins.llm import llm
response = requests.get("https://example.com/", timeout=30)
response.raise_for_status()
service = TurndownService(
{
"baseUrl": response.url,
"includeLinkUrls": False,
"includeImageUrls": False,
}
).use(llm)
print(service.turndown(response.text))Both URL options default to True. Setting them to False keeps link text and image descriptions, reducing the amount of text sent to an LLM. Literal URLs inside page text or code remain intact.
Full Changelog: v0.2.1...v0.2.2