Skip to content

v0.2.2

Latest

Choose a tag to compare

@EvickaStudio EvickaStudio released this 10 Sep 14:23
· 3 commits to main since this release
f116ca8

What's Changed

Lexidown 0.2.2 introduces an opt-in llm preset for converting scraped HTML into cleaner Markdown for LLM context.

  • Add LLM content extraction plugin with preprocessing hook by @EvickaStudio in #7.
  • Reduce common page noise using HTML structure and accessibility attributes, without website-specific rules.
  • Preserve headings, lists, fenced code blocks, and declared TeX math, with readable fallbacks for complex tables.
  • Omit embedded SVG and image data while retaining useful text descriptions.
  • Add includeLinkUrls and includeImageUrls options to keep labels while omitting destination URLs.
  • Support custom DOM cleanup through an optional preprocess(root) callback.

Usage

The preset is included in the standard Lexidown installation. This example uses the optional requests package to fetch a page:

import requests

from lexidown import TurndownService
from lexidown.plugins.llm import llm

response = requests.get("https://example.com/", timeout=30)
response.raise_for_status()

service = TurndownService(
    {
        "baseUrl": response.url,
        "includeLinkUrls": False,
        "includeImageUrls": False,
    }
).use(llm)

print(service.turndown(response.text))

Both URL options default to True. Setting them to False keeps link text and image descriptions, reducing the amount of text sent to an LLM. Literal URLs inside page text or code remain intact.

Full Changelog: v0.2.1...v0.2.2