0.3.0
π New Features
Web Page to Markdown Conversion
Readium now supports direct web page content extraction and conversion to Markdown! This exciting update allows users to:
- π Convert web pages to clean, readable Markdown
- π Process URLs alongside local directories and repositories
- π Configure content extraction with flexible modes:
clean: Extract only main content (default)full: Preserve most page content
Key Enhancements
- Powerful Web Scraping: Leveraging Trafilatura for intelligent content extraction
- Configurable Processing:
- Control table, image, and link inclusion
- Choose between focused and comprehensive extraction modes
- Seamless Integration: New functionality works alongside existing Readium features
π CLI and API Updates
Command Line Examples
# Convert webpage to Markdown
readium https://example.com/docs
# Full content mode
readium https://example.com/docs --url-mode full
# Save to specific output file
readium https://example.com/docs -o webpage.mdPython API
from readium import Readium, ReadConfig
config = ReadConfig(
url_mode='clean', # 'clean' or 'full'
include_tables=True,
include_images=True
)
reader = Readium(config)
summary, tree, content = reader.read_docs('https://example.com/docs')π Processing Modes
-
Clean Mode (Default):
- Focuses on main content
- Removes menus, ads, and navigation elements
- Ideal for documentation and technical content
-
Full Mode:
- Preserves more page structure
- Includes additional elements
- Useful for comprehensive content capture
π¦ Dependencies
- Added [Trafilatura](https://github.com/adbar/trafilatura) for intelligent web content extraction
π Compatibility
- Python 3.10-3.12
- Minimal impact on existing Readium workflows
- Optional web processing functionality