Skip to content

0.3.0

Choose a tag to compare

@pablotoledo pablotoledo released this 26 Mar 09:36
· 20 commits to main since this release
8ee105b

πŸš€ New Features

Web Page to Markdown Conversion

Readium now supports direct web page content extraction and conversion to Markdown! This exciting update allows users to:

  • πŸ”— Convert web pages to clean, readable Markdown
  • πŸ›  Process URLs alongside local directories and repositories
  • πŸŽ› Configure content extraction with flexible modes:
    • clean: Extract only main content (default)
    • full: Preserve most page content

Key Enhancements

  • Powerful Web Scraping: Leveraging Trafilatura for intelligent content extraction
  • Configurable Processing:
    • Control table, image, and link inclusion
    • Choose between focused and comprehensive extraction modes
  • Seamless Integration: New functionality works alongside existing Readium features

πŸ›  CLI and API Updates

Command Line Examples

# Convert webpage to Markdown
readium https://example.com/docs

# Full content mode
readium https://example.com/docs --url-mode full

# Save to specific output file
readium https://example.com/docs -o webpage.md

Python API

from readium import Readium, ReadConfig

config = ReadConfig(
    url_mode='clean',      # 'clean' or 'full'
    include_tables=True,
    include_images=True
)

reader = Readium(config)
summary, tree, content = reader.read_docs('https://example.com/docs')

πŸ” Processing Modes

  • Clean Mode (Default):

    • Focuses on main content
    • Removes menus, ads, and navigation elements
    • Ideal for documentation and technical content
  • Full Mode:

    • Preserves more page structure
    • Includes additional elements
    • Useful for comprehensive content capture

πŸ“¦ Dependencies

πŸ”’ Compatibility

  • Python 3.10-3.12
  • Minimal impact on existing Readium workflows
  • Optional web processing functionality