Skip to content

EdgarTools v5.0.0 - HTMLParser Migration Complete

Choose a tag to compare

@dgunning dgunning released this 08 Dec 21:34
· 1708 commits to main since this release

EdgarTools v5.0.0 - HTMLParser Migration Complete

This major release completes the migration to the new HTMLParser for 10-K and 10-Q document processing, delivering improved accuracy, consistency, and maintainability.

Highlights

  • Complete HTMLParser Migration - TenK and TenQ now use the new edgar.documents.HTMLParser by default
  • Enhanced Section Detection - Pattern-based detection with comprehensive coverage of all 10-K/10-Q item patterns
  • Improved Table of Contents Detection - Configurable strategies for better section boundary identification
  • XBRL Enhancements - Added parent_concept column to statement DataFrames for hierarchical analysis

Key Changes

Document Parsing

  • Complete migration from legacy ChunkedDocument to HTMLParser for 10-K and 10-Q
  • Pattern-based section detection handles all item variations (Items 1-15 for 10-K, Items 1-6 for 10-Q)
  • Part I / Part II boundary detection with proper context tracking
  • Bold paragraph and table cell fallback detection
  • Legacy parser available as fallback (deprecated, removal planned for v6.0)

XBRL Improvements

  • Added parent_concept column to XBRL statement DataFrames
  • Removed misleading "(Standardized)" suffix from statement titles
  • Standardization now applied transparently where mappings exist

Bug Fixes

  • Fixed HTTP 304 Not Modified response handling
  • Fixed missing log import in company_reports modules
  • Fixed table processing bounds checking to prevent IndexError
  • Added missing is_section_header method to HeaderDetectionStrategy

Breaking Changes

  • Document parsing now uses HTMLParser by default
  • Legacy ChunkedDocument available as fallback but deprecated
  • Some section keys have changed format (e.g., part_i_item_1 instead of Item 1)

Migration Guide

Most users won't need to change anything. If you relied on specific section key formats:

# Old format (still works via fallback)
ten_k['Item 1']

# New preferred format
ten_k['Part I, Item 1']
# or
ten_k.sections['part_i_item_1'].text()

Upgrade

pip install --upgrade edgartools

Full Changelog: v4.35.1...v5.0.0