Skip to content

v5.13.0 - Performance & 13F Enhancements

Choose a tag to compare

@dgunning dgunning released this 30 Jan 01:24
· 1310 commits to main since this release

Performance Improvements

13F XML Parsing Optimization (8x faster)

  • Major Performance Gains: Replaced BeautifulSoup with lxml's etree for parsing 13F information tables
    • Parse time: 7.4s → 0.6s for 24K holdings (11.5x faster)
    • Total time: 7.7s → 0.9s (8.2x faster)
    • Real-world impact: 100 filings 12.8min → 1.6min
    • Function calls reduced by 98% (121M → 2.8M)
  • Maintains 100% backward compatibility with all tests passing

New Features

13F Holdings Analysis

  • holdings_view(): Improved default display with configurable limits to prevent terminal flooding
  • compare_holdings(): Compare current vs previous quarter holdings with:
    • Share/value deltas and percentage changes
    • Status labels (NEW/CLOSED/INCREASED/DECREASED/UNCHANGED)
    • Configurable display limits (default 200 rows)
  • holding_history(periods=4): Multi-quarter tracking with:
    • Unicode sparkline trends showing share count evolution
    • Mean-centred ±50% scaling for proportional visualization
    • Chg% column showing first-to-last percentage change
    • Automatic deduplication of amendments
    • Smart display limits (default 100 rows for readability)

Enhanced User Experience

  • Styled Warning Messages: Rich-formatted warnings, errors, and info messages using badge styles
  • Improved Data Quality: Added 336 new companies to reference data (DCX, VTIX, +334 more)
  • Reduced Noise: Fixed spurious warnings when searching for current year filings

Bug Fixes

8-K Earnings Parsing

  • Fixed negative number detection (avoid false positives like "500 (estimated)")
  • Fixed HTML caption insertion to preserve proper table tag structure
  • Fixed potential crash when cell.content is None
  • Fixed dtype error with pandas StringDtype columns

13F Holdings Rendering

  • Fixed NaN and missing column handling with safe _is_number() helper
  • Ensured all expected columns exist with safe defaults
  • Fixed PutCall NaN value handling

Cross-Platform Compatibility

  • Fixed CIK dtype check to handle all string types across different pandas/PyArrow versions
  • Fixed CIK format strings in entity_facts and submissions to handle both int and string types
  • Added proper error logging for parquet file loading

Reference Data & Indexing

  • Fixed cross reference index tests to use specific filing instead of .latest()
  • Updated reference data format to match edgar-storage schema
  • Fixed None exchange display bug in Company.rich()

Internal Improvements

  • Refactored 13F views as view classes with common iteration protocol (iter, getitem)
  • Added lazy loading for _related_filings to avoid unnecessary network requests
  • Filtered _related_filings to 13F forms only to prevent form assertion errors
  • Improved test reliability with specific filing references

📦 Install: `pip install edgartools==5.13.0`