Skip to content

v4.5.0: Header-based column resolution for robust CSV parsing

Choose a tag to compare

@agude agude released this 20 Jan 02:48
· 72 commits to main since this release

Release Notes:

This release refactors the CSV parsing system to use dynamic header-based column resolution instead of
hardcoded indices. This makes the parser resilient to column reordering in future SWITRS data releases and
includes several performance optimizations for processing large files.

There are no breaking changes to the CLI or database schema in this release.

What's New

  • Header-Based Column Resolution: The parser now reads CSV headers at runtime to determine column
    positions, rather than relying on hardcoded indices. This ensures compatibility if CHP reorders columns in
    future SWITRS data exports.

  • Duplicate Header Detection: Added validation that fails fast with a clear error message if a CSV file
    contains duplicate column headers, preventing subtle data ingestion bugs.

  • Performance Optimizations: Pre-calculated column indices during file initialization to eliminate per-row
    dictionary lookups and iteration. For multi-million row SWITRS files, this reduces overhead in the hot path
    of __set_values and date conversion methods.

  • Automatic BOM Handling: Switched file reading to use utf-8-sig encoding, which automatically strips
    the byte-order mark if present. This is the Pythonic approach compared to manual string manipulation.

  • Empty File Handling: Added graceful handling for empty input files (or files containing only a BOM),
    which previously caused an unhandled StopIteration exception.

  • Code Clarity: Renamed RowClass to row_parser in main.py to accurately reflect that these are
    CSVParser instances, not class types.