Repository navigation
v4.5.0: Header-based column resolution for robust CSV parsing
Release Notes:
This release refactors the CSV parsing system to use dynamic header-based column resolution instead of
hardcoded indices. This makes the parser resilient to column reordering in future SWITRS data releases and
includes several performance optimizations for processing large files.
There are no breaking changes to the CLI or database schema in this release.
What's New
-
Header-Based Column Resolution: The parser now reads CSV headers at runtime to determine column
positions, rather than relying on hardcoded indices. This ensures compatibility if CHP reorders columns in
future SWITRS data exports. -
Duplicate Header Detection: Added validation that fails fast with a clear error message if a CSV file
contains duplicate column headers, preventing subtle data ingestion bugs. -
Performance Optimizations: Pre-calculated column indices during file initialization to eliminate per-row
dictionary lookups and iteration. For multi-million row SWITRS files, this reduces overhead in the hot path
of__set_valuesand date conversion methods. -
Automatic BOM Handling: Switched file reading to use
utf-8-sigencoding, which automatically strips
the byte-order mark if present. This is the Pythonic approach compared to manual string manipulation. -
Empty File Handling: Added graceful handling for empty input files (or files containing only a BOM),
which previously caused an unhandledStopIterationexception. -
Code Clarity: Renamed
RowClasstorow_parserinmain.pyto accurately reflect that these are
CSVParserinstances, not class types.