Releases: matplo/rootfileviewer
Release list
v0.13.1 -- final release, renamed to datafileviewer
Final courtesy release -- this project has been renamed to datafileviewer.
rootfileviewer outgrew its name a while ago: it also reads Parquet, HDF5,
numpy, and pandas-readable files through a pluggable backend architecture,
plus TUI features (log-axis toggling, PNG export) that have nothing to do
with ROOT specifically. Development continues under the new identity,
full commit history carried over:
This release (0.13.1) is documentation-only -- no functional changes.
rootfileviewer is now frozen at this version: still installable for
anyone with it pinned, but there will be no further updates here.
pip install 'datafileviewer[all]' for everything this package offered,
under datafileviewer/dfv/dfvt instead of
rootfileviewer/rfv/rfvt.
v0.13.0
Export the current TUI plot as a real PNG -- press p while a plot is showing. The image is self-documenting (node name as the title, a footer naming the source file and the same sampling note shown in the detail panel) and saved to the current directory as <source-file-stem>_<node-name>.png, overwritten on repeat presses.
Uses matplotlib, an optional extra (pip install 'rootfileviewer[matplotlib]' or [all]) following the same lean-by-default pattern as pyarrow/h5py/pandas -- without it, p shows a toast explaining how to install it rather than crashing. If a logarithmic x/y axis is toggled (x/y) when you press p, the PNG reflects that too, using matplotlib's own native log-scale axes.
v0.12.0
Split narrow multi-dim numpy/HDF5 arrays into individually plottable columns, using generic column_N names when there's no real name to use (numpy files have no attribute mechanism at all, unlike HDF5). Applies whenever a multi-dim array's last axis is 20 entries or narrower and doesn't already qualify for HDF5's named-feature split -- wider arrays (e.g. embeddings) still flatten as before, since that's presumed to be genuinely homogeneous data.
Reported and verified against a real file: a (6013910, 4) array held 4 distinct quantities (an integer-like ID and 3 spatial coordinates) that were previously flattened together into one meaningless combined histogram; now each is its own selectable, plottable column_0..column_3.
v0.11.0
Add x/y keys in the interactive TUI to toggle a logarithmic x-/y-axis on the current plot, working identically across every supported format (ROOT, Parquet, HDF5, numpy, pandas).
For a branch/column/dataset/array, x does a genuine rebin with logarithmically-spaced bin edges (not just a compressed x-axis on the same linear bins), so the shape stays meaningful for data spanning multiple decades -- e.g. a pt spectrum. Requires all values to be strictly positive; otherwise reports a clean plot error rather than crashing. y toggles a log count-axis and never has that restriction (counts are never negative). A TH1/TProfile histogram has no raw data left to rebin, so x there re-renders its existing bins on a log-looking axis instead.
plotext has no native log-axis support, so this is implemented by plotting log10(x) and relabeling the ticks with the real values -- verified this actually looks right (plotext places bars at their literal linear position, so log-spaced values alone would render bunched together without this).
v0.10.0
Add HDF5 named-feature dataset support: a dataset whose last axis has named columns via an attribute (a <dataset-name>_features convention, checked on the dataset itself or its parent group) is now split into individually selectable/plottable columns instead of being flattened into one meaningless combined histogram. Detection is conservative -- only applied when a candidate attribute's length exactly matches the dataset's last axis -- so files without this convention are completely unaffected. Reported and verified against a real physics dataset with a (9764, 7) 'jet' dataset holding 7 distinct named quantities.
v0.9.1
Fix a crash (OverflowError deep in plotext) when plotting a float16 branch/column/dataset/array containing large-but-finite values. np.histogram() computes bin edges in the input's own dtype, so two large float16 edges could overflow when added together to get a bin center -- now the array is upcast to float64 before histogramming, and genuine non-finite (inf/nan) values are filtered with a clear note rather than crashing or raising an obscure numpy error. Format-agnostic fix (shared core.py code) reported against a real HDF5 file.
v0.9.0
Add pandas-readable file support: .csv, .pkl/.pickle (pickled DataFrame), .feather, and .jsonl/.ndjson (JSON Lines) -- one-shot, terse, and interactive TUI modes, including value-distribution plotting, with ragged (list-valued) columns flattened and plotted the same way as jagged ROOT branches. pandas is an optional extra (pip install 'rootfileviewer[pandas]' or '[all]'); the base install stays lean.
Also fixes a latent bug in the missing-dependency error message: it now names the file's actual extension (e.g. '.csv') rather than the backend's internal name, which mattered once one backend (pandas) started covering several different extensions at once.
This completes the multi-format backend work: rootfileviewer now reads ROOT, Parquet, HDF5, numpy (.npy/.npz), and pandas-readable files through the same pluggable architecture, with jagged/ragged/variable-length data flattened and plotted consistently across all of them.
v0.8.1
Add numpy file support (.npy/.npz) -- one-shot, terse, and interactive TUI modes, including value-distribution plotting, with ragged (variable-length) arrays flattened and plotted the same way as jagged ROOT branches. No new dependency needed -- numpy is already a bundled dependency, so this works with the base install, no extras group required.
v0.8.0
Add HDF5 file support (.h5/.hdf5), same as Parquet: one-shot, terse, and interactive TUI modes, including value-distribution plotting for datasets -- with jagged (variable-length) datasets flattened and plotted the same way as jagged ROOT branches. h5py is an optional extra (pip install 'rootfileviewer[hdf5]' or '[all]'); the base install stays lean.
v0.7.1
Fix: jagged Parquet list columns (e.g. list, a per-event list of track energies) are now flattened and plotted, like a jagged ROOT branch, instead of being rejected as unsupported nesting. Only struct/map/union columns are still rejected (flattening those would silently mix unrelated fields into one meaningless distribution).