Releases: llnl-asr/dftracer-utils
Release list
v0.0.13
Added
- A columnar
DataFrame/Series/LazyFrameengine that is a drop-in for
pandas and polars:import dftracer.utils.pandas as pd(or.polars as pl)
and most code runs unchanged on the SIMD kernels. The pandas surface covers
loc/iloc/at/iatandset_index(the index is a named column,
copy-on-write assignment),groupbyobjects with theaggforms, column
selection (groupby(k)["v"]), a Series key, group-wise transforms
(cumsum,shift,rank,head,nth,ffill,bfill,rolling,
expanding,ewm,take,sample,resample),apply/mapcompiled
into the engine (Python per row only as a last resort, with a warning), the
stranddtaccessors,tz_localize/tz_convert,merge,nlargest,
value_counts,mode,compare,pivot_tableanddescribe; the polars
spellings sit alongside (select(Expr),with_columns,over, thestr
namespace).Seriesmasks combine with&,|and~;==/!=return
a mask (a Series is unhashable, as in pandas and polars). - A hash join on
DataFrame,LazyFrame, the C ABI and Python (join/
merge; inner, left, right, outer, semi, anti and cross). A plan's join sends its
build keys to the scan, which prunes the chunks that cannot match. - Plans (
LazyFrame):explain(),schema()andoutput_schema()without
running;memory_budget/auto_spillbounding every breaker; a source of
your own (Source/dftu_source_vt, registered by name); a plugin's own
plan step (dftu_node_register,LazyFrame.op); aframe_opstep for
every registry table op (unnest,partition_id,compare_agg,window,
gap_fill,asof,interval,concat,union,pivotandto_dummies). - Aggregates:
prod,cumprod, exact groupmedian/quantile,
unique(subset), a per-columnreduce, a whole-framegroup_by(),
first/lastexact across a parallel merge; group keys of any type
(Binary, Float16, null keys as their own group withdropna=False). - Plugin ABI: a plugin transforms the batch every later plugin receives
(transform), reports and releases what it holds against the memory
budget (bytes/reclaim), and a plugin node or slice under a plan is
measured by the same budget;abi_versionis a hash of the header, so a
plugin built against another version is refused at load. benchmarks/dataframe_vs_pandas_polars.py: the engine against pandas,
polars and DuckDB on the same Arrow tables, eagerly and as a plan, with a
correctness check of every result against ours, the cores each engine kept
busy and--memoryfor the peak resident set per op.- The trace
View(C++) andTraceViewer(Python) are aLazyFrameover a
trace scan: generic ops (filter, select, sort, head, join, ...) chain on
it, and the scan absorbs what it can at plan time (filters, projections,
group keys, a trailing row window). Trace terminals (call_tree,
flamegraph,containment,flamegraph_partial,aggregate_partial,
sink_json,sink_trace,materialize) are lazy results, and
collect_allruns several of them over one shared scan. Python adds lazy
[],LazyScalarreductions andLazyResult. - Python
Indexer(bloom=BloomConfig(fields=...))names extra args fields to
index. - The index covers every flat args field by default: numbers by per-chunk
min/max, strings by a per-chunk bloom up to 256 distinct values, so a
filter on any arg prunes chunks (equality prunes on min/max as well).
dftracer_index --no-auto-dimensionsopts out. On a 2M-event trace the
index is 1.2 MB, against 8.7 MB for the previous named defaults. An
existing index rebuilds once.
Changed
-
Precompiled headers are off by default (
DFTRACER_UTILS_ENABLE_PCH). GCC's
.gchcannot be cached by ccache and evicted the rest of the cache; pass
-DDFTRACER_UTILS_ENABLE_PCH=ONto keep them for local builds. -
CI: every merge into
developpublishes a<tag>.postN.dev0prerelease to
PyPI; wheel and Valgrind jobs run sharded (one job per Python version, three
Python and six C++ Valgrind shards); push workflows run onmainand
developonly, so a pull request no longer runs twice. -
The dataframe engine is measured (10M rows, Apple M4 Pro): ahead of pandas
on every benchmark row, of polars on every row but two at the noise floor,
of DuckDB on every row but one within a millisecond of it. The group-by
runs a plain loop over batches of rows with a direct table for dense
integer keys, a word table for string keys and, with many groups on a
string key, a scatter into per-thread partitions; the join uses a
direct-address table and 32-bit index lists; the sort is a sample sort;
filters, comparisons, string predicates, casts, gathers, rolling windows,
///%/**and the dictionary encoder run in parallel; a quantile
reads its column in place and sorts one bucket. -
A plan over a resident frame runs whole-column for every op (an op with no
eager form runs its own cursor over the frame as one morsel); it no longer
streams through a spool or spills to disk. -
Memory detection reads the free and inactive pages on macOS (the auto
budget assumed 1 GB there). -
Series.rolling(...).mean()and the other windows write their output in
place and run a chunk per thread. -
Breaking (C++):
Viewterminals follow theLazyFramenames:
export_json/export_trace/export_countersbecamesink_json/
sink_trace/sink_counters,merge_partials_to_tablebecame
merge_partials,limit/offsetbecamehead/slice,occ_cell
becameresolutionandschema()becamecolumn_info().collect()
returns theDataFrame;lazy()returns the plan.call_tree,
flamegraph,containmentand the partials return plans to collect.
Several outputs over one scan usecollect_allorTraceSession;
caller folds attach withView::branch. -
Breaking (Python):
TraceVieweris aLazyFrame;AggregatedTraceViewer
andSessionVieware gone.DaskTraceViewer.occ_cellbecameresolution,
and itsoffset/limittrim the merged result instead of each shard's. -
The server,
dftracer_view,dftracer_run, the statistics tools and the
C ABI run on the sameView; the separate builder API is internal. -
A shard set aggregates its shards concurrently, as many at once as the
spill budget gives each at least 64 MB, andcollect_allsplits each
plan's spill budget across the plans it runs together, so concurrent work
stays within the budget.
Removed
- Breaking: the previous plugin ABI. A plugin built against it does not
load; rebuild againstdftracer/utils/plugins/abi/plugin.h. - Breaking (C++):
AggregatedView,ViewSession's public constructor,
View::joinandtrace/views/result_batch.h(collect_batch); use
Viewplans,View::branchandLazyFrame::join.
Fixed
-
Prerelease Linux wheels were versioned
.post1.devNbecause the manylinux
container did not receive the computed version; the version is passed in. -
A sketch quantile (
pctin a plan,DDSketch) returned-infonce a
bucket held more than 65535 values; buckets are 32-bit now. -
Column-column arithmetic dropped nulls; a Bool column was gathered by
byte instead of by bit;group_byreturned groups out of first-seen order
after a parallel run; a plugin node was answered asynchronously when the
data was resident. -
df.groupby(series)raised aSystemError: the frame'sintest left a
pending error for a non-string key. -
A scan of an unindexed multi-member trace read the lines at each member
boundary twice. -
Index extra args fields were dropped by the fold-based build, so
dftracer_index --dimensionsdid nothing; nested fields were dropped in the
first-touch build; merged chunk stats ordered numeric min/max as text; and
a float literal (pid == 1.0) probed the bloom as"1.000000"and pruned
matching chunks. -
Session containment branches ignored
phase(), sophase(Events)kept
aggregated records; session aggregations with string-arg predicates or
transformed keys wrote rollups that later reads served wrongly. -
A blocking
get()on a finished coroutine task could miss its result;
LazyFrame::collect()on a temporary plan read the freed plan. -
A plan's export sink flushes when the export ends.
-
Data races: libdeflate chose its kernels on first use from several threads
at once, and the reader's member decode cache read an entry's ready flag
outside the lock that wrote it. -
An expression the column type cannot take (a string column compared with a
number) produced a null column that crashed a latergroup_by; it now
raises an error. -
Interactive web trace viewer gains a counter timeline track (with malformed-value
handling), per-counter pid/tid breakdown, bounded-density serving, and active-time
statistics. -
Rectangle selection in the viewer scopes analysis to a time range, lanes, and rows;
aggregated (ph=3) events are visualized with uniform extrapolation and labeled
"aggregated" in tooltips. -
DLIO event category and name can be remapped through an event-map file.
-
Python:
DaskTraceVieweris exported from thedaskmodule; newTraceViewer/View
bindings expose the query DSL and portable C-API glue. -
Query DSL: subsumption-based simplification, string-match operators, and
resolved.*virtual fields. -
Materialized views with rollups and tier-served aggregation; a new
AggregationFold
and extended aggregation operators. -
Concurrency-aware occupancy aggregates
busy,concurrency,utilizationand
activemeasure wall-clock busy time and parallelism instead of double-counting
overlapping durations. They are field-less (al...
v0.0.12
What's Changed
- fix(ci): fix ci errors and perf by @rayandrew in #89
- perf(valgrind): shard the C++ run, fix server test timeout on arm64 by @rayandrew in #91
- fix(dfanalyzer): fix deadlock in dfanalyzer multi-processes by @rayandrew in #86
- feat: interactive web trace viewer for dftracer_server by @rayandrew in #92
- feat(web): timeline node grouping by @rayandrew in #93
- feat(viz): app-span events and timelapse axis for multi-run traces by @rayandrew in #94
- fix(comparator): fix wrong counting files by @rayandrew in #87
- feat(indexer): detect and rebuild stale indexes on source change by @rayandrew in #95
- fix: consolidate stale-index rebuilds across consumers and fix multi-run breaks by @rayandrew in #96
- chore: bump version 0.0.12 by @rayandrew in #97
Full Changelog: v0.0.11...v0.0.12
Release v0.0.11
What's Changed
- resolve ObjectPool SIGSEGV on AArch64 + harden CI wheel/coverage pipeline by @rayandrew in #78
- feat(comparator): add dlio preset and consolidate constants by @rayandrew in #82
- Add Valgrind memory checking (C++, Python, MPI) and fix the bugs it found by @rayandrew in #79
- chore!: QoL, error handling, dedup + splits, and concurrency fixes by @rayandrew in #84
- chore: bump version to 0.0.11 by @rayandrew in #85
Full Changelog: v0.0.10...v0.0.11
Release v0.0.10
What's Changed
- dfanalyzer parity: add distributed HLM filtering and time bucketing by @rayandrew in #76
- chore: bump version to 0.0.10 by @rayandrew in #77
Full Changelog: v0.0.9...v0.0.10
Release v0.0.9
What's Changed
- fix(wheel): update tooling version by @rayandrew in #75
Full Changelog: v0.0.8...v0.0.9
Release v0.0.8
What's Changed
- add portable dependencies wheel support by @rayandrew in #73
Full Changelog: V0.0.8...v0.0.8
Release V0.0.8
What's Changed
- chore(utils): add portable to_chars_double fallback for macOS by @rayandrew in #72
Full Changelog: v0.0.7...V0.0.8
Release v0.0.7
What's Changed
- chore(ci): update ccache key to use matrix.os instead of runner.os by @rayandrew in #71
Full Changelog: v0.0.6...v0.0.7
Release v0.0.6
What's Changed
- update cibuildwheel action version by @rayandrew in #27
- feature: composable utilities by @rayandrew in #29
- add watchdog CLI args by @rayandrew in #33
- Consolidated changes to support call_tree API by @amarathe84 in #34
- feat!(pipeline): C++20 coroutine-driven tasks executor and scheduler by @rayandrew in #36
- fix(compile): fix compilation warnings by @rayandrew in #38
- Quality-of-Life Features: Statistics, Views, Manifest Index, and Trace Reorganization by @rayandrew in #42
- Dftracer trace replay functionality with updated interfaces to the call-tree functionality by @amarathe84 in #35
- feat!: pipeline and qol improvement by @rayandrew in #44
- fix(channel): fix channel lifetime management by @rayandrew in #45
- feat: improve line counting by @rayandrew in #49
- feat!: unified sidecar indexes (checkpointing and bloom filter) by @rayandrew in #50
- feat!(runtime): async submit with TaskHandle, Python callable support, and streaming iterators by @rayandrew in #51
- feat!: add Arrow data interchange and utility Python bindings via
nanoarrowby @rayandrew in #52 - feat!(utilities): query DSL for filtering traces by @rayandrew in #53
- chore: move to coveralls.io for coverage collections by @rayandrew in #54
- feat(utils): sequential fallback for small files by @rayandrew in #55
- feat(comparator): add pairwise traces comparator by @rayandrew in #57
- feat(perf): comprehensive memory and throughput optimizations by @rayandrew in #60
- fix(indexing): prevent double counting when index exists by @rayandrew in #61
- fix(gzip): handle concatenated gzip member boundaries in indexer and reader by @rayandrew in #62
- feat(aggregator): support profile/system counter aggregation and custom metric Arrow output by @rayandrew in #63
- feat(rocksdb): migrate SQLite indexing to RocksDB by @rayandrew in #64
- feat(perf): performance improvements for parallel reading, indexing, and aggregation by @rayandrew in #65
- feat(dlio): DLIO config generator by @rayandrew in #66
- feat(aggregator): offset metrics, per-event-name system metrics, and time-bucket persistence by @rayandrew in #68
- feat(dfanalyzer): integration of bridge module, exact-interval reindex, cat normalization by @rayandrew in #69
- chore: add version bump script and update version to 1.0.0 by @rayandrew in #70
New Contributors
- @amarathe84 made their first contribution in #34
Full Changelog: v0.0.5...v0.0.6
Release v0.0.5
What's Changed
- fix performance issue on continued reading in gzip stream by @rayandrew in #26
- fix promise fullfillment and exception handling in pipeline executors by @rayandrew in #24
- set default checkpoint size from defined constant by @rayandrew in #25
- fix wheel not linking libraries by @rayandrew in #23
Full Changelog: v0.0.4...v0.0.5