Skip to content

Releases: llnl-asr/dftracer-utils

v0.0.13

Choose a tag to compare

@rayandrew rayandrew released this 30 Sep 18:30

Added

  • A columnar DataFrame / Series / LazyFrame engine that is a drop-in for
    pandas and polars: import dftracer.utils.pandas as pd (or .polars as pl)
    and most code runs unchanged on the SIMD kernels. The pandas surface covers
    loc / iloc / at / iat and set_index (the index is a named column,
    copy-on-write assignment), groupby objects with the agg forms, column
    selection (groupby(k)["v"]), a Series key, group-wise transforms
    (cumsum, shift, rank, head, nth, ffill, bfill, rolling,
    expanding, ewm, take, sample, resample), apply / map compiled
    into the engine (Python per row only as a last resort, with a warning), the
    str and dt accessors, tz_localize / tz_convert, merge, nlargest,
    value_counts, mode, compare, pivot_table and describe; the polars
    spellings sit alongside (select(Expr), with_columns, over, the str
    namespace). Series masks combine with &, | and ~; == / != return
    a mask (a Series is unhashable, as in pandas and polars).
  • A hash join on DataFrame, LazyFrame, the C ABI and Python (join /
    merge; inner, left, right, outer, semi, anti and cross). A plan's join sends its
    build keys to the scan, which prunes the chunks that cannot match.
  • Plans (LazyFrame): explain(), schema() and output_schema() without
    running; memory_budget / auto_spill bounding every breaker; a source of
    your own (Source / dftu_source_vt, registered by name); a plugin's own
    plan step (dftu_node_register, LazyFrame.op); a frame_op step for
    every registry table op (unnest, partition_id, compare_agg, window,
    gap_fill, asof, interval, concat, union, pivot and to_dummies).
  • Aggregates: prod, cumprod, exact group median / quantile,
    unique(subset), a per-column reduce, a whole-frame group_by(),
    first / last exact across a parallel merge; group keys of any type
    (Binary, Float16, null keys as their own group with dropna=False).
  • Plugin ABI: a plugin transforms the batch every later plugin receives
    (transform), reports and releases what it holds against the memory
    budget (bytes / reclaim), and a plugin node or slice under a plan is
    measured by the same budget; abi_version is a hash of the header, so a
    plugin built against another version is refused at load.
  • benchmarks/dataframe_vs_pandas_polars.py: the engine against pandas,
    polars and DuckDB on the same Arrow tables, eagerly and as a plan, with a
    correctness check of every result against ours, the cores each engine kept
    busy and --memory for the peak resident set per op.
  • The trace View (C++) and TraceViewer (Python) are a LazyFrame over a
    trace scan: generic ops (filter, select, sort, head, join, ...) chain on
    it, and the scan absorbs what it can at plan time (filters, projections,
    group keys, a trailing row window). Trace terminals (call_tree,
    flamegraph, containment, flamegraph_partial, aggregate_partial,
    sink_json, sink_trace, materialize) are lazy results, and
    collect_all runs several of them over one shared scan. Python adds lazy
    [], LazyScalar reductions and LazyResult.
  • Python Indexer(bloom=BloomConfig(fields=...)) names extra args fields to
    index.
  • The index covers every flat args field by default: numbers by per-chunk
    min/max, strings by a per-chunk bloom up to 256 distinct values, so a
    filter on any arg prunes chunks (equality prunes on min/max as well).
    dftracer_index --no-auto-dimensions opts out. On a 2M-event trace the
    index is 1.2 MB, against 8.7 MB for the previous named defaults. An
    existing index rebuilds once.

Changed

  • Precompiled headers are off by default (DFTRACER_UTILS_ENABLE_PCH). GCC's
    .gch cannot be cached by ccache and evicted the rest of the cache; pass
    -DDFTRACER_UTILS_ENABLE_PCH=ON to keep them for local builds.

  • CI: every merge into develop publishes a <tag>.postN.dev0 prerelease to
    PyPI; wheel and Valgrind jobs run sharded (one job per Python version, three
    Python and six C++ Valgrind shards); push workflows run on main and
    develop only, so a pull request no longer runs twice.

  • The dataframe engine is measured (10M rows, Apple M4 Pro): ahead of pandas
    on every benchmark row, of polars on every row but two at the noise floor,
    of DuckDB on every row but one within a millisecond of it. The group-by
    runs a plain loop over batches of rows with a direct table for dense
    integer keys, a word table for string keys and, with many groups on a
    string key, a scatter into per-thread partitions; the join uses a
    direct-address table and 32-bit index lists; the sort is a sample sort;
    filters, comparisons, string predicates, casts, gathers, rolling windows,
    // / % / ** and the dictionary encoder run in parallel; a quantile
    reads its column in place and sorts one bucket.

  • A plan over a resident frame runs whole-column for every op (an op with no
    eager form runs its own cursor over the frame as one morsel); it no longer
    streams through a spool or spills to disk.

  • Memory detection reads the free and inactive pages on macOS (the auto
    budget assumed 1 GB there).

  • Series.rolling(...).mean() and the other windows write their output in
    place and run a chunk per thread.

  • Breaking (C++): View terminals follow the LazyFrame names:
    export_json / export_trace / export_counters became sink_json /
    sink_trace / sink_counters, merge_partials_to_table became
    merge_partials, limit / offset became head / slice, occ_cell
    became resolution and schema() became column_info(). collect()
    returns the DataFrame; lazy() returns the plan. call_tree,
    flamegraph, containment and the partials return plans to collect.
    Several outputs over one scan use collect_all or TraceSession;
    caller folds attach with View::branch.

  • Breaking (Python): TraceViewer is a LazyFrame; AggregatedTraceViewer
    and SessionView are gone. DaskTraceViewer.occ_cell became resolution,
    and its offset / limit trim the merged result instead of each shard's.

  • The server, dftracer_view, dftracer_run, the statistics tools and the
    C ABI run on the same View; the separate builder API is internal.

  • A shard set aggregates its shards concurrently, as many at once as the
    spill budget gives each at least 64 MB, and collect_all splits each
    plan's spill budget across the plans it runs together, so concurrent work
    stays within the budget.

Removed

  • Breaking: the previous plugin ABI. A plugin built against it does not
    load; rebuild against dftracer/utils/plugins/abi/plugin.h.
  • Breaking (C++): AggregatedView, ViewSession's public constructor,
    View::join and trace/views/result_batch.h (collect_batch); use
    View plans, View::branch and LazyFrame::join.

Fixed

  • Prerelease Linux wheels were versioned .post1.devN because the manylinux
    container did not receive the computed version; the version is passed in.

  • A sketch quantile (pct in a plan, DDSketch) returned -inf once a
    bucket held more than 65535 values; buckets are 32-bit now.

  • Column-column arithmetic dropped nulls; a Bool column was gathered by
    byte instead of by bit; group_by returned groups out of first-seen order
    after a parallel run; a plugin node was answered asynchronously when the
    data was resident.

  • df.groupby(series) raised a SystemError: the frame's in test left a
    pending error for a non-string key.

  • A scan of an unindexed multi-member trace read the lines at each member
    boundary twice.

  • Index extra args fields were dropped by the fold-based build, so
    dftracer_index --dimensions did nothing; nested fields were dropped in the
    first-touch build; merged chunk stats ordered numeric min/max as text; and
    a float literal (pid == 1.0) probed the bloom as "1.000000" and pruned
    matching chunks.

  • Session containment branches ignored phase(), so phase(Events) kept
    aggregated records; session aggregations with string-arg predicates or
    transformed keys wrote rollups that later reads served wrongly.

  • A blocking get() on a finished coroutine task could miss its result;
    LazyFrame::collect() on a temporary plan read the freed plan.

  • A plan's export sink flushes when the export ends.

  • Data races: libdeflate chose its kernels on first use from several threads
    at once, and the reader's member decode cache read an entry's ready flag
    outside the lock that wrote it.

  • An expression the column type cannot take (a string column compared with a
    number) produced a null column that crashed a later group_by; it now
    raises an error.

  • Interactive web trace viewer gains a counter timeline track (with malformed-value
    handling), per-counter pid/tid breakdown, bounded-density serving, and active-time
    statistics.

  • Rectangle selection in the viewer scopes analysis to a time range, lanes, and rows;
    aggregated (ph=3) events are visualized with uniform extrapolation and labeled
    "aggregated" in tooltips.

  • DLIO event category and name can be remapped through an event-map file.

  • Python: DaskTraceViewer is exported from the dask module; new TraceViewer/View
    bindings expose the query DSL and portable C-API glue.

  • Query DSL: subsumption-based simplification, string-match operators, and
    resolved.* virtual fields.

  • Materialized views with rollups and tier-served aggregation; a new AggregationFold
    and extended aggregation operators.

  • Concurrency-aware occupancy aggregates busy, concurrency, utilization and
    active measure wall-clock busy time and parallelism instead of double-counting
    overlapping durations. They are field-less (al...

Read more

v0.0.12

Choose a tag to compare

@rayandrew rayandrew released this 14 Jul 13:11
v0.0.12

What's Changed

  • fix(ci): fix ci errors and perf by @rayandrew in #89
  • perf(valgrind): shard the C++ run, fix server test timeout on arm64 by @rayandrew in #91
  • fix(dfanalyzer): fix deadlock in dfanalyzer multi-processes by @rayandrew in #86
  • feat: interactive web trace viewer for dftracer_server by @rayandrew in #92
  • feat(web): timeline node grouping by @rayandrew in #93
  • feat(viz): app-span events and timelapse axis for multi-run traces by @rayandrew in #94
  • fix(comparator): fix wrong counting files by @rayandrew in #87
  • feat(indexer): detect and rebuild stale indexes on source change by @rayandrew in #95
  • fix: consolidate stale-index rebuilds across consumers and fix multi-run breaks by @rayandrew in #96
  • chore: bump version 0.0.12 by @rayandrew in #97

Full Changelog: v0.0.11...v0.0.12

Release v0.0.11

Choose a tag to compare

@hariharan-devarajan hariharan-devarajan released this 06 Jul 02:23

What's Changed

  • resolve ObjectPool SIGSEGV on AArch64 + harden CI wheel/coverage pipeline by @rayandrew in #78
  • feat(comparator): add dlio preset and consolidate constants by @rayandrew in #82
  • Add Valgrind memory checking (C++, Python, MPI) and fix the bugs it found by @rayandrew in #79
  • chore!: QoL, error handling, dedup + splits, and concurrency fixes by @rayandrew in #84
  • chore: bump version to 0.0.11 by @rayandrew in #85

Full Changelog: v0.0.10...v0.0.11

Release v0.0.10

Choose a tag to compare

@hariharan-devarajan hariharan-devarajan released this 09 Jun 05:59

What's Changed

  • dfanalyzer parity: add distributed HLM filtering and time bucketing by @rayandrew in #76
  • chore: bump version to 0.0.10 by @rayandrew in #77

Full Changelog: v0.0.9...v0.0.10

Release v0.0.9

Choose a tag to compare

@hariharan-devarajan hariharan-devarajan released this 07 Jun 22:39

What's Changed

Full Changelog: v0.0.8...v0.0.9

Release v0.0.8

Choose a tag to compare

@hariharan-devarajan hariharan-devarajan released this 07 Jun 01:50

What's Changed

Full Changelog: V0.0.8...v0.0.8

Release V0.0.8

Choose a tag to compare

@hariharan-devarajan hariharan-devarajan released this 04 Jun 05:09

What's Changed

  • chore(utils): add portable to_chars_double fallback for macOS by @rayandrew in #72

Full Changelog: v0.0.7...V0.0.8

Release v0.0.7

Choose a tag to compare

@hariharan-devarajan hariharan-devarajan released this 23 May 03:08

What's Changed

  • chore(ci): update ccache key to use matrix.os instead of runner.os by @rayandrew in #71

Full Changelog: v0.0.6...v0.0.7

Release v0.0.6

Choose a tag to compare

@hariharan-devarajan hariharan-devarajan released this 22 May 03:43

What's Changed

  • update cibuildwheel action version by @rayandrew in #27
  • feature: composable utilities by @rayandrew in #29
  • add watchdog CLI args by @rayandrew in #33
  • Consolidated changes to support call_tree API by @amarathe84 in #34
  • feat!(pipeline): C++20 coroutine-driven tasks executor and scheduler by @rayandrew in #36
  • fix(compile): fix compilation warnings by @rayandrew in #38
  • Quality-of-Life Features: Statistics, Views, Manifest Index, and Trace Reorganization by @rayandrew in #42
  • Dftracer trace replay functionality with updated interfaces to the call-tree functionality by @amarathe84 in #35
  • feat!: pipeline and qol improvement by @rayandrew in #44
  • fix(channel): fix channel lifetime management by @rayandrew in #45
  • feat: improve line counting by @rayandrew in #49
  • feat!: unified sidecar indexes (checkpointing and bloom filter) by @rayandrew in #50
  • feat!(runtime): async submit with TaskHandle, Python callable support, and streaming iterators by @rayandrew in #51
  • feat!: add Arrow data interchange and utility Python bindings via nanoarrow by @rayandrew in #52
  • feat!(utilities): query DSL for filtering traces by @rayandrew in #53
  • chore: move to coveralls.io for coverage collections by @rayandrew in #54
  • feat(utils): sequential fallback for small files by @rayandrew in #55
  • feat(comparator): add pairwise traces comparator by @rayandrew in #57
  • feat(perf): comprehensive memory and throughput optimizations by @rayandrew in #60
  • fix(indexing): prevent double counting when index exists by @rayandrew in #61
  • fix(gzip): handle concatenated gzip member boundaries in indexer and reader by @rayandrew in #62
  • feat(aggregator): support profile/system counter aggregation and custom metric Arrow output by @rayandrew in #63
  • feat(rocksdb): migrate SQLite indexing to RocksDB by @rayandrew in #64
  • feat(perf): performance improvements for parallel reading, indexing, and aggregation by @rayandrew in #65
  • feat(dlio): DLIO config generator by @rayandrew in #66
  • feat(aggregator): offset metrics, per-event-name system metrics, and time-bucket persistence by @rayandrew in #68
  • feat(dfanalyzer): integration of bridge module, exact-interval reindex, cat normalization by @rayandrew in #69
  • chore: add version bump script and update version to 1.0.0 by @rayandrew in #70

New Contributors

Full Changelog: v0.0.5...v0.0.6

Release v0.0.5

Choose a tag to compare

@hariharan-devarajan hariharan-devarajan released this 22 Oct 03:39

What's Changed

  • fix performance issue on continued reading in gzip stream by @rayandrew in #26
  • fix promise fullfillment and exception handling in pipeline executors by @rayandrew in #24
  • set default checkpoint size from defined constant by @rayandrew in #25
  • fix wheel not linking libraries by @rayandrew in #23

Full Changelog: v0.0.4...v0.0.5