Skip to content

Release 0.6.2

Choose a tag to compare

@dylanuys dylanuys released this 05 Apr 17:51
· 252 commits to main since this release
d5d1369

Parallelize gasbench data loading & fix memory leaks

Problem

Image and video benchmarks run unacceptably slowly when data was coming from NAS (not noticeable on local setups), and also occasionally OOM deep into runs. T

Root causes identified:

  1. Sequential disk I/O from network volumes — The DatasetIterator reads each image/video file one-by-one in the producer thread. Each read from a NAS incurs network latency, paid serially N times.

  2. "Drain-all-futures" stall — PrefetchPipeline accumulated num_workers * 2 futures then blocked on ALL of them (for future in futures: future.result()). The pipeline stalls on the slowest task even when other workers are idle.

  3. Only 3 worker threads — With I/O-bound work (network volume reads + PIL decode), 3 threads underutilize available concurrency.

  4. Memory leak from large images — Datasets with very large source images (100+ megapixels observed in logs) cause multi-GB memory spikes because image bytes are held in multiple places simultaneously: the sample dict, the result dict, and the batch queue. No explicit cleanup of PIL Image objects in multi-threaded workers.

Changes

gasbench/src/gasbench/dataset/iterator.py

  • Added lazy_read: bool parameter to DatasetIterator
  • When True, image samples yield {"image_path": ...} instead of reading file bytes; video samples yield {"video_path": ...} for file-based videos (frame directories are already lazy)
  • Iterating the dataset becomes near-instant (path collection only, no I/O)

gasbench/src/gasbench/benchmarks/image_bench.py

  • Rewrote PrefetchPipeline with three fixes:
    • Parallel I/O: New _read_and_preprocess() does file read + PIL decode + augmentation as a single unit inside worker threads — 8 threads read from the network volume concurrently
    • Bounded sliding window: Uses wait(FIRST_COMPLETED) with max_in_flight = num_workers * 4 = 32 instead of submit-all. Prevents unbounded memory growth from completed-but-unconsumed futures
    • Sample metadata stripping: Drops heavy keys (image, image_bytes, image_path) from result dicts immediately after preprocessing — tracker only needs metadata fields
  • Default num_workers increased from 3 → 8
  • DatasetIterator created with lazy_read=True
  • executor.shutdown() now uses cancel_futures=True for clean teardown

gasbench/src/gasbench/benchmarks/video_bench.py

  • Same rewrite applied to VideoPrefetchPipeline
  • Default num_workers increased from 3 → 4 (fewer than image due to heavier per-sample memory)
  • max_in_flight = num_workers * 3 = 12 (tighter bound for video frames)
  • Strips video_bytes and video_path from result dicts

gasbench/src/gasbench/processing/media.py

  • Added explicit image.close() in process_image_sample() after extracting the numpy array — prevents PIL Image objects from lingering in multi-threaded workers

Expected impact

Metric Before After
Image I/O concurrency 1 (serial) 8 threads
Video I/O concurrency 1 (serial) 4 threads
Pipeline stall pattern Drain all 6, block on slowest FIRST_COMPLETED, no stalls
Peak in-flight samples (image) 6 32 (bounded)
Peak in-flight samples (video) 6 12 (bounded)
Image bytes in result dict Held until tracker consumes Stripped immediately
PIL Image cleanup GC-dependent Explicit .close()
Est. image benchmark time ~5 hours (52 datasets) ~1-2 hours