Release 0.6.2
Parallelize gasbench data loading & fix memory leaks
Problem
Image and video benchmarks run unacceptably slowly when data was coming from NAS (not noticeable on local setups), and also occasionally OOM deep into runs. T
Root causes identified:
-
Sequential disk I/O from network volumes — The
DatasetIteratorreads each image/video file one-by-one in the producer thread. Each read from a NAS incurs network latency, paid serially N times. -
"Drain-all-futures" stall —
PrefetchPipelineaccumulatednum_workers * 2futures then blocked on ALL of them (for future in futures: future.result()). The pipeline stalls on the slowest task even when other workers are idle. -
Only 3 worker threads — With I/O-bound work (network volume reads + PIL decode), 3 threads underutilize available concurrency.
-
Memory leak from large images — Datasets with very large source images (100+ megapixels observed in logs) cause multi-GB memory spikes because image bytes are held in multiple places simultaneously: the sample dict, the result dict, and the batch queue. No explicit cleanup of PIL Image objects in multi-threaded workers.
Changes
gasbench/src/gasbench/dataset/iterator.py
- Added
lazy_read: boolparameter toDatasetIterator - When
True, image samples yield{"image_path": ...}instead of reading file bytes; video samples yield{"video_path": ...}for file-based videos (frame directories are already lazy) - Iterating the dataset becomes near-instant (path collection only, no I/O)
gasbench/src/gasbench/benchmarks/image_bench.py
- Rewrote
PrefetchPipelinewith three fixes:- Parallel I/O: New
_read_and_preprocess()does file read + PIL decode + augmentation as a single unit inside worker threads — 8 threads read from the network volume concurrently - Bounded sliding window: Uses
wait(FIRST_COMPLETED)withmax_in_flight = num_workers * 4 = 32instead of submit-all. Prevents unbounded memory growth from completed-but-unconsumed futures - Sample metadata stripping: Drops heavy keys (
image,image_bytes,image_path) from result dicts immediately after preprocessing — tracker only needs metadata fields
- Parallel I/O: New
- Default
num_workersincreased from 3 → 8 DatasetIteratorcreated withlazy_read=Trueexecutor.shutdown()now usescancel_futures=Truefor clean teardown
gasbench/src/gasbench/benchmarks/video_bench.py
- Same rewrite applied to
VideoPrefetchPipeline - Default
num_workersincreased from 3 → 4 (fewer than image due to heavier per-sample memory) max_in_flight = num_workers * 3 = 12(tighter bound for video frames)- Strips
video_bytesandvideo_pathfrom result dicts
gasbench/src/gasbench/processing/media.py
- Added explicit
image.close()inprocess_image_sample()after extracting the numpy array — prevents PIL Image objects from lingering in multi-threaded workers
Expected impact
| Metric | Before | After |
|---|---|---|
| Image I/O concurrency | 1 (serial) | 8 threads |
| Video I/O concurrency | 1 (serial) | 4 threads |
| Pipeline stall pattern | Drain all 6, block on slowest | FIRST_COMPLETED, no stalls |
| Peak in-flight samples (image) | 6 | 32 (bounded) |
| Peak in-flight samples (video) | 6 | 12 (bounded) |
| Image bytes in result dict | Held until tracker consumes | Stripped immediately |
| PIL Image cleanup | GC-dependent | Explicit .close() |
| Est. image benchmark time | ~5 hours (52 datasets) | ~1-2 hours |