0.4.0
Pre-release
Pre-release
What's Changed
- Sub-file chunk loading bounded by max_batch_bytes by @gitbisector in #90
- Add Windows ARM64 wheel builds by @chinazhangchao in #97
- Fit planner: deterministic per-file chunk budgets (device_memory_budget) by @gitbisector in #91
- Improve DirectStorage fallback and CUDA runtime discovery by @chinazhangchao in #98
- fix(tests): keep the 3FS mock reader on host memory by @gitbisector in #100
- fix: skip cuFileDriverOpen() when /dev/nvidia-fs0 is missing to prevent fd 0 corruption by @yinjuncheng in #102
- Account for yield clones in fit planner by @takeshi-yoshimura in #103
- Reuse chunk budgets for stable buffer allocations by @takeshi-yoshimura in #104
- Fix GDS capability checks before library initialization by @takeshi-yoshimura in #105
- Clamp queue_size to the device memory budget instead of failing by @takeshi-yoshimura in #106
- Per-tensor residency predicate for the fit planner by @gitbisector in #108
- Prepare the 0.4.0 release by @takeshi-yoshimura in #107
New Contributors
- @chinazhangchao made their first contribution in #97
- @yinjuncheng made their first contribution in #102
Full Changelog: 0.3.3...0.4.0