Skip to content

Bug Fixes / Improvements / Optimizations

Choose a tag to compare

@FearL0rd FearL0rd released this 03 Feb 21:36
· 27 commits to main since this release
8a3085d

Critical Bug Fixes:
Memory Leak: weakref.finalize now uses weakref.ref(target_model) instead of direct reference
Race Condition: Worker threads now return exceptions instead of raising them, allowing proper aggregation and cleanup
Batch Validation: get_batch_size() now validates consistent batch dimensions across tensor lists
Double Device Transfer: Removed .to(device) from split_kwargs(), now only happens in worker()
Attribute Handling: Cleaner separation of in_features vs in_channels in layer cloning
Improvements Added:
CUDA Streams: Each GPU now uses its own CUDA stream for true parallel execution
Auto VRAM Balance: New auto_vram_balance parameter adjusts splits based on available VRAM (70% user preference / 30% availability)
Gradient Checkpointing: Automatically disabled on replicas to save VRAM
Accelerate Hooks: Clears _hf_hook and hooks attributes for compatibility with accelerate offloading
Enhanced Cache Clearing: Added rope_cache and freqs_cis_cache to FLUX cleanup
Thread Safety: Proper executor context management and result validation
Performance Optimizations:
Non-blocking transfers where safe
Synchronization only on CUDA/XPU devices
Batch size auto-adjustment prevents OOM on uneven GPU setups