Skip to content

v 1.0.0

Latest

Choose a tag to compare

@mkurman mkurman released this 25 Feb 14:02

Optimize ReasonFlow for improved performance and code structure.

This release introduces significant performance optimizations and code
restructuring the ReasonFlow framework, making it more efficient and
maintainable.

Performance Improvements:

  • Optimized embedding generation and noise application

    • Reduced redundant copies of input embeddings
    • More efficient noise application using vectorized operations
    • Better memory management with tensor preallocation
  • Improved token selection and hidden state processing

    • Implemented weighted averaging for multiple best thinkers
    • Enhanced token selection with better numerical stability
    • Added fast paths for common cases like single-thinker
    • Reduced unnecessary CUDA synchronization points
  • Enhanced KV-cache handling

    • Created a specialized path for the single best thinker
    • Reduced frequency of CUDA memory cleanup
    • Explicit device and type management for fewer transfers
  • Optimized thinker selection algorithm

    • Fast paths for the single thinker and zero diversity weight cases
    • Improved similarity calculation with batched matmul
    • Better numerical stability in normalization
    • Reduced redundant calculations
  • Improved uncertainty measurement

    • Conditional use of log_softmax for better numerical stability
    • Optimized handling for large batch sizes
    • More efficient certainty calculation
    • Reduced memory footprint during computation
  • Enhanced hooks for hidden states

    • Added LRU caching for module name lookups
    • Kept tensors on the original device when possible
    • Extracted only needed portions of hidden states
    • Added robust error handling with fast recovery

Code Structure Improvements:

  • Split monolithic implementation into modular components
  • Created a dedicated generation directory with specialized modules
  • Better isolation of functionality for easier maintenance
  • Improved documentation throughout the codebase
  • Removed unused and redundant code (like apply_temperature_based_sampling)
  • Updated README with performance considerations and architectural details

These changes maintain the same functionality while significantly improving
performance, especially for larger models, more thinkers, and longer
generation sequences.

🚀 Performance tests show up to 30% speedup on generation tasks with
multiple thinkers, with the greatest gains seen in high-complexity scenarios.