Releases
v0.2.5
Compare
Sorry, something went wrong.
No results found
Added
Multi-Model Proxy Server (Experimental) : Enabling multiple LLMs through a single unified API endpoint
Single OpenAI-compatible endpoint for all models
Request routing based on model name
Save and reuse proxy configurations
Dynamic Model Management : Add or remove models at runtime without restarting the proxy
Live model registration and unregistration
Pre-registration with verification lifecycle
Graceful handling of model failures without affecting other models
Model state tracking (pending, running, sleeping, stopped)
Model Sleep/Wake for GPU Memory Management : Efficient GPU resource distribution
Sleep Level 1: CPU offload for faster wake-up
Sleep Level 2: Full memory discard for maximum savings
Real-time memory usage tracking and reporting
Models maintain their ports while sleeping
Test Coverage : Added comprehensive tests for multi-model proxy and model registry
Changed
Improved error handling with detailed logs when PyTorch is not installed
Better server cleanup and process management
Fixed
UI navigation improvements and minor display fixes
You can’t perform that action at this time.