Improved routing with continuous QoS latency weighting and a configurable latency target in Settings.
Corrected the Cerebras GLM 4.7 free-tier context limit to 8,192 tokens.
Added documentation and regression tests for the new routing behavior.
Improved routing with continuous QoS latency weighting and a configurable latency target in Settings.
Corrected the Cerebras GLM 4.7 free-tier context limit to 8,192 tokens.
Added documentation and regression tests for the new routing behavior.