Releases
v0.10.0
Compare
Sorry, something went wrong.
No results found
0.10.0 (2026-09-29)
⚠ BREAKING CHANGES
metal-agent: serve engines on loopback behind an authenticated ingress (#1945 )
Features
controller: front Metal InferenceServices with an authenticated relay (#1944 ) (37be11b )
metal-agent: serve engines on loopback behind an authenticated ingress (#1945 ) (cf55427 )
metal-agent: tensorfold runtime for MLX speculative decoding (#1923 ) (8c9820c )
Bug Fixes
api: let InferenceService select Metal runtimes and fall back to the agent flag (#1912 ) (19191b4 )
metal-agent: enforce a typed allowlist on extraArgs (#1941 ) (c09a4be )
metal-agent: enforce allowed roots on local model paths (#1939 ) (bcdd27f )
metal-agent: fail fast when llama-server exits during startup (#1917 ) (23f1fc6 )
metal-agent: fail fast when mlx-server exits during startup (#1929 ) (c55f535 )
metal-agent: fail fast when vllm-swift exits during startup (#1928 ) (46e4e96 )
metal-agent: load local-path GGUF sources in place instead of downloading (#1920 ) (6c045b4 )
metal-agent: never overwrite or delete a Service or EndpointSlice it does not own (#1932 ) (799194a )
metal-agent: size file:// model sources in the memory estimate (#1930 ) (39b1ef6 )
metal-agent: spell mlock as --load-mode on llama.cpp 0.5.0+ (#1916 ) (f760639 )
metal: stop routing to a Metal endpoint nothing is serving (#1921 ) (75c792f )
You can’t perform that action at this time.