v0.6.2
Added three dense text-embedding models: mixedbread-ai/mxbai-embed-large-v1, Snowflake/snowflake-arctic-embed-l-v2.0, and nomic-ai/modernbert-embed-base.
Self-hosted startup is more reliable: the configuration service can serve health checks while NATS connects, single-profile bundles can scale from zero for requests without an explicit GPU selection, and CUDA images include the build tool needed for SGLang's first-use kernels. Qwen3-VL embedding models also accept and validate their configured output dimension.