You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Prompt cache: two sushi processes on the same model no longer share one SSD cache folder, which could return
another session's answer; a restart on a nearly full disk keeps the cache, and loading a second model no longer
deletes the first one's entries.
Crash fixes: an empty /v1/completions prompt, deeply nested JSON and oversized WebSocket messages are refused
with a 400 instead of stopping the server.
GLM-5.3-Flash: after a client disconnects, the next request on a streamed GLM no longer fails; sushi pull
now fetches the DFlash2 assistant, and deleting its BF16 source keeps DFlash2 on.
Tools and structured output: earlier tool calls with empty or non-object arguments render in the model's own
format, anyOf/oneOf/$ref parameters get their real types, and JSON-schema output stops cleanly on numbers and
bounded arrays; logit_bias: null is accepted again.
CLI: sushi launch quotes model names, sushi serve without --model honours the sampling flags, update
no longer logs the API key, and sushi run accepts long pasted lines and filters terminal escapes from model output.
mlx-serve: the guest manifest now lists GLM-5.3-Flash.