Skip to content

v0.2.0

@lrozewicz lrozewicz tagged this 06 Aug 07:28
~14.3 tok/s decode with DSpark k=2 (was ~9.8), quality semantics unchanged:
the FP4 confidence gate was firing on ~38% of prose steps because the V2
gate hook evaluated verify rows the sampler discards. Prefill ~753 tok/s on
a 23.7K prompt (was ~500).

Every <think> block now reaches the `reasoning` field, including blocks the
model reopens mid-answer, which used to leak a raw </think> into agent
clients in both thinking and non-thinking mode.

Serving images vllm-moet-gb10:v028 through v030. See CHANGELOG.md for the
full list and docs/CHANGES-VS-UPSTREAM.md for the delta against
kacper-daftcode/vLLM-Moet.
Assets 2
Loading