Repository navigation
Releases: Xiaooolong/vev
Releases · Xiaooolong/vev
Release list
Vev 0.1.1
- The questions of a request run as one batch instead of one forward pass each. On one H800, 100 questions about a
short text take 0.23 s instead of 3.75 s withvev-4b; images go through the image processor once per request. - Requests with one question give bit-identical probabilities to 0.1.0. With several questions, probabilities move by
up to a few hundredths (bf16 rounding); spec §6 and the conformance tolerance are updated with the measured values. flash-linear-attention0.5.2 is a dependency on Linux, as in the evaluation environment; the server logs which
DeltaNet kernels it uses.torchvisionis a dependency.vev serve --revision; startup fails fast without CUDA; 500 responses no longer include the exception text.evals.run --one-question-per-request.
The weights are unchanged from 0.1.0 (collection): vev-4b · vev-4b-lora · vev-9b · vev-9b-lora.
Vev 0.1.0
First public release, a research preview.
vev serve: HTTP server with the/v1/systemonerequest and response shapes. Text, JSON and images instate;
noul,choiceandscorequestions.- Models:
vev-4bandvev-9b, LoRA fine-tunes of Qwen3.5-4B and Qwen3.5-9B, released as merged weights and as
adapters. Weights are CC BY-NC 4.0. - Conformance suite that runs against any
/v1/systemoneimplementation. - Evaluation harness, dataset converters and the training code used for the release.
Weights and data on Hugging Face (collection):
vev-4b · vev-4b-lora · vev-9b · vev-9b-lora · vev-anchor-v3 (training anchor data).