entail 1.2.0
Released 2026-09-26. The sites the first pre-registered replay found unread (a LoRA adapter's settings file, a
request's template settings at the reasoning parser, the prefix-cache key and the beam reorder, Triton kernel
launches, rotary pairing), each written from the real bug and measured on it, and a second pre-registered replay
with the vocabulary frozen at this version.
Added
- A LoRA adapter's
adapter_config.jsonas a declaration file (adapter_config_contract.py,
data/adapter_config_keys.json): every PEFT key, which ones each consumer reads (PEFT for transformers and
diffusers; vLLM 0.30's PEFTHelper; SGLang 0.5.20's LoRAConfig, with code lines), and one rule: a key declared with
a value that changes how the weights apply, that the consumer does not read, isbrokenat that consumer's load
boundary, orresolvedwhere the consumer can carry it. Neutral values, keys the weights carry and training-time
keys decide nothing; a key the consumer refuses loudly passes with a note. Adapterssglang_lora(carries
use_rslorainto the adapter's scaling, computed in the core: sglang#40835 served rsLoRA adapters 4-8x too weak)
andvllm_lora(reports what vLLM drops:rank_pattern,alpha_pattern,lora_bias, ...).entail checkon an
adapter folder decides it per engine. - A request's template settings at the reasoning parser (
request_contract.setting_names,
data/request_settings.json): a setting the template honoured under one name (enable_thinking) that the
parser reads under another (thinking) leaves the parser on its default (vllm#43728:content: null). The
table names, per vLLM version and reasoning parser, the names each parser reads (code lines); the vLLM serve
adapter wraps the parser class the server builds per request and hands the request's value to the parser under
a name it reads (resolved), or reports it. A parser that reads no name of the setting, or a setting the
template did not read either, decides nothing. - A store's key against the fields that shaped the item (
cache_key_contract.py,data/cache_key_fields.json):
vLLM's prefix-cache block hash keys a request by its tokens, embeddings digest, multimodal hashes, LoRA name and
cache salt, but not byprompt_is_token_ids(which positions take the embeddings), so two requests that differ
only in that mask share a key (vllm#56655, fix unmerged at 0.30.0). The adaptervllm_cache_keydecides at
Request.__init__and repairs by appending a digest of the block's mask to the hash's extra keys and remaking
the request's hashes (resolved). The same rule covers a permutation: transformers' beam search reorders the
cache under the names it knows (transformers_beam); a model whose cache lives under another name is reported
(5.12.1 reorderedpast_key_valuesonly: transformers#46612; 5.17.0 reorders every name). - What a Triton kernel is told about its tensors (
kernel_launch_contract.py, adaptertriton_launch, engine-
independent: one hook onJITFunction.run, so every@triton.jitkernel launched eagerly by any engine; kernels
Inductor generates for a compiled forward are not seen): a tensor strided in its innermost dimension handed to a
kernel that was not told that stride (no integer argument among its value parameters and stride-named
constexprs equals it) and whose parameters name no stride at all (stride,_s0,sxm,ld...) isbroken
(the kernel reads it as if contiguous); handed to a kernel that names strides but was not told this one, it is
unknown(said once). Each (kernel, stride pattern) is
decided once per process, and a kernel is looked at for its first eight strided patterns; compile-only warm-ups
are not launches. sglang#21843 (fused_gdn_gating read interleaved a/b) is the case the rule comes from; there
the kernel takes row strides, so the decision isunknownat the kernel's boundary. Rotary.pairing(vocabulary v8): how a rotary embedding pairs the dimensions it rotates,split(i with
i + d/2, the Llama convention) orinterleaved(2i with 2i+1, GPT-J's). Declared by a config key
(rope_interleave,rope_interleaved,is_neox_style) or, failing that, by the architecture's reference
implementation (data/rotary_pairing.json: transformers 5.17.0 configuration and modeling files with lines; GLM,
Cohere, Ernie 4.5, GPT-J and DeepSeek-V3 interleaved, Llama, Qwen and Gemma split).rotary_pairing_contract.py
decides the rotary modules vLLM built for the language model (is_neox_style, adaptervllm_pairing) against
the declaration -resolvedby setting the modules' convention,brokenwhere the model's MRoPE module
dispatches to a kernel that pairs split-wise whatever the layer says (vLLM's Triton MRoPE kernel before 0.27.0
with the custom op enabled: vllm#42016, #49290). Not compared: a multimodal model's vision tower (its own
reference), a DSA indexer (its own key), and modules that pair both ways in one language model (unknown, nothing
set). An architecture the table does not know, without a key, decides nothing.
Fixed
- The KV contract's
kv_neededrule breaks only when a sequence holds fewer slots than its tokens. Holding more
than one allocation unit over its tokens was alsobroken, and under speculative decoding an engine
legitimately does that: it reserves lookahead slots and keeps the blocks of the drafts it rejected (vLLM 0.30
with ngram speculation saidbrokenon a healthy run; vLLM 0.23 with extract_hidden_states likewise, seen in
the second replay). Fewer slots than tokens is the loss and is still reported. - The vLLM serve adapter's wrapper of
ParserManager.get_parserbinds its arguments by name and passes the rest
through: with a fixed signature it raisedTypeErroron vLLM 0.23.0 (whoseget_parsertakesis_harmony)
and the API server died - the one run entail broke in the second pre-registered replay. - A model loaded by hub id resolves to its cached snapshot folder again, so the Vocab and Stops checks decide
instead of saying "no local folder to read": huggingface_hub 1.32 refuses a cached snapshot that lacks files the
engine never fetched (.gitattributes, evaluation results), and every hub-id load on vLLM and transformers
was "could not be checked". The folder of the cachedconfig.jsonis used when the snapshot lookup refuses;
nothing is downloaded.
Changed
- The log and record files are kept open per process, one write and one flush per line, instead of being opened
and closed per line: on a 9P mount (a project under WSL's/mnt/c) the open and close cost 4.6 ms per line, and a
boundary that runs per request (the prefix-cache key) made a batch-32 decode 6% slower; kept open it is 0.18 ms
there and 0.005 ms on ext4. vLLM's CUDA-graph path with every adapter on: 1.005x, 1.009x and 0.999x at batch 1,
8 and 32 (control runs without entail 0.994-1.003x).
Docs
- README (EN/KO): the "4 of 4" sentence is marked as the bugs the facts were written from, and the pre-registered
replay with the vocabulary frozen at 1.1.0 is reported next to it: 150 issues screened, 17 passed, 15 reproduced,
7 in the class by two blind raters, 0 of the 7 detected, 0 false alarms on the 8 outside the class. Known gaps
list the facts and sites it exposed, and the hub-id loads that the Vocab and Stops checks cannot decide. - README (EN/KO): the second pre-registered replay, with the vocabulary frozen at this version's code (
ce79b19):
the next 150 issues screened, 20 passed, 15 reproduced, 8 in the class by two blind raters (kappa 0.72 over
seven categories, 0.68 in-class versus not; 18 of 86 settled by a third), 0 of the 8 detected (rule of three: at
most 3 of 8), 0 false alarms on the 7 reproduced outside the class, one run broken by entail (the serve wrapper,
fixed above). Known gaps name what it left unread, first the ids a built tokenizer produces.