Skip to content

entail 1.2.0

Choose a tag to compare

@wwoosshh wwoosshh released this 26 Sep 09:55
· 60 commits to main since this release

Released 2026-09-26. The sites the first pre-registered replay found unread (a LoRA adapter's settings file, a
request's template settings at the reasoning parser, the prefix-cache key and the beam reorder, Triton kernel
launches, rotary pairing), each written from the real bug and measured on it, and a second pre-registered replay
with the vocabulary frozen at this version.

Added

  • A LoRA adapter's adapter_config.json as a declaration file (adapter_config_contract.py,
    data/adapter_config_keys.json): every PEFT key, which ones each consumer reads (PEFT for transformers and
    diffusers; vLLM 0.30's PEFTHelper; SGLang 0.5.20's LoRAConfig, with code lines), and one rule: a key declared with
    a value that changes how the weights apply, that the consumer does not read, is broken at that consumer's load
    boundary, or resolved where the consumer can carry it. Neutral values, keys the weights carry and training-time
    keys decide nothing; a key the consumer refuses loudly passes with a note. Adapters sglang_lora (carries
    use_rslora into the adapter's scaling, computed in the core: sglang#40835 served rsLoRA adapters 4-8x too weak)
    and vllm_lora (reports what vLLM drops: rank_pattern, alpha_pattern, lora_bias, ...). entail check on an
    adapter folder decides it per engine.
  • A request's template settings at the reasoning parser (request_contract.setting_names,
    data/request_settings.json): a setting the template honoured under one name (enable_thinking) that the
    parser reads under another (thinking) leaves the parser on its default (vllm#43728: content: null). The
    table names, per vLLM version and reasoning parser, the names each parser reads (code lines); the vLLM serve
    adapter wraps the parser class the server builds per request and hands the request's value to the parser under
    a name it reads (resolved), or reports it. A parser that reads no name of the setting, or a setting the
    template did not read either, decides nothing.
  • A store's key against the fields that shaped the item (cache_key_contract.py, data/cache_key_fields.json):
    vLLM's prefix-cache block hash keys a request by its tokens, embeddings digest, multimodal hashes, LoRA name and
    cache salt, but not by prompt_is_token_ids (which positions take the embeddings), so two requests that differ
    only in that mask share a key (vllm#56655, fix unmerged at 0.30.0). The adapter vllm_cache_key decides at
    Request.__init__ and repairs by appending a digest of the block's mask to the hash's extra keys and remaking
    the request's hashes (resolved). The same rule covers a permutation: transformers' beam search reorders the
    cache under the names it knows (transformers_beam); a model whose cache lives under another name is reported
    (5.12.1 reordered past_key_values only: transformers#46612; 5.17.0 reorders every name).
  • What a Triton kernel is told about its tensors (kernel_launch_contract.py, adapter triton_launch, engine-
    independent: one hook on JITFunction.run, so every @triton.jit kernel launched eagerly by any engine; kernels
    Inductor generates for a compiled forward are not seen): a tensor strided in its innermost dimension handed to a
    kernel that was not told that stride (no integer argument among its value parameters and stride-named
    constexprs equals it) and whose parameters name no stride at all (stride, _s0, sxm, ld...) is broken
    (the kernel reads it as if contiguous); handed to a kernel that names strides but was not told this one, it is
    unknown (said once). Each (kernel, stride pattern) is
    decided once per process, and a kernel is looked at for its first eight strided patterns; compile-only warm-ups
    are not launches. sglang#21843 (fused_gdn_gating read interleaved a/b) is the case the rule comes from; there
    the kernel takes row strides, so the decision is unknown at the kernel's boundary.
  • Rotary.pairing (vocabulary v8): how a rotary embedding pairs the dimensions it rotates, split (i with
    i + d/2, the Llama convention) or interleaved (2i with 2i+1, GPT-J's). Declared by a config key
    (rope_interleave, rope_interleaved, is_neox_style) or, failing that, by the architecture's reference
    implementation (data/rotary_pairing.json: transformers 5.17.0 configuration and modeling files with lines; GLM,
    Cohere, Ernie 4.5, GPT-J and DeepSeek-V3 interleaved, Llama, Qwen and Gemma split). rotary_pairing_contract.py
    decides the rotary modules vLLM built for the language model (is_neox_style, adapter vllm_pairing) against
    the declaration - resolved by setting the modules' convention, broken where the model's MRoPE module
    dispatches to a kernel that pairs split-wise whatever the layer says (vLLM's Triton MRoPE kernel before 0.27.0
    with the custom op enabled: vllm#42016, #49290). Not compared: a multimodal model's vision tower (its own
    reference), a DSA indexer (its own key), and modules that pair both ways in one language model (unknown, nothing
    set). An architecture the table does not know, without a key, decides nothing.

Fixed

  • The KV contract's kv_needed rule breaks only when a sequence holds fewer slots than its tokens. Holding more
    than one allocation unit over its tokens was also broken, and under speculative decoding an engine
    legitimately does that: it reserves lookahead slots and keeps the blocks of the drafts it rejected (vLLM 0.30
    with ngram speculation said broken on a healthy run; vLLM 0.23 with extract_hidden_states likewise, seen in
    the second replay). Fewer slots than tokens is the loss and is still reported.
  • The vLLM serve adapter's wrapper of ParserManager.get_parser binds its arguments by name and passes the rest
    through: with a fixed signature it raised TypeError on vLLM 0.23.0 (whose get_parser takes is_harmony)
    and the API server died - the one run entail broke in the second pre-registered replay.
  • A model loaded by hub id resolves to its cached snapshot folder again, so the Vocab and Stops checks decide
    instead of saying "no local folder to read": huggingface_hub 1.32 refuses a cached snapshot that lacks files the
    engine never fetched (.gitattributes, evaluation results), and every hub-id load on vLLM and transformers
    was "could not be checked". The folder of the cached config.json is used when the snapshot lookup refuses;
    nothing is downloaded.

Changed

  • The log and record files are kept open per process, one write and one flush per line, instead of being opened
    and closed per line: on a 9P mount (a project under WSL's /mnt/c) the open and close cost 4.6 ms per line, and a
    boundary that runs per request (the prefix-cache key) made a batch-32 decode 6% slower; kept open it is 0.18 ms
    there and 0.005 ms on ext4. vLLM's CUDA-graph path with every adapter on: 1.005x, 1.009x and 0.999x at batch 1,
    8 and 32 (control runs without entail 0.994-1.003x).

Docs

  • README (EN/KO): the "4 of 4" sentence is marked as the bugs the facts were written from, and the pre-registered
    replay with the vocabulary frozen at 1.1.0 is reported next to it: 150 issues screened, 17 passed, 15 reproduced,
    7 in the class by two blind raters, 0 of the 7 detected, 0 false alarms on the 8 outside the class. Known gaps
    list the facts and sites it exposed, and the hub-id loads that the Vocab and Stops checks cannot decide.
  • README (EN/KO): the second pre-registered replay, with the vocabulary frozen at this version's code (ce79b19):
    the next 150 issues screened, 20 passed, 15 reproduced, 8 in the class by two blind raters (kappa 0.72 over
    seven categories, 0.68 in-class versus not; 18 of 86 settled by a third), 0 of the 8 detected (rule of three: at
    most 3 of 8), 0 false alarms on the 7 reproduced outside the class, one run broken by entail (the serve wrapper,
    fixed above). Known gaps name what it left unread, first the ids a built tokenizer produces.