Repository navigation
entail 1.3.0
Released 2026-09-27. Reference comparisons where a declaration lives only in code (a tokenizer run against the
declared one, a custom op's kernel against its own definition, a parser's stream against its whole-text parse, a
multimodal placeholder's origin, a completion's logprobs), definitions for five engine functions that carry none,
checks that reach an engine's warm-up and capture (warm-up probes, a Triton launch's layout run twice), and the
engine's own paths held against each other at start. In two further pre-registered replays of real engine bugs, the
code frozen at 07fceac and at this version's 2aa975b, it detected 0 of 5 and 0 of 6 in-class bugs (0 of 26 over
four replays). Read it as a
light pre-deployment check that repairs the classes it knows and says where meaning broke, not as protection against
unseen bugs; see Measured for this release and Known issues.
Added
-
Warm-up probes (M19 L3.3a): the calls an engine makes before serving (vLLM's profile run and warm-ups, SGLang's
capture warm-ups, recognised as a run of calls with one repeated row before the first real decision) decide the
kernel-against-definition comparison, on made-up token values with the engine's own shapes, dtypes and strides,
once per power-of-two size class and at the engine's own row count (512 rows at most). A repair can therefore be in
place before a CUDA graph captures the kernel: it is offered when no graph captured the kernel at a size class not
yet verified, and a repair now applies to every module of the same configuration, not only the one compared. -
A Triton launch run twice (M19 L3.3b;
kernel_layout_variant): the first launch of a pattern whose innermost
stride the kernel is not told runs on copies, as given and relaid (innermost dimension contiguous, outer strides
kept; wholly contiguous for a kernel that takes no stride), the relaid copy twice for the kernel's own noise. A
difference is resolved by launching that pattern on relaid copies from then on and copying back only the tensors
the kernel wrote. Not run twice: launches under capture, a written tensor overlapping another argument, tensors
past the budget, layouts that cannot be kept, comparisons where every value is zero. -
PathAgreement(vocabulary v10) andpath_contract.pywith the adaptersvllm_pathsandsglang_paths(M19
L3.3c): right after vLLM'sLLMor SGLang'sEngineis built, three fixed probe texts go through the public
generate API and three pairs of the engine's own paths are compared - decode against a fresh prefill of the same
tokens, a request alone against the same requests batched, a cold run against one that reads the prefix cache.
Rulepaths_disagree(broken): a confident prediction (margin over 1.0) that changes, or a kept token's
probability that moves by more than 0.25, thresholds set from healthy engines. The cache is cleared afterwards.
ENTAIL_NO_PATHS=1turns it off. -
Exact definitions for index bookkeeping (M19 L3.3d): vLLM 0.30's
prepare_pos_seq_lens(positions and sequence
lengths) andBlockTables.compute_slot_mappings(the KV slot of every token) held against entail's definitions
element by element (compare_exact), on fresh output buffers so the engine's persistent buffers are written only
by the real call. Definitions can now declare the arguments they write, compare whole, use fresh buffers, compare
exactly, and wrap class methods. The two were chosen because every model of a 14-model census ran them. -
Definitions for engine functions that carry none (M19 L3):
entail/definitions.pywrites, in plain PyTorch, what
vLLM 0.30'sfused_experts(unquantized or INT8 W8A8: weight scales applied by their own shape, activations
quantized as the config says), vLLM 0.30'sw8a8_triton_block_scaled_mmand SGLang 0.5.20'sfused_gdn_gating
compute - the functions' own arguments, outputs of the same shape and dtype; a call a definition does not cover
raises NotImplementedError and isunknown.adapters/function_reference.pywraps each function's name as soon as
its module has loaded (start-up shim) and holds it against its definition on the first real call with the kernel
reference rule (a 64-row slice, the definition in its noise dtype and in float32; calls whose rows are all one row,
an engine's dummy batch, decide nothing and are not counted towards giving up); a mismatch is resolved by sending
the name to the definition (float32, the function's output dtype) from that call on, unless the function was
already called inside a CUDA graph capture in the process; a repaired function captured later puts its definition
in the graph when the definition is capturable, else the kernel, said once as broken. Measured on an RTX 4070 Ti:
vllm#58532 (INT8 MoE, per-channel weight scales, static activation scale), vllm#52576 (block FP8, BLOCK_SIZE_K 256)
and sglang#21843 (GDN gate on inputs with an inner stride of 2) resolved on their first call, the output's error
against a float64 formula 0.956 -> 0.0020, 0.612 -> 0.0027, 0.937 -> 1.5e-7; seven healthy controls of the same
functions pass with the kernel's output untouched. Not caught: a bad launch configuration used only by a later,
larger call (decided once, on the first call). On healthy engines (entail on against off, six greedy probes):
SGLang Qwen3.5-4B with and without CUDA graphs and vLLM Qwen3-4B-FP8 on the Triton block-FP8 path (Marlin
disabled; vLLM picks Marlin on sm_89 by default, which never reaches the function), eager and default mode: the
function is reached, passes, and the output is unchanged (6/6). -
Kernel reference repair (M19 L3): a custom op whose kernel differs from its own native definition is sent to that
definition (forward_native) - from the very call that was compared (the slice is now decided before the real
input is computed) and for every later call of that module in the process. Resolved in the record, with the
resolution; offered only when the engine runs without CUDA graphs (captured graphs replay the kernel), and not
underENTAIL_POLICY=refuseorKernelReference=refuse. Measured: vllm#42016 (GLM-OCR, vLLM 0.22.0, eager)
resolved, eager output now equal to the native mode's; a healthy Llama-3.2-3B changed nothing. -
Tokenization(vocabulary v9) andtokenizer_contract.py(M18.1): the tokenizer the engine built is run against
the tokenizer the folder declares, on ten fixed probe texts, and the ids must be the same. The declaration can be
run: tokenizer.json by thetokenizerslibrary (a sentencepiece-only folder isunknownuntil that reference is
measured); tokenizer_config.json's
added_tokens_decoder(else tokenizer.json'sadded_tokens, else added_tokens.json) names every added token with
its id, and each is looked up in the built tokenizer. Rulestokenizer_idsandadded_token_id,broken(reported,
the run goes on); a declaration that cannot be run (no file, no library, a tiktoken.model without its
pre-tokenization pattern: the added tokens are still compared) isunknownonce per folder. The size check
(Vocab) passed every tokenizer built from the right file by the wrong class; this is the check those bugs
needed: transformers#46489 (deepseek-coder as LlamaTokenizer, 5.10.2), #45812 (Granite as GPT2Tokenizer, 5.8.0),
#45356 (Kimi-K2.5's</think>given<|media_end|>'s id, 5.4.0), #46710 (DeepSeek-R1-Distill's declared class
replaced, 5.12.1). Thetransformers_tokenizeradapter runs it after the size check;entail checkruns it
statically when transformers can build the tokenizer. The declared tokenizer's probe ids are kept per folder in
entail_logs/tokenizer_ids.json, so a process after the first only encodes the probes with the engine's tokenizer.
A difference that the folder's own declared flag explains -legacy: false(oradd_prefix_space) next to a
tokenizer.json exported the other way, on the text the flag speaks of (after a special token, or at the start)
and exactly as the flag's pipeline gives it - is the sources disagreeing (unknown, the flag recorded as the
conflicting source,ENTAIL_SOURCE_CONFLICT=stophonoured); every other difference isbroken. A difference the
user's own build settings explain (legacy=,add_prefix_space=, ... given to from_pretrained) is the user's
choice. Measured (retrospective: the rule was written from these bugs): the four bugs at their reported versions
are allbrokenat the tokenizer boundary (Kimi by 18 of 23 declared added tokens with other ids, the others by 4
to 9 of the 10 probe texts); on 5.17.0 three pass and Kimi isunknown(tiktoken: added tokens compared, texts
not). 38 popular folders on 5.17.0 (11 distinct tokenizers): 36 pass, 2 broken - the one Llama-2-era tokenizer in
the set (TinyLlama-1.1B-Chat, a tiny test folder): transformers 5 rebuilds a legacy-export tokenizer.json as
Metaspace and never doubles the▁before text that starts with whitespace, whateverlegacysays. The 300-folder
static corpus on 5.17.0: 201 pass, 9 broken, 5 unknown, 85 without a tokenizer to compare; the 9 are six of that
Llama-2 shape and three folders that declareLlamaTokenizerFastover a byte-level BPE tokenizer.json
(DeepSeek-R1-0528-Qwen3-8B, deepseek-coder-7b-instruct-v1.5, an MLX export), which 5.17.0 builds as a Llama
pipeline: "How are you doing?" decodes back as "Howareyoudoing?". Cost: the first process on a folder +241 ms at
the median and +894 ms at most (the reference is built), later processes +2 ms at the median, +36 ms at the 90th
percentile. The sentencepiece reference is not compared until it is measured on sentencepiece-only folders
(special-token strings would differ); tokenizer.json's padding and truncation are cleared in the reference and
BPE dropout is not compared; every declared added token is looked up; the machine cache is keyed by the library
version too, and lives per start folder. -
KernelReference(vocabulary v9) andkernel_reference_contract.pywith the vLLM adapter
vllm_kernel_reference(M18.2): a custom op's dispatched kernel against the op's own native definition, run on
the same input. vLLM's CustomOp carries its meaning asforward_nativeand dispatches toforward_cuda; after
the model is built, every op dispatching to a kernel path is wrapped, and on its first real call per (op class
and module, configuration, input pattern) the kernel and the definition are run on a 64-row slice of the real
input (clones cut before the kernel touched its arguments; the engine's tensors are untouched), the definition in
the input dtype and in float32, and the op's own tensors put back afterwards (a definition may convert its cache
to the query's dtype). Rulekernel_reference_mismatch, decided value by value and per output tensor in its own
dtype: a value non-finite on one side only, or differing from the float32 definition by more than FACTOR times
the definition's own rounding noise at that value (backed by the tensor's typical noise) plus ATOL_ULPS units in
the last place of the output dtype at that value;broken(reported, the run goes on). Afterwards the original
method is put back, so the steady state costs nothing; under a stop policy the decision raises once. Not
compared, each saidunknownonce: ops that overrideforward(the mamba mixers), ops holding the engine's
state (a KV cache, an index buffer, a forward that reads the forward context), every op when the process is one
rank of several, every op enabled under torch.compile (traced, the wrapper hands the call to the kernel), ops in
vLLM's registry not reached from the model's modules, arguments that share no token dimension or cannot be cut
and are too large to clone, and definitions that refuse the input. Calls inside vLLM's own dummy runs (profile,
capture warm-ups) are neither compared nor counted, and an input that decides nothing (zeros, one repeated row,
an identity such as rotary at position 0) leaves the wrapper on for the next real call (64 such calls at most).
Measured (retrospective: the rule was written from this bug): vllm#42016 (GLM-OCR on vLLM 0.22.0, the Triton
MRoPE kernel pairing split-wise for a model that pairs interleaved) isbrokenatMRotaryEmbeddingon its
first real input - max |kernel - definition| 10.5 at a scale of 10.9, allowed 0.588 - with no architecture
table, and passes on 0.30.0; 8 popular models on vLLM 0.30.0 withenforce_eager: 35 decisions, 34 pass and one
unknown(144quant_fp8instances held by linear-kernel helpers, not reached from the model's modules), every
decision on the first real input. Of the 34, 13 compare an independent kernel (rotary 6, activations 7, the
activations bitwise equal to the definition) and 21 hold the definition against itself, which the record says:
in eager mode vLLM 0.30'sRMSNorm.forward_cudareturnsforward_native(the fused kernels are reached under
torch.compile, where entail does not compare). The worst value's ratio to its allowance is at most 0.095
(median 0.062). FACTOR 8 and ATOL_ULPS 4 are headroom, not derived from that distribution (the kernels'
largest error equals the definition's own largest rounding step there, which any FACTOR of 1 or more admits);
every decision records that ratio so a later measurement can fix them from data. -
Parse(vocabulary v9),parse_contract.pyand the vLLM adaptervllm_parse(M18.3): a chat parser's streamed
message against its parse of the same complete text, and its tool calls against the tools the request declared.
The class vLLM's server builds a parser from per request (ParserManager.get_parser, 0.30's unified parsers with
parse_deltaandparse) is returned wrapped: its instances accumulate whatparse_deltahands on (content,
reasoning, tool-call names and argument pieces), and when the stream finishes a fresh instance parses the whole
text and the two must agree exactly (arguments as JSON values): rulestream_differs_from_full. A tool call
that names an undeclared tool, lacks a parameter the declared tool requires, or carries a key the tool's
parameters.propertiesdo not have when the tool forbids additional properties (additionalProperties: false;
JSON Schema allows them by default, so under a tool that did not forbid them an extra key is a note on a pass)
istool_args_outside_schema, on both paths; whether the key is in the model's text (the model's call does not
fit the declared tool, or the parser reshaped it) or not (the parser added it) is said. Bothbroken(reported;
the client already has the streamed message). An output that did not finish by itself - the request's token
limit reached, the reasoning block still open, a forced tool whose arguments never became JSON - is where vLLM
documents its two paths to differ, so a difference there isunknown; reasoning the stream sent again as
content (vLLM's fallback) is a note. Deltas are accumulated with nothing recorded, the comparison runs once at
the end of the stream, and a silent pass is counted, not recorded. Measured on vLLM 0.30.0's own parsers
(retrospective: the rules were written from these bugs), driven as the server drives them with a stand-in
tokenizer and no prompt: vllm#49316 (kimi_k2: the streamed path skips the schema's type coercion, 4 of 4 texts),
#49412 (qwen3: the content around tool calls is dropped on the whole-text path, 2 of 3; the third differs in
surrounding whitespace only) and #47986 (deepseek_v4: tool_b unwrapped with tool_a's schema, with tool_b
declared precisely so that a correct parser passes the same rule) arebroken; the well-formed texts raise
nothing. Content that differs in surrounding whitespace only is a note on a pass, not broken: on a live vLLM
server (Qwen3-0.6B, qwen3 reasoning parser, hermes tool parser, 18 streamed and whole requests) every tool-call
stream streamed two newlines and parsed nothing whole, and a newline loses no meaning. -
Placeholder(vocabulary v9),placeholder_contract.pyand the vLLM adaptervllm_multimodal(M18.4): where
vLLM binds a multimodal item's placeholder against the markup the model's config declares
(vision_start_token_idbeforeimage_token_id, the Qwen-VL family): an image placeholder run not preceded by
the declared start token came from the prompt's text, not from the template - a literal<|image_pad|>typed by
the user took the image (vllm#57740). Ruleplaceholder_outside_markup,broken; a model that declares no markup
decides nothing, and vLLM's own profiling prompts (placeholder runs from token 0, no template) are not decided.
Measured: Qwen2.5-VL-3B-Instruct on vLLM 0.30.0 with the report's two message orders - the attack order is
broken("the image placeholder bound at tokens 20..275 is preceded by id 220"), the control order passes. -
The SGLang serve adapter also decides a chat completion's logprobs against its message (M18.4,
parse_contract.check_logprobs): the logprob tokens must decode to the content the client gets; with
separate_reasoningSGLang's logprobs covered the whole raw output,<think>span and markers included, while
message.contentheld the parsed answer (sglang#25055). Rulelogprobs_cover_other_text,broken. Measured:
SGLang 0.5.20, Qwen3-0.6B with the qwen3 reasoning parser, one request withlogprobsandseparate_reasoning:
broken("the 155 logprob tokens cover the reasoning span (545 characters and its markers) as well as the
content (12 characters)"). -
The false-alarm yardstick now covers a live vLLM server with every adapter on (18 streamed and whole chat
requests through a reasoning parser and a tool parser: nothing broken), ngram speculative decoding (nothing
broken: the M17.6 narrowing ofkv_neededholds) and a hybrid Mamba-attention model with several KV groups
(nothing broken); prefill-decode disaggregation is not measured on one card.
Changed
- Kernel reference slices keep their strides (
kernel_reference_contract.kept):clone()made a view with gaps
contiguous, so a kernel that misreads a layout read the slice right. - vLLM's dummy runs are marked in the second GPU runner's graph capture and memory profiling too (
capture_model,
profile_cudagraph_memory, which run outside_dummy_run): their warm-up calls had used up TRIES before the first
real request. - SGLang's start-up path check runs only when asked for (
ENTAIL_PATHS=1; otherwise the start boundary says once
why it did not compare). SGLang runs no prefill while it starts, so the probe requests were the engine's first
prefills, and a defect on that path stopped the engine before the caller's first request: SGLang 0.5.20 picks
flashinfer for Phi-3.5-mini-instruct (head_dim 96), whose state merge does not take that head size, and any prompt
of 128 tokens or more stops the scheduler, with entail off as well. vLLM's path check stays on. - A kernel comparison records its worst value-to-allowance ratio with a floor for an all-zero allowance (it was
written as 0 when the float32 allowance underflowed).
Measured for this release (one RTX 4070 Ti; the research workspace's testbed/results/m19/l4/SUMMARY.md)
- Healthy runs: 38 popular models on transformers 5.17, vLLM 0.30 and SGLang 0.5.20, 102 valid runs: entail broke no
run; no check added since 1.2.0 saidbrokenorrefused; outputs identical in 97 of 98 comparisons without a
repair (the one difference is an engine's own nondeterminism). Six runs saybrokenat the tokenizer boundary,
all from two Llama-2-era folders (TinyLlama-1.1B-Chat and a tiny test folder) on every engine: a real difference -
transformers 5 builds these tokenizers so that text starting with a space loses one space, where the folder's
tokenizer.json, the model's sentencepiece file and transformers 4.57 agree (transformers#47700 describes it). - Request throughput with everything on against entail not installed, vLLM's default path (torch.compile and CUDA
graphs), Qwen3-4B: 1.0007x, 1.0028x and 1.0110x at batch 1, 8 and 32 (two identical states differ by up to 0.6%). - Load: from an installed copy, +1.6 to +1.7 s on vLLM for 0.6-3B models (13-15% of their load), mostly the start-up
path check (ENTAIL_NO_PATHS=1turns it off); the hook alone -0.03 to +0.25 s. - Detection on unseen bugs (pre-registered, code frozen): the third replay (frozen at 1.2.0's successor
07fceac)
detected 0 of the 5 reproduced in-class bugs; the fourth (frozen at this version's code, a new population of 437
issues from the six months before) detected 0 of the 6, 0 of the 5 low-level ones, and raised one false alarm
(below). Over four replays: 0 of 26.
Known issues
- The start-up path check counts non-finite log-probabilities as agreement: a model whose every log-probability was
NaN (vLLM 0.16, NVFP4 with float16 activations) passed all three pairs. - False alarm on encoder-decoder models on vLLM: whisper-large-v3-turbo gets nine
brokenlines at the KV cache
boundary (container:vllm.allocate_slots, cache group 1 holds 0 slots for its tokens) although its output is
right - the rule does not know that a cross-attention cache group follows the encoder, not the decoder's tokens.
The run goes on (the default policy reports). - A tokenizer built from a GGUF file is
unknownat the tokenizer boundary (the GGUF file's own tokenizer
declaration is not read), so a GGUF tokenizer built as another type than the file declares is not caught
(transformers#41494). - On vLLM's default path (torch.compile) custom ops are compiled and not compared with their definitions; the
comparison runs in eager mode. Kernels called from C++ (Marlin) are not reached.