Skip to content

entail 1.3.0

Choose a tag to compare

@wwoosshh wwoosshh released this 27 Sep 13:18
· 104 commits to main since this release

Released 2026-09-27. Reference comparisons where a declaration lives only in code (a tokenizer run against the
declared one, a custom op's kernel against its own definition, a parser's stream against its whole-text parse, a
multimodal placeholder's origin, a completion's logprobs), definitions for five engine functions that carry none,
checks that reach an engine's warm-up and capture (warm-up probes, a Triton launch's layout run twice), and the
engine's own paths held against each other at start. In two further pre-registered replays of real engine bugs, the
code frozen at 07fceac and at this version's 2aa975b, it detected 0 of 5 and 0 of 6 in-class bugs (0 of 26 over
four replays). Read it as a
light pre-deployment check that repairs the classes it knows and says where meaning broke, not as protection against
unseen bugs; see Measured for this release and Known issues.

Added

  • Warm-up probes (M19 L3.3a): the calls an engine makes before serving (vLLM's profile run and warm-ups, SGLang's
    capture warm-ups, recognised as a run of calls with one repeated row before the first real decision) decide the
    kernel-against-definition comparison, on made-up token values with the engine's own shapes, dtypes and strides,
    once per power-of-two size class and at the engine's own row count (512 rows at most). A repair can therefore be in
    place before a CUDA graph captures the kernel: it is offered when no graph captured the kernel at a size class not
    yet verified, and a repair now applies to every module of the same configuration, not only the one compared.

  • A Triton launch run twice (M19 L3.3b; kernel_layout_variant): the first launch of a pattern whose innermost
    stride the kernel is not told runs on copies, as given and relaid (innermost dimension contiguous, outer strides
    kept; wholly contiguous for a kernel that takes no stride), the relaid copy twice for the kernel's own noise. A
    difference is resolved by launching that pattern on relaid copies from then on and copying back only the tensors
    the kernel wrote. Not run twice: launches under capture, a written tensor overlapping another argument, tensors
    past the budget, layouts that cannot be kept, comparisons where every value is zero.

  • PathAgreement (vocabulary v10) and path_contract.py with the adapters vllm_paths and sglang_paths (M19
    L3.3c): right after vLLM's LLM or SGLang's Engine is built, three fixed probe texts go through the public
    generate API and three pairs of the engine's own paths are compared - decode against a fresh prefill of the same
    tokens, a request alone against the same requests batched, a cold run against one that reads the prefix cache.
    Rule paths_disagree (broken): a confident prediction (margin over 1.0) that changes, or a kept token's
    probability that moves by more than 0.25, thresholds set from healthy engines. The cache is cleared afterwards.
    ENTAIL_NO_PATHS=1 turns it off.

  • Exact definitions for index bookkeeping (M19 L3.3d): vLLM 0.30's prepare_pos_seq_lens (positions and sequence
    lengths) and BlockTables.compute_slot_mappings (the KV slot of every token) held against entail's definitions
    element by element (compare_exact), on fresh output buffers so the engine's persistent buffers are written only
    by the real call. Definitions can now declare the arguments they write, compare whole, use fresh buffers, compare
    exactly, and wrap class methods. The two were chosen because every model of a 14-model census ran them.

  • Definitions for engine functions that carry none (M19 L3): entail/definitions.py writes, in plain PyTorch, what
    vLLM 0.30's fused_experts (unquantized or INT8 W8A8: weight scales applied by their own shape, activations
    quantized as the config says), vLLM 0.30's w8a8_triton_block_scaled_mm and SGLang 0.5.20's fused_gdn_gating
    compute - the functions' own arguments, outputs of the same shape and dtype; a call a definition does not cover
    raises NotImplementedError and is unknown. adapters/function_reference.py wraps each function's name as soon as
    its module has loaded (start-up shim) and holds it against its definition on the first real call with the kernel
    reference rule (a 64-row slice, the definition in its noise dtype and in float32; calls whose rows are all one row,
    an engine's dummy batch, decide nothing and are not counted towards giving up); a mismatch is resolved by sending
    the name to the definition (float32, the function's output dtype) from that call on, unless the function was
    already called inside a CUDA graph capture in the process; a repaired function captured later puts its definition
    in the graph when the definition is capturable, else the kernel, said once as broken. Measured on an RTX 4070 Ti:
    vllm#58532 (INT8 MoE, per-channel weight scales, static activation scale), vllm#52576 (block FP8, BLOCK_SIZE_K 256)
    and sglang#21843 (GDN gate on inputs with an inner stride of 2) resolved on their first call, the output's error
    against a float64 formula 0.956 -> 0.0020, 0.612 -> 0.0027, 0.937 -> 1.5e-7; seven healthy controls of the same
    functions pass with the kernel's output untouched. Not caught: a bad launch configuration used only by a later,
    larger call (decided once, on the first call). On healthy engines (entail on against off, six greedy probes):
    SGLang Qwen3.5-4B with and without CUDA graphs and vLLM Qwen3-4B-FP8 on the Triton block-FP8 path (Marlin
    disabled; vLLM picks Marlin on sm_89 by default, which never reaches the function), eager and default mode: the
    function is reached, passes, and the output is unchanged (6/6).

  • Kernel reference repair (M19 L3): a custom op whose kernel differs from its own native definition is sent to that
    definition (forward_native) - from the very call that was compared (the slice is now decided before the real
    input is computed) and for every later call of that module in the process. Resolved in the record, with the
    resolution; offered only when the engine runs without CUDA graphs (captured graphs replay the kernel), and not
    under ENTAIL_POLICY=refuse or KernelReference=refuse. Measured: vllm#42016 (GLM-OCR, vLLM 0.22.0, eager)
    resolved, eager output now equal to the native mode's; a healthy Llama-3.2-3B changed nothing.

  • Tokenization (vocabulary v9) and tokenizer_contract.py (M18.1): the tokenizer the engine built is run against
    the tokenizer the folder declares, on ten fixed probe texts, and the ids must be the same. The declaration can be
    run: tokenizer.json by the tokenizers library (a sentencepiece-only folder is unknown until that reference is
    measured); tokenizer_config.json's
    added_tokens_decoder (else tokenizer.json's added_tokens, else added_tokens.json) names every added token with
    its id, and each is looked up in the built tokenizer. Rules tokenizer_ids and added_token_id, broken (reported,
    the run goes on); a declaration that cannot be run (no file, no library, a tiktoken.model without its
    pre-tokenization pattern: the added tokens are still compared) is unknown once per folder. The size check
    (Vocab) passed every tokenizer built from the right file by the wrong class; this is the check those bugs
    needed: transformers#46489 (deepseek-coder as LlamaTokenizer, 5.10.2), #45812 (Granite as GPT2Tokenizer, 5.8.0),
    #45356 (Kimi-K2.5's </think> given <|media_end|>'s id, 5.4.0), #46710 (DeepSeek-R1-Distill's declared class
    replaced, 5.12.1). The transformers_tokenizer adapter runs it after the size check; entail check runs it
    statically when transformers can build the tokenizer. The declared tokenizer's probe ids are kept per folder in
    entail_logs/tokenizer_ids.json, so a process after the first only encodes the probes with the engine's tokenizer.
    A difference that the folder's own declared flag explains - legacy: false (or add_prefix_space) next to a
    tokenizer.json exported the other way, on the text the flag speaks of (after a special token, or at the start)
    and exactly as the flag's pipeline gives it - is the sources disagreeing (unknown, the flag recorded as the
    conflicting source, ENTAIL_SOURCE_CONFLICT=stop honoured); every other difference is broken. A difference the
    user's own build settings explain (legacy=, add_prefix_space=, ... given to from_pretrained) is the user's
    choice. Measured (retrospective: the rule was written from these bugs): the four bugs at their reported versions
    are all broken at the tokenizer boundary (Kimi by 18 of 23 declared added tokens with other ids, the others by 4
    to 9 of the 10 probe texts); on 5.17.0 three pass and Kimi is unknown (tiktoken: added tokens compared, texts
    not). 38 popular folders on 5.17.0 (11 distinct tokenizers): 36 pass, 2 broken - the one Llama-2-era tokenizer in
    the set (TinyLlama-1.1B-Chat, a tiny test folder): transformers 5 rebuilds a legacy-export tokenizer.json as
    Metaspace and never doubles the ▁ before text that starts with whitespace, whatever legacy says. The 300-folder
    static corpus on 5.17.0: 201 pass, 9 broken, 5 unknown, 85 without a tokenizer to compare; the 9 are six of that
    Llama-2 shape and three folders that declare LlamaTokenizerFast over a byte-level BPE tokenizer.json
    (DeepSeek-R1-0528-Qwen3-8B, deepseek-coder-7b-instruct-v1.5, an MLX export), which 5.17.0 builds as a Llama
    pipeline: "How are you doing?" decodes back as "Howareyoudoing?". Cost: the first process on a folder +241 ms at
    the median and +894 ms at most (the reference is built), later processes +2 ms at the median, +36 ms at the 90th
    percentile. The sentencepiece reference is not compared until it is measured on sentencepiece-only folders
    (special-token strings would differ); tokenizer.json's padding and truncation are cleared in the reference and
    BPE dropout is not compared; every declared added token is looked up; the machine cache is keyed by the library
    version too, and lives per start folder.

  • KernelReference (vocabulary v9) and kernel_reference_contract.py with the vLLM adapter
    vllm_kernel_reference (M18.2): a custom op's dispatched kernel against the op's own native definition, run on
    the same input. vLLM's CustomOp carries its meaning as forward_native and dispatches to forward_cuda; after
    the model is built, every op dispatching to a kernel path is wrapped, and on its first real call per (op class
    and module, configuration, input pattern) the kernel and the definition are run on a 64-row slice of the real
    input (clones cut before the kernel touched its arguments; the engine's tensors are untouched), the definition in
    the input dtype and in float32, and the op's own tensors put back afterwards (a definition may convert its cache
    to the query's dtype). Rule kernel_reference_mismatch, decided value by value and per output tensor in its own
    dtype: a value non-finite on one side only, or differing from the float32 definition by more than FACTOR times
    the definition's own rounding noise at that value (backed by the tensor's typical noise) plus ATOL_ULPS units in
    the last place of the output dtype at that value; broken (reported, the run goes on). Afterwards the original
    method is put back, so the steady state costs nothing; under a stop policy the decision raises once. Not
    compared, each said unknown once: ops that override forward (the mamba mixers), ops holding the engine's
    state (a KV cache, an index buffer, a forward that reads the forward context), every op when the process is one
    rank of several, every op enabled under torch.compile (traced, the wrapper hands the call to the kernel), ops in
    vLLM's registry not reached from the model's modules, arguments that share no token dimension or cannot be cut
    and are too large to clone, and definitions that refuse the input. Calls inside vLLM's own dummy runs (profile,
    capture warm-ups) are neither compared nor counted, and an input that decides nothing (zeros, one repeated row,
    an identity such as rotary at position 0) leaves the wrapper on for the next real call (64 such calls at most).
    Measured (retrospective: the rule was written from this bug): vllm#42016 (GLM-OCR on vLLM 0.22.0, the Triton
    MRoPE kernel pairing split-wise for a model that pairs interleaved) is broken at MRotaryEmbedding on its
    first real input - max |kernel - definition| 10.5 at a scale of 10.9, allowed 0.588 - with no architecture
    table, and passes on 0.30.0; 8 popular models on vLLM 0.30.0 with enforce_eager: 35 decisions, 34 pass and one
    unknown (144 quant_fp8 instances held by linear-kernel helpers, not reached from the model's modules), every
    decision on the first real input. Of the 34, 13 compare an independent kernel (rotary 6, activations 7, the
    activations bitwise equal to the definition) and 21 hold the definition against itself, which the record says:
    in eager mode vLLM 0.30's RMSNorm.forward_cuda returns forward_native (the fused kernels are reached under
    torch.compile, where entail does not compare). The worst value's ratio to its allowance is at most 0.095
    (median 0.062). FACTOR 8 and ATOL_ULPS 4 are headroom, not derived from that distribution (the kernels'
    largest error equals the definition's own largest rounding step there, which any FACTOR of 1 or more admits);
    every decision records that ratio so a later measurement can fix them from data.

  • Parse (vocabulary v9), parse_contract.py and the vLLM adapter vllm_parse (M18.3): a chat parser's streamed
    message against its parse of the same complete text, and its tool calls against the tools the request declared.
    The class vLLM's server builds a parser from per request (ParserManager.get_parser, 0.30's unified parsers with
    parse_delta and parse) is returned wrapped: its instances accumulate what parse_delta hands on (content,
    reasoning, tool-call names and argument pieces), and when the stream finishes a fresh instance parses the whole
    text and the two must agree exactly (arguments as JSON values): rule stream_differs_from_full. A tool call
    that names an undeclared tool, lacks a parameter the declared tool requires, or carries a key the tool's
    parameters.properties do not have when the tool forbids additional properties (additionalProperties: false;
    JSON Schema allows them by default, so under a tool that did not forbid them an extra key is a note on a pass)
    is tool_args_outside_schema, on both paths; whether the key is in the model's text (the model's call does not
    fit the declared tool, or the parser reshaped it) or not (the parser added it) is said. Both broken (reported;
    the client already has the streamed message). An output that did not finish by itself - the request's token
    limit reached, the reasoning block still open, a forced tool whose arguments never became JSON - is where vLLM
    documents its two paths to differ, so a difference there is unknown; reasoning the stream sent again as
    content (vLLM's fallback) is a note. Deltas are accumulated with nothing recorded, the comparison runs once at
    the end of the stream, and a silent pass is counted, not recorded. Measured on vLLM 0.30.0's own parsers
    (retrospective: the rules were written from these bugs), driven as the server drives them with a stand-in
    tokenizer and no prompt: vllm#49316 (kimi_k2: the streamed path skips the schema's type coercion, 4 of 4 texts),
    #49412 (qwen3: the content around tool calls is dropped on the whole-text path, 2 of 3; the third differs in
    surrounding whitespace only) and #47986 (deepseek_v4: tool_b unwrapped with tool_a's schema, with tool_b
    declared precisely so that a correct parser passes the same rule) are broken; the well-formed texts raise
    nothing. Content that differs in surrounding whitespace only is a note on a pass, not broken: on a live vLLM
    server (Qwen3-0.6B, qwen3 reasoning parser, hermes tool parser, 18 streamed and whole requests) every tool-call
    stream streamed two newlines and parsed nothing whole, and a newline loses no meaning.

  • Placeholder (vocabulary v9), placeholder_contract.py and the vLLM adapter vllm_multimodal (M18.4): where
    vLLM binds a multimodal item's placeholder against the markup the model's config declares
    (vision_start_token_id before image_token_id, the Qwen-VL family): an image placeholder run not preceded by
    the declared start token came from the prompt's text, not from the template - a literal <|image_pad|> typed by
    the user took the image (vllm#57740). Rule placeholder_outside_markup, broken; a model that declares no markup
    decides nothing, and vLLM's own profiling prompts (placeholder runs from token 0, no template) are not decided.
    Measured: Qwen2.5-VL-3B-Instruct on vLLM 0.30.0 with the report's two message orders - the attack order is
    broken ("the image placeholder bound at tokens 20..275 is preceded by id 220"), the control order passes.

  • The SGLang serve adapter also decides a chat completion's logprobs against its message (M18.4,
    parse_contract.check_logprobs): the logprob tokens must decode to the content the client gets; with
    separate_reasoning SGLang's logprobs covered the whole raw output, <think> span and markers included, while
    message.content held the parsed answer (sglang#25055). Rule logprobs_cover_other_text, broken. Measured:
    SGLang 0.5.20, Qwen3-0.6B with the qwen3 reasoning parser, one request with logprobs and separate_reasoning:
    broken ("the 155 logprob tokens cover the reasoning span (545 characters and its markers) as well as the
    content (12 characters)").

  • The false-alarm yardstick now covers a live vLLM server with every adapter on (18 streamed and whole chat
    requests through a reasoning parser and a tool parser: nothing broken), ngram speculative decoding (nothing
    broken: the M17.6 narrowing of kv_needed holds) and a hybrid Mamba-attention model with several KV groups
    (nothing broken); prefill-decode disaggregation is not measured on one card.

Changed

  • Kernel reference slices keep their strides (kernel_reference_contract.kept): clone() made a view with gaps
    contiguous, so a kernel that misreads a layout read the slice right.
  • vLLM's dummy runs are marked in the second GPU runner's graph capture and memory profiling too (capture_model,
    profile_cudagraph_memory, which run outside _dummy_run): their warm-up calls had used up TRIES before the first
    real request.
  • SGLang's start-up path check runs only when asked for (ENTAIL_PATHS=1; otherwise the start boundary says once
    why it did not compare). SGLang runs no prefill while it starts, so the probe requests were the engine's first
    prefills, and a defect on that path stopped the engine before the caller's first request: SGLang 0.5.20 picks
    flashinfer for Phi-3.5-mini-instruct (head_dim 96), whose state merge does not take that head size, and any prompt
    of 128 tokens or more stops the scheduler, with entail off as well. vLLM's path check stays on.
  • A kernel comparison records its worst value-to-allowance ratio with a floor for an all-zero allowance (it was
    written as 0 when the float32 allowance underflowed).

Measured for this release (one RTX 4070 Ti; the research workspace's testbed/results/m19/l4/SUMMARY.md)

  • Healthy runs: 38 popular models on transformers 5.17, vLLM 0.30 and SGLang 0.5.20, 102 valid runs: entail broke no
    run; no check added since 1.2.0 said broken or refused; outputs identical in 97 of 98 comparisons without a
    repair (the one difference is an engine's own nondeterminism). Six runs say broken at the tokenizer boundary,
    all from two Llama-2-era folders (TinyLlama-1.1B-Chat and a tiny test folder) on every engine: a real difference -
    transformers 5 builds these tokenizers so that text starting with a space loses one space, where the folder's
    tokenizer.json, the model's sentencepiece file and transformers 4.57 agree (transformers#47700 describes it).
  • Request throughput with everything on against entail not installed, vLLM's default path (torch.compile and CUDA
    graphs), Qwen3-4B: 1.0007x, 1.0028x and 1.0110x at batch 1, 8 and 32 (two identical states differ by up to 0.6%).
  • Load: from an installed copy, +1.6 to +1.7 s on vLLM for 0.6-3B models (13-15% of their load), mostly the start-up
    path check (ENTAIL_NO_PATHS=1 turns it off); the hook alone -0.03 to +0.25 s.
  • Detection on unseen bugs (pre-registered, code frozen): the third replay (frozen at 1.2.0's successor 07fceac)
    detected 0 of the 5 reproduced in-class bugs; the fourth (frozen at this version's code, a new population of 437
    issues from the six months before) detected 0 of the 6, 0 of the 5 low-level ones, and raised one false alarm
    (below). Over four replays: 0 of 26.

Known issues

  • The start-up path check counts non-finite log-probabilities as agreement: a model whose every log-probability was
    NaN (vLLM 0.16, NVFP4 with float16 activations) passed all three pairs.
  • False alarm on encoder-decoder models on vLLM: whisper-large-v3-turbo gets nine broken lines at the KV cache
    boundary (container:vllm.allocate_slots, cache group 1 holds 0 slots for its tokens) although its output is
    right - the rule does not know that a cross-attention cache group follows the encoder, not the decoder's tokens.
    The run goes on (the default policy reports).
  • A tokenizer built from a GGUF file is unknown at the tokenizer boundary (the GGUF file's own tokenizer
    declaration is not read), so a GGUF tokenizer built as another type than the file declares is not caught
    (transformers#41494).
  • On vLLM's default path (torch.compile) custom ops are compiled and not compared with their definitions; the
    comparison runs in eager mode. Kernels called from C++ (Marlin) are not reached.