Releases: wwoosshh/Entail
Release list
entail 1.2.0
Released 2026-09-26. The sites the first pre-registered replay found unread (a LoRA adapter's settings file, a
request's template settings at the reasoning parser, the prefix-cache key and the beam reorder, Triton kernel
launches, rotary pairing), each written from the real bug and measured on it, and a second pre-registered replay
with the vocabulary frozen at this version.
Added
- A LoRA adapter's
adapter_config.jsonas a declaration file (adapter_config_contract.py,
data/adapter_config_keys.json): every PEFT key, which ones each consumer reads (PEFT for transformers and
diffusers; vLLM 0.30's PEFTHelper; SGLang 0.5.20's LoRAConfig, with code lines), and one rule: a key declared with
a value that changes how the weights apply, that the consumer does not read, isbrokenat that consumer's load
boundary, orresolvedwhere the consumer can carry it. Neutral values, keys the weights carry and training-time
keys decide nothing; a key the consumer refuses loudly passes with a note. Adapterssglang_lora(carries
use_rslorainto the adapter's scaling, computed in the core: sglang#40835 served rsLoRA adapters 4-8x too weak)
andvllm_lora(reports what vLLM drops:rank_pattern,alpha_pattern,lora_bias, ...).entail checkon an
adapter folder decides it per engine. - A request's template settings at the reasoning parser (
request_contract.setting_names,
data/request_settings.json): a setting the template honoured under one name (enable_thinking) that the
parser reads under another (thinking) leaves the parser on its default (vllm#43728:content: null). The
table names, per vLLM version and reasoning parser, the names each parser reads (code lines); the vLLM serve
adapter wraps the parser class the server builds per request and hands the request's value to the parser under
a name it reads (resolved), or reports it. A parser that reads no name of the setting, or a setting the
template did not read either, decides nothing. - A store's key against the fields that shaped the item (
cache_key_contract.py,data/cache_key_fields.json):
vLLM's prefix-cache block hash keys a request by its tokens, embeddings digest, multimodal hashes, LoRA name and
cache salt, but not byprompt_is_token_ids(which positions take the embeddings), so two requests that differ
only in that mask share a key (vllm#56655, fix unmerged at 0.30.0). The adaptervllm_cache_keydecides at
Request.__init__and repairs by appending a digest of the block's mask to the hash's extra keys and remaking
the request's hashes (resolved). The same rule covers a permutation: transformers' beam search reorders the
cache under the names it knows (transformers_beam); a model whose cache lives under another name is reported
(5.12.1 reorderedpast_key_valuesonly: transformers#46612; 5.17.0 reorders every name). - What a Triton kernel is told about its tensors (
kernel_launch_contract.py, adaptertriton_launch, engine-
independent: one hook onJITFunction.run, so every@triton.jitkernel launched eagerly by any engine; kernels
Inductor generates for a compiled forward are not seen): a tensor strided in its innermost dimension handed to a
kernel that was not told that stride (no integer argument among its value parameters and stride-named
constexprs equals it) and whose parameters name no stride at all (stride,_s0,sxm,ld...) isbroken
(the kernel reads it as if contiguous); handed to a kernel that names strides but was not told this one, it is
unknown(said once). Each (kernel, stride pattern) is
decided once per process, and a kernel is looked at for its first eight strided patterns; compile-only warm-ups
are not launches. sglang#21843 (fused_gdn_gating read interleaved a/b) is the case the rule comes from; there
the kernel takes row strides, so the decision isunknownat the kernel's boundary. Rotary.pairing(vocabulary v8): how a rotary embedding pairs the dimensions it rotates,split(i with
i + d/2, the Llama convention) orinterleaved(2i with 2i+1, GPT-J's). Declared by a config key
(rope_interleave,rope_interleaved,is_neox_style) or, failing that, by the architecture's reference
implementation (data/rotary_pairing.json: transformers 5.17.0 configuration and modeling files with lines; GLM,
Cohere, Ernie 4.5, GPT-J and DeepSeek-V3 interleaved, Llama, Qwen and Gemma split).rotary_pairing_contract.py
decides the rotary modules vLLM built for the language model (is_neox_style, adaptervllm_pairing) against
the declaration -resolvedby setting the modules' convention,brokenwhere the model's MRoPE module
dispatches to a kernel that pairs split-wise whatever the layer says (vLLM's Triton MRoPE kernel before 0.27.0
with the custom op enabled: vllm#42016, #49290). Not compared: a multimodal model's vision tower (its own
reference), a DSA indexer (its own key), and modules that pair both ways in one language model (unknown, nothing
set). An architecture the table does not know, without a key, decides nothing.
Fixed
- The KV contract's
kv_neededrule breaks only when a sequence holds fewer slots than its tokens. Holding more
than one allocation unit over its tokens was alsobroken, and under speculative decoding an engine
legitimately does that: it reserves lookahead slots and keeps the blocks of the drafts it rejected (vLLM 0.30
with ngram speculation saidbrokenon a healthy run; vLLM 0.23 with extract_hidden_states likewise, seen in
the second replay). Fewer slots than tokens is the loss and is still reported. - The vLLM serve adapter's wrapper of
ParserManager.get_parserbinds its arguments by name and passes the rest
through: with a fixed signature it raisedTypeErroron vLLM 0.23.0 (whoseget_parsertakesis_harmony)
and the API server died - the one run entail broke in the second pre-registered replay. - A model loaded by hub id resolves to its cached snapshot folder again, so the Vocab and Stops checks decide
instead of saying "no local folder to read": huggingface_hub 1.32 refuses a cached snapshot that lacks files the
engine never fetched (.gitattributes, evaluation results), and every hub-id load on vLLM and transformers
was "could not be checked". The folder of the cachedconfig.jsonis used when the snapshot lookup refuses;
nothing is downloaded.
Changed
- The log and record files are kept open per process, one write and one flush per line, instead of being opened
and closed per line: on a 9P mount (a project under WSL's/mnt/c) the open and close cost 4.6 ms per line, and a
boundary that runs per request (the prefix-cache key) made a batch-32 decode 6% slower; kept open it is 0.18 ms
there and 0.005 ms on ext4. vLLM's CUDA-graph path with every adapter on: 1.005x, 1.009x and 0.999x at batch 1,
8 and 32 (control runs without entail 0.994-1.003x).
Docs
- README (EN/KO): the "4 of 4" sentence is marked as the bugs the facts were written from, and the pre-registered
replay with the vocabulary frozen at 1.1.0 is reported next to it: 150 issues screened, 17 passed, 15 reproduced,
7 in the class by two blind raters, 0 of the 7 detected, 0 false alarms on the 8 outside the class. Known gaps
list the facts and sites it exposed, and the hub-id loads that the Vocab and Stops checks cannot decide. - README (EN/KO): the second pre-registered replay, with the vocabulary frozen at this version's code (
ce79b19):
the next 150 issues screened, 20 passed, 15 reproduced, 8 in the class by two blind raters (kappa 0.72 over
seven categories, 0.68 in-class versus not; 18 of 86 settled by a third), 0 of the 8 detected (rule of three: at
most 3 of 8), 0 false alarms on the 7 reproduced outside the class, one run broken by entail (the serve wrapper,
fixed above). Known gaps name what it left unread, first the ids a built tokenizer produces.
entail 1.1.0
Released 2026-09-26. Five facts from the low-level study (codebook v2): classes of wrong output that 1.0 did not read,
each measured on the real bug it comes from.
Added
-
Fact vocabulary v5:
Identity(TIME) — what a stored or cached item stands for, so a store keyed by identity does
not serve one sequence's KV under another's key. Ruleidentity_staleinidentity_contract.py; the repair is to
forget the stale identities and let the store remake them. -
vLLM adapter
vllm_identity: wrapsScheduler._update_request_as_sessionand checks a request's prefix-cache
block hashes against the hashes its current tokens give, from the truncation point on. This is the class behind
vllm#49377 and #49449 (a streaming-session rebuild leaves stale block hashes; the fix PRs are unmerged, so it is
live in vLLM 0.30.0). Measured end to end on SmolLM2-135M: the stale hash is caught and repaired, and the wrong
output (a false 16-token cache hit) becomes the correct recomputed output. entail 1.0.0-1.0.2 passed it (E3 miss). -
Fact vocabulary v6:
TokenType(MAPPING) — the token type id a position is given by its role (padding).
request_contract.pad_typecompares the id a server gave the padding with the tokenizer'spad_token_type_id.
vLLM adaptervllm_scoring: wraps the scoring processor's padding of token type ids (vllm#58138: a cross-encoder's
padding was given the document's segment, and /rerank scores moved). The repair (give the padding the declared
id) is offered only where the consumer can carry it; vLLM 0.30 keeps token types as the index of the first 1, so
it cannot (capability row, measured): the decision is broken under the default policy and the padded request is
refused before scoring underENTAIL_ON_BROKEN=stop. -
Fact vocabulary v6:
KernelConfig(LAYOUT) — the tile a kernel steps K in, against the block the weights are
quantized in; the tile must be a divisor of the block (tile_contract.py, ruletile_over_block). SGLang adapter
sglang_fp8_tile: wraps the block-FP8 Triton matmul and checks the config map it picks from, once per map; the
repair clamps the tile to the block (the engine's own default). sglang#39626 (a hand-supplied K tile of 64 over a
block of 32 gave 64 where 288 was right): measured on 0.5.20, the tile is clamped and the kernel returns 288.
SGLang's shipped tuned configs (1,887 entries) all divide, so ordinary runs decide nothing. -
Records name a transformers config by its class and quote the file's own
_name_or_pathas the file's claim,
since that field can be stale (a checkpoint copied from another model). -
Fact vocabulary v6:
Vocab(MAPPING) — the tokenizer's base vocabulary.vocab_contract.pyreads what a model
folder declares (tokenizer.json, vocab.txt, vocab.json, a sentencepiece model, config vocab_size, the embedding's
rows) and decides the tokenizer the engine built against it, with two rules and no threshold: the tokenizer's ids
must fit the embedding, and when the folder carries two vocabularies the engine must hold the model's. Adapter
transformers_tokenizerwrapsPreTrainedTokenizerBase.from_pretrained(vLLM and SGLang build their tokenizers
through it too);entail checkruns the same rule statically. transformers#48967 (a folder with vocab.txt of
100,000 and a tokenizer.json of 32,000; transformers 5 built the 32,000 one and the ids changed): reported as
broken, refused before the first id underENTAIL_ON_BROKEN=stop, exit 1 fromentail check. -
Start-up hook: the target table had the tokenizer module keyed twice, so the second adapter was silently dropped;
one entry now, and a test guards the table against repeated keys. -
Rotarygains optional fields (v6): yarn'sbeta_fast,beta_slow,attention_factor(also spelt
attn_factor),mscale,mscale_all_dim,truncate; longrope'slong_factor/short_factoras a SHA-256
digest and their count;partial_rotary_factor(a top-level key, or insiderope_parameters);local_thetaand
local_factorfor a model that alternates two RoPEs (Gemma 3:rope_local_base_freq, orrope_parameterssplit
into full_attention/sliding_attention - the local layers declare no scaling, so an engine that scales them is
caught; a stated local RoPE without a factor islocal_factor1.0, a definite value);mrope_sectionand
mrope_interleaved(the Qwen-VL family;mropeis not a rope type - transformers 5 normalises the old spelling
{"type": "mrope"}todefault, and the fact of mrope is its section); and the rope typeproportional
(Gemma 4).original_max_position_embeddingsis read from the config's top level when the scaling dict has none
(Phi). A localpartial_rotary_factoror scaling type different from the global one is reported as beyond the
vocabulary, not dropped. Over the 230 most-downloaded models, RoPE declarations outside the vocabulary went from
34 to 3 (twoattn_factor, a name no engine reads, and one localpartial_rotary_factor; all three are said as
not compared), and a RoPE key the ENGINE's config holds that the vocabulary cannot carry is now reported at the
RoPE boundary instead of dropped. An alias is added only for a name a consumer reads. -
sglang_fp8_tilealso wraps the fused-MoE config lookup (try_get_optimal_moe_config), which has no sanitiser:
SGLang 0.5.20 ships an H100 config for E=512, N=256, fp8 block [128, 128] whose BLOCK_SIZE_K is 256; at the
kernel level that returns 256 where 512 is right, and the clamp restores 512. The dense hot path now costs a
dict lookup per call and writes no record line once a map is decided. -
Vocab: only a base vocabulary larger than the embedding is broken; an added token past the rows (gemma-3-1b-it's
image token on the text-only model) is noted on a passing decision. The embedding is the tensor with at least the
config's vocabulary of rows; a stale second tokenizer source that is not the model's vocabulary is unknown, not
judged. Tokenizers built outsidePreTrainedTokenizerBase.from_pretrained(Mistral, tiktoken, GGUF) are not
checked at run time;entail checksays so when Mistral files are present. -
Fact vocabulary v7:
Stops(MAPPING) - the token ids a generation ends with (and begins with, and is padded
with), as each file states them: generation_config.json, config.json and the tokenizer's eos_token. Every engine
builds its stop set from a different subset (data/stops_sources.json: transformers from generation_config.json
alone, vLLM from the tokenizer's eos plus generation_config.json, SGLang from config.json, generation_config.json
and the tokenizer's eos its scheduler matches), so an end one file declares can be one the engine never sees and the model runs past
the end of its answer (Llama 3, April 2024).stops_contract.pytakes the union: a consumer whose set lacks a
declared end is resolved by adding it (adapterstransformers_stops,vllm_stops,sglang_stops), a declared
id past the tokenizer is broken;entail checkdecides the set each engine would build. The tokenizer's own
declaration (eos_tokenin tokenizer_config.json) is read as an id by that file's added-token table, without
building a tokenizer. Measured: a Llama-3.2-3B-Instruct copy whose generation_config.json names only
<|end_of_text|>ran every answer to the token limit on transformers 5.17 and stops at the end with entail; and
nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 as shipped does the same (its config.json and auto-written
generation_config.json name</s>, its tokenizer and chat template<|im_end|>): three answers ran to 160
tokens without entail, and stopped at 46, 63 and 56 with it. Over 230 popular folders,entail checkfinds no
id past a tokenizer and would add a dropped end on transformers for 9 folders and on vLLM for 2; SGLang's
scheduler also matches the tokenizer's eos, so nothing is added there.
Changed
- Config coverage: a key the class does not take but the vocabulary maps and compares elsewhere (Qwen2.5 and Qwen3
writerope_scaling: null) is a pass that names it, not an unknown; before, it was 240 of the 394 unknown lines
in 114 healthy runs, on models where nothing was in doubt. The misspelling rule compares a top-level key with
top-level vocabulary keys only: a field that lives insiderope_scaling/rope_parameters(factor,
beta_fast,mrope_interleaved...) is no target, so SmolLM2'srope_interleavedis an unread key, not a
misspeltmrope_interleaved(which 1.1.0.dev had called broken on all three engines). Checked over the 558
distinct keys of 230 popular configs: no top-level key is within the rule's distance of a vocabulary key. - A decision recorded once for an owner (
enforce(once_for=...)) also works when the owner is a value (a folder
path, a (class, name) pair): before, such an owner was silently not remembered, and the tokenizer's pass was
recorded at every one of SGLang's tokenizer builds; the same config's coverage decision is now recorded once per
process instead of at every build (vLLM and SGLang build the same config several times). - Load cost: a model folder's vocabulary sources are read once per process and once per machine (a stamp of the
watched files keys an in-process cache andentail_logs/vocab_sources.json), tokenizer.json is counted only
when a second tokenizer source exists to compare with, and the tokenizer's highest id comes from its added-token
table instead of a fullget_vocab(). On Qwen3-4B the library's time at load went from 264/872/999 ms
(transformers/vLLM/SGLang) to 32/185/94 ms once the cache is filled, 150/289/60 ms on a machine's first run; over
102 healthy runs the share of load time is 1.2% at the median, 7.8% at the 90th percentile and up to 37% on toy
models that load in under a second (testbed/results/m15/E2_SUMMARY.md,E2_RECOST_SUMMARY.md).
Older facts and files still read (`REA...
entail 1.0.2
What an external review of 1.0.1 found, and what its readers will ask first.
Fixed
- Config keys (
load.keys_taken): a key the class does not take counted as renamed when its value appeared in any
field the class knows, so an unread key holding a 1 or atruewas silently taken as read, and a misspelt
vocabulary key with such a value (tie_word_embedding: true) escaped the misspelling rule. Now a misspelling of a
key the vocabulary maps is lost whatever its value holds, and a value counts as evidence of a rename only when it
is distinctive (a float, a string of four characters or more, an integer of 256 or more). The names a class
renames on the way in (attribute_map, GPT-2'shidden_sizeforn_embd) are taken by name, and the token ids
anduse_cachethatGenerationConfig.from_model_configreads off any config are listed as read elsewhere. The
stricter rule reports more keys as read by nothing entail knows (head_dim,max_window_layerson classes that
keep them as plain attributes, which their model code may read): on the 230 popular configs, 165 carry one such
unknownline instead of 92.ENTAIL_QUIET=unknownkeeps it off the console. entail checkbuilds the config withtrust_remote_code=Falseandlocal_files_only=True: a model with its own
code fails at once instead of prompting (which waited 15 s per model where there was no terminal), and nothing is
fetched.
Added
ENTAIL_QUIET=unknown: non-blockingunknowndecisions go to the log and the record only; the console says so
once per process. 1.0.1 printed one such line per boundary that could not be decided (69 lines over 81 loads of
30 popular models).- README: what the package does to an environment (the start-up hook and how to remove it, the log folder and how to
turn it off, the engine versions each adapter was measured against, the import name); the Korean README carries
the 1.0.1 and 1.0.2 changes. - The research workspace behind the numbers - the measurement scripts, the result files, the theory, design and
roadmap documents the code cites - is public at https://github.com/wwoosshh/entail-research.
entail 1.0.1
False alarms found when 1.0.0 was run on 30 popular models, three engines each (81 runs: it never broke a
run and every output was identical, but 17 runs carried a report that was wrong; testbed/results/m10/E2_SUMMARY.md
in the research repository). Nothing in the vocabulary or the verdicts changes; what changes is what counts as
evidence at five boundaries, and how much is said.
Fixed
- Config keys (
load.config_keys): a key the model's config class does not take wasbrokenwhatever it was, and
popular models carry keys nobody reads (swiglu_limit,task_specific_params; 45 of the 81 runs). Now only a
misspelling of a key entail's vocabulary maps isbroken(rope_scaleforrope_scaling: nobody can read it, so
it is lost here). A key spelt right that the class does not take is decided where a consumer of its fact reads
it; any other unread key is oneunknownline that names the keys, blocking only in debug mode. On 215 unread
keys of 92 popular models the new rule flags none; the rolebench 15 case staysbroken. - Tied embeddings (
load.tie): vLLM 0.30 setstie_word_embeddingsto False when the checkpoint ships an
lm_head.weight, loads it and re-ties it when it equals the embedding; transformers 5.17 compares the two the same
way. 1.0.0 read the False as the loader's choice and reported Qwen3 0.6B, 1.7B and quantised exports (a stored copy
of the tied head) asbroken. Now the adapters read what the loader left in the model - the head sharing the
embedding's tensor, or its own - after its step (vLLM afterprocess_weights_after_loading, transformers at the
tie_weightscall that has the weights), and the checkpoint is only sampled. A copy satisfies a declared tie; a
different head is what those loaders run, so the declaration is reported false, and withuse_datathe head that
runs passes as the data used. The static check (entail check), which has no model in memory, compares the
checkpoint's head with the embedding byte for byte (observe.head). SGLang 0.5.20 ties regardless, and is
decided as before. - Chat template (
readers.HfTemplate):chat_template.jinjatakes precedence over the entry in
tokenizer_config.json, as transformers reads them; the entry is no longer a second declaration. A checkpoint
whose two copies differed by blank lines only was reportedbrokenat every request. An entry that differs from
the file beyond blank lines and trailing spaces is noted as not what runs. - Evidence for a repair (
load.attention,load.tool_parser): a mismatch the capability table knows only from
reading code (SGLang flashinfer's sliding window) is reported asunknown- "inferred from its code, not
measured; nothing is switched on it" - instead of switching the backend on it. A declared sliding window that
is not belowmax_position_embeddings(Phi-3.5, Phi-4-mini: 262144 over 131072) never binds and is not read as a
requirement. - SGLang KV contract under speculative decoding: the scheduler reserves draft slots ahead of the tokens, which the
contract reported as reserved and written slots disagreeing at every step. Speculative batches are now skipped
and said so once per process. - diffusers pipelines named by a hub id are checked from the folder in the local huggingface_hub cache; 1.0.0 checked
only local folders. - Weights the layout step cannot read (a conv1d in a hybrid model) are one line per reason, not one per weight (42
lines per load of Nemotron-H).
entail 1.0.0
A redesign. 0.3.0 was a set of checks and resolvers for cases that had been measured one by one; 1.0 is one
mechanism for all of them: facts that say what a value means, read from what declares them, compared where they are
used, under one policy and in one record. Everything listed was measured on the engines and versions in the README
(one RTX 4070 Ti); the README's "How it was measured" and "Known gaps" give the results and the limits.
Facts and verdicts
- A closed vocabulary of what a value means (version 4): layout and quantization, RoPE (with llama3's frequency
factors) and position frames, valid ranges and KV extents, model properties, prediction type and latent scale,
chat template, key coverage, reduction state, epochs, assumptions and precedence. Each fact carries where it came
from and how certain it is. - Five verdicts at every boundary:
pass,resolved,broken,refused,unknown.
Where facts come from
- Model folders and files:
config.json, the tokenizer's chat template, safetensors metadata, GGUF keys, diffusers
configs. - Manifests for files that declare nothing:
entail inferwrites a draft,entail pinmarks it reviewed; a manifest
next to the file, or in a folder named byENTAIL_MANIFESTS. - Which consumer honours which fact is a table with its evidence; only measured entries are used for repairs.
Where they are compared
- At load: attention properties against each backend, RoPE (a declared key the vocabulary cannot carry is reported
as not compared), config keys nobody reads, tied embeddings against the checkpoint, vLLM's weights after
repacking against the signatures of the steps that repack them, weights against the checkpoint file
(ENTAIL_SOURCE=1). - In containers: the KV cache contract on transformers, vLLM and SGLang; buffers read after an in-place write; CUDA
graphs and compiled code reused for inputs they were not made for. - Per request, on vLLM's OpenAI server: the chat template, the reasoning history and tool-call format a model
declares, request fields and template settings nothing reads. Where transformers applies a chat template - a
script'sapply_chat_template, SGLang's server - the template and the reasoning history; and SGLang's server
when it renders with a conversation template of its own. The core also has a rule for the context a prompt
needs against what the model declares; no engine adapter calls it yet. - In code, in debug mode:
@entail.boundarydeclares what each argument means; strided layouts, quantized values and
chunk-relative positions are converted where the reader needs them. - Image models on ComfyUI and diffusers: prediction type and latent scale against the sampler and the VAE, and whether
a LoRA reaches the model it is applied to.
Policy and record
- A mismatch is repaired first. What nothing can repair is reported as
brokenand the run goes on;
ENTAIL_ON_BROKEN=stop(orName=stopinENTAIL_FACT_POLICY) stops before any output.ENTAIL_POLICY=refuse
repairs nothing and reports everything. - Everything entail says goes to
entail_logs/in the folder a program starts from: a log and a JSON record of every
decision, from every process an engine starts.
When the output is still wrong
entail locatenames the first boundary that did not keep a fact, or says that every checked boundary held.diagnose.watchanddiagnose.comparecompare a layer with a reference on the same inputs;diagnose.propagating()
follows declared facts through tensor operations and names the one that made a fact untrue.pytest --entailruns each test that way, andentail_conditionsruns a test once per condition a model's
declarations put at stake.
For new code (experimental)
entail.frontend: role-typed values and keyword-only operations, checked when a program is traced; a Qwen3 decode
step written with it, lowered to torch, FlexAttention or a Triton kernel.
Changed from 0.3.0
- A mismatch nothing repairs is reported and the run goes on; 0.3.0 stopped. Set
ENTAIL_ON_BROKEN=stopto stop. - A checkpoint that declares nothing is reported as unknown. 0.3.0 judged the prediction type from the first model
call; a checkpoint whose marker was lost now needs a manifest. - A LoRA that reaches only part of the model is
broken(0.3.0: a one-line note), and a sampling setting the user
chose against the checkpoint's declaration is reported instead of passed over. - The fault injection, the layout ledger and the bookkeeping probe used to measure entail left the package, and with
themENTAIL_SEED,ENTAIL_LEDGERandENTAIL_PROBE.
Install: pip install entail-ai (the import name is entail). See the README for how it was measured and the known gaps.
v0.3.0
ComfyUI support: two sampling problems that ComfyUI finishes as a success.
A v-prediction checkpoint that lost its marker. ComfyUI decides v-prediction for SD1/SD2/SDXL checkpoints from one v_pred key in the file. A merge or conversion that drops it makes ComfyUI sample a v-prediction model as eps: coloured noise or black images, run marked as a success. entail reads the first model call of each sampling (the noisiest step, so no extra forward pass): an eps model returns the noise it was given, a v model does not. The model is then sampled the way it behaves, as a ModelSamplingDiscrete node would, and one line says so. A sampling node in the workflow that contradicts the model stops the run instead: an explicit choice is not overridden silently.
Measured on a real ComfyUI 0.34.1: NoobAI-XL-Vpred with the marker removed came out 67-102/255 away from the right images without entail, 12-20 with it (the rest is the zero-terminal-SNR setting, which behaviour cannot reveal). A v_prediction node left on in front of an eps checkpoint stops at the sampler (3/3; without entail a flat grey image).
A sampling node's schedule that outlives its workflow. ComfyUI 0.34.1's dynamic VRAM loader backs model buffers up by attribute path. So after a run with a ModelSamplingDiscrete (or similar) node, the checkpoint kept sampling with that node's schedule once the node was gone: another image, then black, until a restart, with or without entail. The other way round, a node used after a plain run silently got the plain schedule. entail keeps each schedule with the object that set it: the loader's backup goes back to its own object, and the first model call after the buffers change checks them against what the object's own setter registered.
Measured: both directions now match a fresh session pixel for pixel. The first-call check alone left 2-11/255, so the guard does the exact part and the check is the fallback.
Images are identical with entail on and off for eps and v checkpoints and for an Anima (flow model) workflow with a LoRA, at the same speed.
pip install -U entail-aiAlso: ENTAIL_SKIP=comfyui:install_buffer_guard leaves out single entries, to measure what the rest does without them.
v0.2.0
ComfyUI support: a LoRA that cannot reach the model it is applied to now stops the workflow before sampling, with the reason.
ComfyUI skips every LoRA module with no counterpart in the loaded model (one console line each) and finishes the run as a success, so a LoRA made for another base model silently does nothing. Measured on a real ComfyUI 0.34.1: an Anima LoRA in an SDXL workflow logged 840 such lines and changed the image by 0.8/255 (the matching LoRA: 35.2).
With ENTAIL=load, entail counts how many LoRA modules reach the model and the text encoder using ComfyUI's own key map:
- none: stops with what the LoRA declares it was trained for and which model it met
- some: one line, the run goes on
Measured: both wrong directions (Anima <-> SDXL) stop; 22 right pairings pass with no false alarm; images are identical with entail on and off.
pip install -U entail-aiAlso: entail doctor recognises ComfyUI when run from its folder.
v0.1.0
First release of entail: keep what a value means intact across LLM inference-stack boundaries.
pip install entail-ai
entail doctor
ENTAIL=load <your vLLM / SGLang / transformers command>- Resolvers: RoPE values given under their transformers-4 names after the config is built (
rope_theta,rope_scaling;from_pretrainedkeywords, attributes, vLLM--hf-overrides, SGLang--json-model-override-args) are written whereconfig.jsonwould put them. Attention backends that drop a declared model property are switched to one measured to honour it. - Checks: attention properties against backends, swallowed config keys, tied embeddings, vLLM weight layout/stride/transform, weights against the checkpoint, the KV cache contract (transformers, vLLM, SGLang).
- A start-up hook (
entail-autoinstall.pth) letsENTAIL=loadreach engine worker processes; it does nothing unlessENTAILis set.
Research prototype (alpha). Tested with transformers 5.12.1 / 5.17.0, vLLM 0.30.0, SGLang 0.5.20.