Releases: sqliteai/warp
Release list
v0.6.7
The engine decodes exactly as 0.6.6 did — same logits, same container
format — and two things underneath it changed. The second model this project
ships numbers for is now usable the way K3 is: converted with a chat format
it can actually read, and served over HTTP rather than from the command line
only. And the automatic memory budget stopped being able to choose a value
ten times slower than a smaller one.
waste_memplan gained a field, so this is not a drop-in header. A caller
that only reads the struct is fine; one that allocates or copies it must
recompile. serve/engine.py mirrors the new layout.
The chat half came out of one support report — a Kimi-Linear container on a
16 GB MacBook where waste run worked, waste chat answered oddly, and
python3 -m serve printed its banner and exited. The budget half came out
of an experiment that failed: collapsing K3's experts down to one, which
does not work and is recorded below, is what made a small enough working set
to expose the ceiling.
Added
-
A second prompt format for the server (
serve/chatfmt.py,
docs/SERVE.md). At startup the richer format is asked for first: XTML
when the container's tokenizer carries<|open|>,<|sep|>,<|close|>
and<|end_of_msg|>as single tokens, and otherwise the container's own
chat.json— the same four prefix/suffix stringswaste chathas always
read. A container is now addressed identically over HTTP and on the
command line, and a hand-editedchat.jsonis honoured by both.Plain means plain: system / user / assistant turns, blocking and
streaming, with the stop token taken from the template's assistant suffix
rather than guessed. Everything four strings cannot express is refused
with a 400 naming the field —tools,reasoning_effort, an image part,
a tool-result turn. None of it is dropped silently, on the same reasoning
the effort mapping already followed: a server that ignores
reasoning_effortreports a different amount of reasoning than it did.chat.jsonis validated harder here than by the CLI's reader, which has a
person watching and an interrupt key. Serving requires anopen, auser
turn, and an assistant suffix carrying a control token — without the last
one every reply runs tomax_tokensand reportsfinish_reason: "length",
which reads as a broken model rather than a broken template. And every
<|…|>in the file is resolved against the real vocabulary before the
format is used at all, because markup the tokenizer does not have encodes
as ordinary text and the model then reads its own turn structure as prose
and answers anyway, plausibly and wrongly. The rendering keeps the split
that makes the XTML path safe: the template's strings go out as markup
segments, the caller's content never does.Measured on a Kimi-Linear container: a multi-turn conversation
round-trips, streaming deltas arrive, and both refusals come back as 400s
namingtoolsandreasoning_effort. The serve suite is 211 checks, up
from 174. -
tests/sweep.ctakes atopk=arm, lowering only, since the scratch
is sized at load from the manifest'stop_k. One load with the arms
interleaved, which is what made §56's curve trustworthy after two earlier
attempts were spoiled by comparing arms across process lifetimes — the
failuresweep.cexists to prevent and that §32 and §33 already record. -
tools/convert.pyinstalls achat.jsonper architecture, with
examples/chat-kimi-linear.jsonalongside K3's. The architecture was
already recognised; it simply had nothing to install and said so, which
left every Kimi-Linear container to be finished by hand — and the obvious
hand fix, copying the ChatMLexamples/chat.json, is wrong in a way
nothing reports:<|im_start|>is not in Kimi's vocabulary and encodes as
six ordinary tokens.So the converter also refuses to install a template whose markup the
release's tokenizer does not carry, naming the missing markers. Only
when there is a specials list to check against: a release without
tokenizer_config.jsonis no evidence, and refusing on none would be
worse than the unconditional copy it replaces.
tests/test_convert_chat.pycovers both claims per architecture.
Fixed
-
The automatic budget could choose a value 10x slower than a smaller one
(src/waste.c, LEARNED.md §57). It stepped down a whole
working set at a time until the total fit under 7/8 of usable RAM. 7/8
of 64 GB is 56, inside the 46-52 GB band §39 measured as an eightfold
collapse: the engine stays inside its budget, the machine does not, and a
cache hit becomes a page fault.It had stayed harmless by luck. K3 at top-16 asks for 80.77 GB, cannot have
it, and the step-down lands onfloor + 1x= 46.39 GB — the measured
optimum, reached for the wrong reason. Lowernum_experts_per_tokento 8
and a token's working set halves, three multiples fit under the old
ceiling, and the default took 54.77 GB and ran at 0.08 tok/s against 0.77
at 46 GB, with a higher hit rate and a lower RSS — which is what
paging looks like from inside the process.The ceiling is now 3/4, measured rather than assumed: 46 GB is the
largest budget on this machine known to be on the good side and 52 the
smallest known to be on the bad one. K3 at top-16 still resolves to
46.39 GB, Kimi-Linear is untouched, a 128 GB machine still gets the full
floor + 3x, and K3 at top-8 now resolves to 46.18 GB and 0.88 tok/s
on the path a user gets by typing nothing. -
The server refused to start on any container without XTML markers
(#34).ChatServer.__init__
resolved the four control tokens and let theEngineErrorout, so the
process died before binding a port — taking/health,/v1/modelsand
/v1/completionswith it, none of which need a chat format at all, and
reporting only that<|open|>had come out as five tokens. The reason is
now held rather than raised, and a container with neither format serves
everything except/v1/chat/completions, which returns 400 with
code: "unsupported_chat_format"and both reasons — no XTML markers, and
what was wrong with thechat.json.
Changed
-
waste_memplanreportsworking_set_bytes— one token's expert
traffic,top_krecords per MoE layer. Callers recovered it as
(recommended_bytes - floor_bytes) / 3, which stopped being true here:
recommended_bytesis now capped at the container's whole expert set,
because a cache holding every expert cannot be improved by growing. On an
ordinary container nothing moves — K3's bank is 952 GB against a 3x working
set of 52 GB — but the two are no longer three times apart in general, so
the quantity the rule is built on is reported instead of re-derived.
waste plan --jsoncarries it. -
README.mdfigures re-measured on this commit: K3's floor 29.06 →
29.19 GB, its default budget 46.25 → 46.39 GB, Kimi-Linear's floor 1.28 →
1.32 GB and its decode 10.65 → 10.62 tok/s. Drift from earlier commits,
not from any change here; the K3 decode range is left as it was, because
its low end was measured under conditions this pass did not reproduce.
Measured and not adopted
-
Merging a layer's experts into one, or into sixteen, is not a
compression of a MoE — it is a deletion of it (§53-§55). Built because it
was asked for: 982 GB becomes 30 GB, decode goes 0.60 to 1.58 tok/s, and
the model emits<|close|>forever. No weighting helps, and the reason is
geometric — distinct experts are mutually orthogonal (cos 0.0006), so
their average has1/sqrt(E)of their norm and is 99.8% orthogonal to
every one of them. A gain sweep confirms it from the other side: scaling
the merged expert to zero is better than using it. Clustering into 16
does not rescue it, and pruning to the 16 busiest — which beats every
merge — still answers that the capital of Italy is Paris. The tooling
stays on thek3-minibranch rather than main. -
Fewer experts per token is worth taking; the default still does not
take it (§56).num_experts_per_token8 instead of 16 is 1.49x at
KL 0.037, with top-16's greedy continuation reproduced; top-4 is 1.78x,
keeps the right argmax, and stops following the prompt within a few
tokens. It is a quality trade, so it is documented inREADME.mdand left
to the caller rather than changed under anyone. -
§4D re-priced: batching is worth less in this regime, not more (§58).
docs/EFFICIENCY.md§1's 1.62x reproduces exactly at top-16 (1.60x)
and falls to 1.28x at top-8 — batching takes its gain from the I/O and
truncatingtop_khas already taken half of it, so the two overlap rather
than compose. The ceiling holds for batching across independent streams
too, which §4D never separated from grouping within one:vq_applycosts
one pass per (token, expert) pair however they are grouped, and that is
64.2% of a step, so no scheme beats 1.56x. -
The trunk's contextual sparsity is real and unusable (§59, §60). A
quarter of the shared expert's intermediate channels carry 99% of the
layer's output — but the trunk is 49.9% attention against 16.5% FFN,
so perfect sparsity where the technique fits is worth 1.14x. And the
channel identity is near-random across tokens: Jaccard 0.196-0.271 against
a 0.143 chance baseline, 2-4% of the set common to eight consecutive
tokens, 70-83% of all channels appearing in at least one of them. No
static core to prune, no cheap prediction to make. Taken together these
put the honest ceiling from 0.88 tok/s at about 2x, and name what the
rest would cost: 6.1x of bandwidth efficiency, which is a rewrite of the
forward pass, and 11.2x of bytes per token, f...
v0.6.6
The engine decodes exactly as 0.6.5 did and no container format moved. Two
things make it a tag. A host can now say which CPUs the compute pool runs
on, which is the first answer this project has to a machine whose cores are
not interchangeable. And tools/diskbench — the tool whose entire job is to
certify the storage a container will be streamed from — was measuring the
page cache on Linux, so it had been answering that question wrong for
everyone who is not on macOS.
Callers must recompile against this header. cpu_list was added to
waste_cfg between n_threads and cache_policy, so a caller built
against 0.6.5's struct reads cache_policy and every field after it from
the wrong offset. serve/engine.py's mirror moved with it; an out-of-tree
ctypes or FFI binding has to move too. The library is 0.6.x and promises no
stable ABI yet, but a silent misread is worth the sentence.
Added
-
--cpus LISTon the CLI and the server,waste_cfg.cpu_listin the
API,WASTE_CPUSin the environment — a Linux-style cpu list (0-5,
0-2,6-8) that the compute pool binds to. Linux and Windows; macOS has
no call that binds a thread to a core, so a list there is refused with
WASTE_E_UNSUPPORTEDrather than ignored.It exists because on a machine whose cores are not interchangeable,
placement is worth more than the thread count.
Issue #23 measured
Kimi-Linear-48B on a Ryzen 9 9900X — two 6-core CCDs, separate 32 MB L3 —
and found six threads on one CCD 16-25% faster than the same six split
across both, at identicalbytes_readand identical hit counts. Handing
those six threads all 24 CPUs to migrate between costs a further ~10%.
Not reproduced here: this repo has no multi-CCD machine, and the numbers
above are the reporter's. What is checked here is the mechanism —
tests/test_cpus.creads back every participant's affinity mask.No default changed. The engine still names no CPUs and leaves
placement to the OS.docs/LEARNED.md§47 measured the tempting default
— cap the pool at the fast cores — as a 25% gain on Kimi-Linear and a 34%
loss on K3, so it stays a switch.--threads 0with a cpu list means one
thread per CPU listed; an explicit--threadsstill wins. The thread
that calls into the engine is bound too, on its first parallel region,
because it is one of the workers; the expert cache's reader threads are
not, because they are blocked inpreadrather than competing for a
core.docs/ENGINE.md, "Thread placement", has the rest. -
The converted K3 container over BitTorrent, in
README.mdahead of
the conversion recipe. The default conversion is deterministic, so the
982 GB directory is byte-identical for everyone who produces it, and
nearly all of what the recipe costs — a 1.42 TB source download, 4.7
hours, and staging storage that has to exist before it can be freed — is
paid to reproduce a fixed artifact. The torrent's own piece hashes verify
it as it arrives. Converting from the published weights stays documented,
for anyone who would rather not trust a third-party copy.
Fixed
-
tools/diskbenchmeasured the page cache on Linux, not the disk
(#22,docs/LEARNED.md
§49). It documented itself as reading with the cache bypassed, and did
neither:nocache()had an#ifdef __APPLE__body and nothing else in
it, andO_DIRECTappeared nowhere in the file. Against a Samsung 970
PRO on Gen3 x4 it reported 44.67 GB/s sequential and 65.72 GB/s random
over a 3.94 GB/s link — 11x and 17x the ceiling. Bypassed: 3.15 and 3.33
GB/s, saturating at two threads, which is what that drive should do.This is LEARNED §14 in the one place §14 did not reach. The engine's own
bypass was written blind and fixed on 2026-07-28; the tool that exists to
characterise the engine's I/O kept reading RAM, which means §46's standing
rule — rundiskbenchand divide before claiming anything is disk-bound —
returned a fiction on Linux for that whole window. No published number
moves: everydiskbenchfigure indocs/GATES.md,docs/EFFICIENCY.md
and LEARNED §44/§46 was measured on macOS, whereF_NOCACHEdid work.The flag alone is not enough, for the reason
bank_openalready knows:
O_DIRECTis accepted at open and refused at transfer (tmpfs does this),
so a bare flag turns a refusing filesystem into a table of zeroes with no
cause given. It followsbank_openinstead — probe with one aligned
transfer, fall back to a plain open plusPOSIX_FADV_RANDOM, and label
every row, because a bench that quietly measures something else is worse
than one that says it could not. The write is bypassed too, and that is
not symmetry for its own sake:F_NOCACHEstops new pages being cached
but does not evict resident ones, so a buffered write leaves the file in
the UBC and every read row below it reports RAM — 8.07 GB/s sequential
with the write bypassed against 26.04 GB/s with it buffered, 1 GB file on
an M5 Pro. Also fixed alongside: a sub-page record rounded to zero and was
divided by, and a failed sequential read ended the loop and silently
shortened the row.Reported, diagnosed and fixed by fab2s. Verified here on macOS as
unchanged within noise; the Linux figures are the reporter's, on hardware
this repo does not have. -
diskbench's tok/s column answered for K3 whatever was being sized.
The derived column carried 12.5 GB/token in its format string — K3's
figure — so on a 48B model at a measured 1.61 GB/token it was ~8x off,
and silent about the assumption, which is what made it a trap rather than
an approximation. It is now the fifth positional argument with no default:
without it the column is not printed. A tool cannot derive bytes-per-token
from a scratch file — that number belongs to a container, andwaste benchalready reports it.docs/GATES.md's Gate H table keeps its
"tok/s @12.5 GB/token" header: that was a K3 decision, and the figure is
stated in the header rather than hidden in a format string.
v0.6.5
Nothing in the engine changed: a binary built from this tag decodes exactly
as 0.6.4 did, and no container format moved. What changed is on either side
of it — converting K3 on the disk you actually have, and how long a reply
the server gives a client that never asked for a length.
Added
-
tools/convert.py --reclaim {off,dry,on}— deletes each source shard
once its last consumer has published, so peak staging is the container
plus the shards still owed rather than the container plus all of them. On
K3 that is the difference between 1.42 TB of staging beside a 982 GiB
container — two disks — and one. It is safe because every tensor has
exactly one consumer, so a shard whose last consumer has finished is never
opened again.--reclaimalso runs the trunk pass first: it consumes
every non-expert tensor, and while it ran last almost no shard was ever
spent. The reordering is neutral —off,dryandonproduce
byte-identical containers.Off by default, and not reversible. A reclaimed shard has to be
downloaded again, andverify_container.pyloses the comparison against
source for good — which is whypipeline.shnow converts one probe layer
and round-trips that while the checkpoint is still whole, before
converting the rest. Stages renumber to six; stage 4 passes--skip-trunk
because stage 2 built it, a K3 trunk being hours to do twice.It refuses before deleting rather than during:
--experts, a
container inside the checkpoint or the reverse, a shard that is neither on
disk nor already reclaimed, and a bank that is not a whole bank —
bank_is_soundwalks the records by their own block counts for 48 bytes
each, which catches a bank truncated by a kill, a full disk, or a torn
rename. Releases are recorded in<src>/.reclaimed, fsynced ahead of the
unlink, because a name recorded but not deleted costs nothing while the
reverse is indistinguishable from an unfinished download.
K3.md has the refusals and the ledger discipline.Proven on a copy of a real Kimi-Linear checkpoint rather than on stubs:
--layers 1,2 --reclaim ondeleted exactly the one shard the dry pass
named (4.7 GiB, 92 -> 87 GB), left the other 19 unchanged, and wrote a
container that reads back 256 records with 0 problems. The second run over
the now-incomplete source refuses — pointing at--skip-trunk— and
deletes nothing while refusing.
Changed
-
The server's default
--max-tokensis 4096, was 512. Clients mostly
do not send the field; Open-WebUI does not unless you set it in the
model's advanced parameters, so the server default was every reply's
length, and a reply that ends at the cap is indistinguishable from a model
that stopped on its own.--ctxnever lifted it and could not: the limit
is clamped to the room left after the prompt, so raising the context only
ever lowers the cap. This is a behaviour change, not a fix — a
deployment that relied on 512 to bound per-request cost should now pass
--max-tokensexplicitly.Engine.generate's own default is untouched:
the server always passes the value, so the two never meet. -
SERVE.md documents Open-WebUI — the base URL, why no
compatibility mode is needed,--host 0.0.0.0for a client in a
container, the background title and tag requests that queue behind the
reply on the lock every generation takes, andreasoning_content, which a
client that does not know the field renders as a server that has stopped.
Fixed
- A resumed
--reclaimrun believed the download ledger over the disk.
ST.have()readsfetch_weights.sh's.download-state, so a shard that
the run had itself consumed still read as present and both refusals were
skipped. Absence is now asked of the filesystem, andhave()only asked
whether a shard that is there finished downloading. The pipeline test
found it; the file-backed test stub could not have, and now mirrors
mxfp4.ST, trap included.
v0.6.4
A container format addition and the measurements that price it. Nothing
changes by default and this is not a speedup release: no shipped container
uses the new format, the new switches are off, and an engine built from this
tag behaves exactly as 0.6.3 did unless it is asked otherwise. The reason to
tag it is that fmt 8 is now allocated and public, and a format code that
lives only in a working tree is one somebody else reassigns.
Added
-
WQ_VQ4P, fmt 8 — 4 residual VQ stages of 64 entries with 6-bit
indices packed four into three bytes. Same 3.00 bits/weight as VQ3R, same
record size, same blocked index layout; what changes is that a 64-entry
stage table is 64 bytes, which is one NEONvqtbl4q, where VQ3R's
256-entry table is sixteen vector registers on a machine that has
thirty-two and cannot be held at all. That is why the VQ3R gather is
scalar. FORMAT.md specifies the packing;
LEARNED.md §41 has the derivation.It is a distinct fmt rather than a manifest flag because a VQ4P payload is
byte-for-byte the size of a VQ3R one, so a reader taking the three bytes
for three one-byte indices would decode silently and wrongly. The engine
additionally refuses a record whose fmt byte disagrees with the manifest's
index_bits, which is the only read-path behaviour that changed. -
tools/convert.py:--entries,--index-bits,--stages 4|6.
Default output is unchanged VQ3R. The parameters travel in the job tuple
rather than a module global: the worker pool usesspawn, so a global
would have written every layer at 256 entries without saying so. -
WASTE_XPAR— one task per routed expert instead of one per row
range, off by default.WASTE_XPAR_BATCHbounds how many experts are
held at once,WASTE_P6_CHUNKsizes the VQ4P apply, and
-DWASTE_P6_SCALARbuilds the kernel's portable path, which is
bit-identical to the NEON one rather than merely close — §43 explains why
an int8 lookup table raises that bar. -
tools/lutbw.c,tools/lutmt.c— the two benches the sections below
rest on: kernel throughput against a working set from 4 MB to 1 GB, and
the same kernels driven through the engine's ownwaste_parallel_for.
Measured
Numbers on this commit, medians of repeated runs, each container on the same
storage as the baseline it is compared against.
- The kernel is 3.88x and that survives a gigabyte of index stream, so it
is not a cache artefact (§46). - In the engine it is 1.24x on Kimi-Linear best-configuration against
best-configuration, and 1.74x against what the engine does untouched;
1.09x on K3 (§46, §47). - Quality costs +2.7% perplexity on Kimi-Linear (10.937 -> 11.237), all
of it from the smaller codebook. Quantizing the runtime lookup table to
int8 — which is what makes a byte shuffle possible at all — measured free,
because the scale is per 32 vector positions rather than global (§41).
The gap between 3.88x on a bench and 1.17x in place is the useful part, and
§47 is the answer: 3.88x is a single-thread ratio, this machine is 6
performance cores and 12 efficiency cores with the pool taking all 18, and
the fast kernel is the one an E-core straggler hurts. On K3 the same
settings invert — six threads are 34% worse than eighteen there. No
default was changed because every one of them is right on one model and
wrong on the other.
Not adopted
- 6x16 codebooks (FAISS FastScan's shape) are a third faster than 4x64
and cost 18% more reconstruction error against 4x64's 7.5%. Speed was not
the binding constraint (§41). - Expert-parallel MoE as a default. Worth 1.24x on Kimi-Linear, a
regression on K3: the batch that gives it parallelism is the batch that
barriers the read-ahead, and no batch size wins both (§44). - Capping the thread pool at the performance cores. A 25% win on
Kimi-Linear and a 34% loss on K3 (§47).
v0.6.3
Fixed
-
The automatic budget sized against the host's RAM inside a container
(#14).waste_physical_ram()
issysconf(_SC_PHYS_PAGES)on Linux, which reports the host'sMemTotal
from inside a cgroup that is allowed a fraction of it, so a--budget-less
open resolvedfloor + 3xagainst memory the kernel would never hand over:
K3 in a 32 GiB cgroup on a large host asks for ~80 GB and is killed. Unlike
the paging cliff of LEARNED.md §16 this has no gradual
form and no cache policy softens it. The ceiling is now
min(physical, cgroup limit)— the smallest finitememory.maxor
memory.highacross the cgroup and its ancestors, since the limit is
hierarchical — and the rest of the resolver is unchanged. See §40.Current pressure (
MemAvailable,memory.current) was considered and
deliberately left out: a budget is resolved once and held for a whole run,
so bounding it by an instantaneous sample would make the same command on
the same machine two different runs. Whether it should trim the working-set
multiplier instead is #14, still open. -
tools/convert.pyspawned--jobs× cores threads
(#13, contributed by
@andrewwhitecdw). torch sizes its intra-op pool fromos.cpu_count(), and
the native VQ encoder readsnthreads=0as "every core" (capped at 64), so
N worker processes meant N×cpus threads competing for the machine: on a
224-core box--jobs 8spawned ~1792 threads and the codebook phase ran
~20x slower than the measured baseline. Each worker is now capped at its
fair share,cpu_count // jobs, for both pools — set in the parent before
the workers spawn, since that is the only point torch reads it, and with
setdefault, so an explicitOMP_NUM_THREADSstays the caller's. -
The trace simulator modelled a different cache than the engine.
tools/routing_stats.py simulatekept a frequency count across evictions
thatec_claimresets and sampled 32 victims whereEC_SAMPLEis 16:
against the same trace it read 36.6% where the engine measured 30.4% —
optimistic, plausible and wrong. Both constants now come fromecache.c,
which brings it within 1.5 points across a 0–30% range, andtests/run.sh
asserts the agreement rather than remembering it. Two smaller things went
with it: the route dump writes the absolute position of the token each row
belongs to, so readers stop re-deriving token boundaries from where the
layer index wraps (a heuristic that is simply wrong on the chunked path,
where rows group by layer), andsimulate --datatakes a container, whose
manifest states the record size the engine actuallypreads. See §37,
The simulator was modelling a different cache.
Added
-
waste_usable_ram(): physical RAM, or a smaller cgroup-v2 limit when one
applies — what a budget of 0 sizes against, and what an embedding host
should size its own ceiling from.waste plan --jsonreports it beside
physical_ram_bytes, which stays what it always was. -
tests/sweep.c, a one-process measurement harness. It loads a
container once and runs the arms back to back, interleaved, resetting the
session and clearing the expert cache between each — a warm cache would
hand the second arm the first one's work and measure the order instead of
the setting. Kimi-Linear, two arms, three repeats: spreads of 2.6% and
0.5%, where nine paired runs of the same comparison across processes
spanned 0.79x to 1.79x. The variance was the harness, not the feature. On
K3 the deterministic columns come out exact and the clock still drifts,
which is the machine's memory system rather than the process — what the
harness buys there is that the noise is visible as noise. §38. -
docs/TECHNICAL.mdandexamples/. The measurement tables move out of
README.mdinto TECHNICAL.md, andexamples/carries three compilable
programs against the public header —api_plan.c(budget arithmetic
without loading),api_text.c,api_vision.c— with a README that walks
through them.
Changed
-
§4's cache floor still reproduces exactly, and has stopped binding.
The oldest load-bearing measurement here — below one token's working set
the hit rate is zero, not low — was re-measured across four cache sizes in
one process, two repeats, hit rate and bytes read identical to the digit
across both:budget expert cache slots hit decode 32 GB 3.32 GB 287 29.1% 0.56–0.58 tok/s 46 GB 17.32 GB 1498 36.2% 0.63 tok/s 52 GB 23.32 GB 2018 38.4% 0.07–0.09 tok/s 58 GB 29.32 GB 2537 41.3% 0.07–0.08 tok/s The 29.1% at 287 slots is not a refutation of §4: with the lookahead off
the same 287 slots give 0.0%, exactly §4's zero. What breaks it is that
a speculative record has to survive one attention rather than one token, so
a cache far too small for a token's working set is ample to hold six
experts. A 3.32 GB cache is now within 10% of a 17.32 GB one, which means
the premise the default resolver is built on — that cache is only worth
buying in whole multiples of a working set — no longer holds. The
resolver is unchanged: that is a decision, not a measurement, and it is
GATES.md Gate 7, open. The cliff is exactly where it was,
and the last two rows say what it is — throughput falls eightfold while the
hit rate rises and the bytes read fall, because the engine is inside its
budget and the machine is not. §39. -
0.6.2's "total bytes read unchanged" for the router lookahead was the
harness. Measured with both arms starting from an identically cleared
cache, the byte economics depend on the cache size: 6.6% fewer bytes at
1498 slots, 8% more at 287, where speculative records are evicted
before use often enough to be re-read. It is a prefetch at small caches and
a scheduling change at large ones, and 0.6.2 measured only the large end.
The feature and its default are unchanged. §38, §39.
Section numbers refer to docs/LEARNED.md, which carries the reasoning and the full measurements behind each entry. The two §37s and two §33s there are noted in that file's preamble and cited by title where it matters.
Full Changelog: v0.6.2...v0.6.3
v0.6.2
Fixed
waste infoandwaste runcrashed on K3 on every x86 build
(#10). The tensors the
loader skips — the vision tower, and anything outsidetensor_prefix—
keptgroupat 0, and the row-scratch sizing divided by it. The
architecture decided what that meant: arm64'ssdivanswers 0 and the
run continues, x86'sidivraises#DE.waste planwas unaffected
because it does not load. §37.WASTE_Q8=0could not load a 4-bit trunk
(#6) — that is, any
container a defaulttools/convert.pyrun produces. The dequantizer
read one byte per weight, true of Q8G alone, while catching every
quantized format. It now decodes throughwaste_deq_row, the one place
that knows all three widths. The same lines also predatedwaste_f16's
subnormal fix and flushed group scales below 6.1e-05 to zero.embed_tokensstays on disk underWASTE_Q8=0, as it does
otherwise: 7.93 → 6.52 GiB of peak RSS on Kimi-Linear, identical logits.
The f32-equivalence check now differs from the default path in the
storage width alone, which is what it claims to compare.
Added
- Router lookahead in the decode path. At the end of a MoE layer, once
its reads are consumed and the disk is about to idle through the next
layer's attention, layer L+1's router runs on layer L's hidden state and
issues speculative reads for its top 6. Demand hit rate 14–19% → 38–40%
with total bytes read unchanged (254.2 → 254.5 GB): the records were
going to be read anyway, and only when changes. Nine paired runs,
median 1.17x.WASTE_LOOKAHEAD=0disables. §34, §35. WASTE_MLOCKwires the trunk and the expert cache;WASTE_MLOCK=cache
wires the cache alone. Off by default — Linux'sRLIMIT_MEMLOCKis
commonly 8 MB. Wiring the trunk is worth 3x in the transition zone around
52 GiB and nothing below it; it does not move the knee. §30, §31, §32.
Changed
tests/run.shgenerates its own PyTorch oracle from the container
under test (16.9 s) instead of diffing against a shipped fixture, which
can only ever be valid for the container that produced it: expert
codebooks are k-means, and the same seed on a different--device
trains different books — one layer of 26 moves the logits by 1.24
against a 1e-3 threshold. The fixture remains as the fallback whereuv
is absent, with its provenance recorded beside it.
(#7), §33.tools/make_test_container.pyemits what a real conversion does: a
Q4G/Q8G/F32trunk rather than Q8G throughout, and--prefixfor a
container whose tensors are not all under itstensor_prefix. Both are
shapes the suite could not previously reach, and both had a live bug
behind them.- Checks that cannot run now say why instead of reporting a refusal as a
divergence:WASTE_Q8=0on K3 wants 211 GB of f32 trunk on a 64 GB
machine, the oracle prompt is Kimi-Linear's, and a cold hotlist run that
already missed nothing demonstrates neither outcome
(#5).
Measured and not adopted
- Cross-layer prefetch from
next_layer_top— gated before building
and refused: 29.0% recall against a 60% break-even. §29, revisited and
superseded by the lookahead above in §34. - The lookahead in the prefill path. Built, measured, removed: a chunk
layer's disk is busy continuously, so a prefetch there does not move a
read into idle time, it moves it in front of another read and pays an
eviction for it — 7% more bytes. Decode keeps it. §36.
Section numbers refer to docs/LEARNED.md, which carries the reasoning and the full measurements behind each entry.
Full Changelog: v0.6.1...v0.6.2
v0.6.1
v0.6.0
Full Changelog: https://github.com/sqliteai/waste/commits/v0.6.0