Skip to content

b10372

Choose a tag to compare

@github-actions github-actions released this 12 Aug 12:06

speculative: self-calibrating scheduler with persistent cache

--spec-sts required a manual offline fitting pipeline; calibration is
now automatic. The server collects (confidence vector, realized
accepted length) pairs from full-block rounds into a ring buffer -
with an exploration round every 16th drafted round verifying the full
block, so scheduled truncation never censors deep positions out of the
data - and refits the paper-3.2.1 sequential cumprod-ECE grid search
in-process every 256 rounds (~3ms on the most recent 1024 rounds).

STS temperatures and the online cost model (t_fix, c_tok) persist to a
per-draft-model cache file (--spec-sched-cache, default 'auto' in the
llama cache dir), so later boots start warm. --spec-sts still accepts
fixed temperatures and disables auto-refit.

Smoke-tested: cold boot refits at 384/640 rounds and writes the cache;
warm boot loads it (sts=8, t_fix=6.2ms).

Website:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI: