Skip to content

b10373

Latest

Choose a tag to compare

@github-actions github-actions released this 12 Aug 14:19

speculative: never-worse scheduler - verify-all cold start, robust fits

Field regression on a 5090 (coding 600 -> 350-400 t/s) exposed three
defects in the self-calibrating scheduler:

  • cold sessions ran admission on the head's raw pessimistic
    confidences; worse, the resulting truncation censored the very
    full-block rounds calibration needs, so the first fit was ~6000
    rounds away. Uncalibrated mode now admits everything (same
    convention as sglang's flat-SPS verify-all default): full fixed-depth
    speed from round one, and every round feeds calibration, putting the
    first fit ~1 long prompt away.
  • early small-window fits produced extreme temperatures (T=0.087 seen,
    crushing deep-position survival). Fits now require >=32 samples of
    each outcome class per position (identity otherwise) and shrink
    geometrically toward identity by window size.
  • the admission scan now demands a meaningful predicted gain before
    truncating (4% margin on the causal early stop, and full admission
    is preferred whenever it scores within the margin of an interior
    argmax): near-flat theta curves - easy content on fast GPUs - admit
    everything, encoding the measured lesson that trading accepted
    tokens for noise-level predicted savings always loses.

Also: profiled SPS curve forced monotone (Windows timing outliers put
size 1 above size 2), and the calibration cache version bumped to v2
so existing caches from the noisy-fit era are discarded on load.

Isolation on RTX PRO 6000 (GSM k8, fresh caches, 4 rounds): OFF 588
t/s mean vs sched+adapt 587 (parity, accept len 7.18 vs 7.08); cold
first round 600.6 (was 350-400 on the affected 5090 setup).

Website:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI: