You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Performance: on Node, each pinned ONNX payload file is hashed once per installed runtime instead of four times per load (prefetch plus three loader requests); unchanged files are recognised by device, inode, size, mtime and ctime, and new processes still verify every byte. Measured warm model load fell from about 2.3 s to about 1.3 s on macOS arm64. The cache no longer calls chmod on files that already have mode 0600, which would have advanced ctime.
Performance: Node ONNX sessions now use os.availableParallelism() intra-op threads. On a 4 performance + 6 efficiency core Apple machine this raised prefill about 25% (about 213-238 to about 290 tokens/s) and lowered decode about 12% (about 28 to about 25 tokens/s); the best value depends on hardware.
Performance: structured calls reuse an engine-private decoder state for their fixed system instruction (text-only, at least 32 tokens, no explicit reuse). Time to first token for a 115-token structured prompt fell from about 460-540 ms to about 130-160 ms on a hit, with identical output in the exercised cases. reuseCacheInfo() gains prefixEntries, prefixBytes, prefixHits and prefixMisses.
Performance: the constrained-decoding token trie uses typed arrays; building it for the 248,056-piece vocabulary took about 270 ms (was about 337 ms) and about 54 MB of external/array-buffer memory instead of about 230 MB of JavaScript heap.