Skip to content

TensorFold 1.0.0

Choose a tag to compare

@ashhart ashhart released this 07 Oct 12:04
· 12 commits to main since this release

TensorFold 1.0.0

TensorFold 1.0.0 ships the native Zig engine and HTTP server. Checkpoint loading, tokenization, chat templates and inference run in the native process, with no Python engine required at runtime.

Supported models

The release includes these qualified model and hardware combinations:

Model Qualified platform Serving scope
Nemotron 3.5 Lightning Apple Metal, M1 through M5 Native text serving and its draft head
Nemotron 3.5 Lightning NVIDIA GB10, CUDA Text serving, greedy and sampled
Qwen3.8 Flash Next Apple Metal, M5 Ultra Greedy text serving from the checkpoint, without a Python kernel recording
GLM-5.3-Flash Apple Metal, two M5 Ultra Macs Paired text serving
Qwen3.5-2B Apple Metal, M5 Max Native text serving

Hardware support is specific to the combinations above. Qwen3.8-27B, Bonsai, Gemma 4, Qwen3.6 and DeepSeek-V4 are being ported for 1.0.x. Qwen3.8-27B passes drafted/plain equality, but remains outside 1.0.0 because its paired served decode and cold-prefill results are slower than Python 0.6.6.

Native runtime and server

Flash Next builds its native layouts from the checkpoint and uses embedded Metal sources. It no longer needs a Python recording before serving.

The server exposes OpenAI chat and completion routes, Anthropic Messages and token counting, tool calls, streaming responses, tokenization. Health, metrics and an optional dashboard expose server status. /v1/decisions checks requests the way 0.6.6 does; label scoring for each family follows in 1.0.x. API-key controls, prompt caching and request cancellation are implemented in the native server.

Metal engines support configurable idle keepalive through --keep-warm. A load-time probe selects a prebuilt packed-kernel library when macOS 26.3's runtime compiler rejects that kernel.

Release archives target macOS arm64 and Linux aarch64 for the GB10. They contain bin/tensorfold-native, runtime information and license notices. CUDA needs a compatible NVIDIA driver and the qualified kernel assets described in the deployment instructions; paired Metal serving needs its transport and peer configuration. See README.md for installation and RUNBOOK.md for deployment.

The shipped commands include --version, --help, capabilities --json, serve MODEL, and pull, models and info for checkpoints in the Hugging Face cache. Use serve MODEL --help and the capabilities output for supported flags and backends.

Concurrency

Nemotron runs concurrent requests in shared rounds on Metal and CUDA, and every stream equals its solo run. On an M5 Max, eight concurrent requests reach 436.7 tok/s together, against 262.8 for one. Flash Next and GLM-5.3 answer one request at a time in 1.0.0; their shared rounds follow in 1.0.x.

Context compaction

--compact-at auto|FRACTION turns on context compaction, which is off by default. A conversation that would overflow is compacted instead of refused. The server keeps the system prompt and the recent turns word for word, turns the older turns into a structured memory note, and updates that note at each later compaction. --compact-keep sets how much recent text stays, and --compact-memory DIR also keeps each note as a Markdown file. The same request gives the same compaction and the same reply.

CUDA in 1.0.0

On the GB10, greedy and sampled replies equal the one-token reference, seeds and every sampling rule behave as in 0.6.6, concurrent streams equal their solo runs, and health and metrics report device memory. Decode runs 1.01x to 1.10x the Python 0.6.6 engine on the same Spark. Other NVIDIA GPUs and the other families stay on the Python line for now.

Moving from Python 0.6

The Python 0.6.6 engine remains on the python-0.6 branch and the v0.6.6 tag. Its historical release entries remain in CHANGELOG.md. The native release's model table and capabilities define its supported options; older Python CLI features and model backends are maintained on the Python line. A pip install of 1.0.0 stops with directions instead of replacing a working 0.6.6 install.

TensorFold is Apache-2.0. Earlier code retains the bundled MIT notice, and model weights retain their own licenses; see LICENSE, NOTICE and THIRD_PARTY_NOTICES.md.

Contributors

Thank you to all 145 public contributors who have sent TensorFold a pull request, a measurement or a bug report. Their work shaped both the Python releases and the native engine:

@0mao0,
@321sssrt-bit,
@aditya1503,
@AdrianBinDC,
@akol1,
@Alexbob0,
@anvilsong,
@Arminova,
@b-ostrov,
@barelyworkingcode,
@Benjamin-Wegener,
@benthecarman,
@benwilson,
@BHCC2025,
@Bizuayeu,
@BlivionIaG,
@BobClawblaw,
@borodach23,
@Boscoeuk,
@boxabirds,
@brandonmmusic-max,
@bunnyfu,
@CerebralCoding,
@cesarswong,
@chadhurley25075-png,
@chaog992,
@Charlie-Louis,
@Chedrian07,
@chris247474,
@christrade215,
@crescit,
@cshintov,
@cwschroeder,
@DakotaTexas,
@danyo1399,
@Deesha08,
@Defilan,
@DevRico003,
@di37,
@drowzeys,
@ecohash-co,
@edurdias,
@eleqtrizit,
@ericlsimplifi,
@EugeneClaw,
@feni6,
@gbgbgbg,
@GDACONSULT,
@gecobattya,
@gilby,
@Gogo6969,
@gprot42,
@GraithSecurity,
@grantoverton,
@grearjake-star,
@greatyingzi,
@harrisonfriia,
@haxudev,
@heitke,
@hichaiuse,
@ivanfioravanti,
@jasontitus,
@jayleaton,
@jeffpeng3,
@jeidbugs404,
@jetnet,
@jkuepker,
@johnymoo,
@JordiPosthumus,
@JRaxworthy,
@jregan-beasley-smc,
@jschmied,
@juliankang4,
@kingjamez,
@kky42,
@lcgutierrez,
@LECYWZA,
@lijian1999,
@liumorrisclaw,
@LXD-8,
@m-naoki-m,
@mapamalu,
@mcclanahanaman,
@MESevenJourney,
@mgoldwasser,
@MiaAI-Lab,
@mikolaj92,
@millaguie,
@Mirrdhyn,
@Moutonc,
@MovieMaker93,
@mrpmorris,
@MV10,
@NeoAiLabs,
@Nipale-ai,
@nood-co1,
@nullburn,
@olexale,
@omar16100,
@optimisme,
@outcastofmusic,
@paragontasx,
@peacockesq,
@philip-pentatonic,
@plotarmordev,
@pmeenan,
@pulseandthread,
@quigles1977,
@rafafortes,
@raymondkpwong,
@robertpitt,
@RoscoeTT,
@salmanarshad321,
@samwang0041-star,
@sanjaibalajee,
@satindergrewal,
@scottleimroth,
@sethforprivacy,
@sfxnz,
@shantanugoel,
@simon-lin88,
@simonmd,
@spenchey,
@squarrier,
@ss-cong,
@styles01,
@SxMShaDoW,
@sxuff,
@taussoe,
@tfolkman,
@ThinkOffApp,
@Thotheris,
@tinyapps,
@tolewis,
@tomByrer,
@tonydehnke,
@tournierjc,
@tpischke,
@urtho,
@vcruz305,
@vinicius-symetrix,
@wojo,
@xjqx2z,
@Yuepixel,
@YvesLaRose.