Repository navigation
v0.2.1
Adds asynchronous whole-word generation to the offline SmolLM2 companion while
preserving the statistical and reranking APIs and SQLite formats. Generation
uses protocol 2, so the matching v0.2.1 workers are required. Model and tokenizer
bytes are unchanged.
The generation request carries context, prefix, session identity, up to three
excluded instant words and a result limit of up to three. Search uses eight
beams, up to eight tokens per word and 64 forward evaluations. Results are
normalized whole words with a probable following boundary. Words need not
belong to the statistical vocabulary. Empty results are valid.
The inference/reset reply deadline increases from 500 ms in v0.2.0 to two
seconds. Startup remains bounded to 30 seconds. Failure requires explicit retry. No network access is needed at runtime.
Generation is not automatically production-qualified. Corpus comparisons are
regression tests with unknown pretraining overlap. Existing reranking quality
regressions remain documented in the historical qualification reports.
Prepared CI bundles are unsigned; platform applications sign their embedded
workers through their own packaging workflows.
This companion release does not publish the Switchify PC desktop RC.
The release workflow validates all platform archives, checksums, database smoke tests and 1,000 warmed generation queries on Windows and macOS. Executables are unsigned for embedding in platform applications. Model weights remain separately checksum-pinned upstream inputs.