v0.8.18
-
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b10276, picking up Qwen3-TTS model-loading
primitives, explicit bundled-MTP loading, automatic model-specific token
suppression, and recent model, multimodal, speculative-decoding, backend,
and runtime fixes. Regenerated matching Dart FFI bindings, migrated model
loading to llama.cpp's load-mode ABI and the penalty sampler to the new
vocabulary-sized ABI, and refreshed thellamadart_llama_cpp_flutterApple
SwiftPM checksum. Speech generation is not yet exposed through the public
Dart API. -
Updated the default LiteRT-LM runtimes to
leehack/litert-lm-native@v0.15.0-native.3and
@litert-lm/core@0.15.0. The native artifact includes a corrected v0.15
streaming callback bridge and an Android Dawn rollback for Mali-G715 GPU
device loss; incompatible callback runtimes now fail safely before
generation, and concrete macOS app, framework, and cache libraries take
precedence over process-linked assets. -
Disabled automatic WebGPU fetch-backed model loading by default. Streamed
loading remains the safe default; controlled range-capable deployments can
opt in explicitly. -
Updated WebGPU bridge assets to
v0.1.26(llama.cppb10276), refreshing
both WebAssembly runtimes while preserving the existing bridge API.