You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This commit was created on GitHub.com and signed with GitHub’s verified signature.
Align native leehack/llamadart-native@v0.4.1 on upstream b29c606e28a01b1bc8c1351026a0fa6e616bf6c4, with matching Dart bindings
and Apple companion 0.0.19. This resolves the 0.8.23 grammar limitation:
native {2000} repetitions are accepted again
(llamadart-native#76).
Aligned the default WebGPU bridge assets to v0.1.44 for matching
Web/native v0.4.1@b29c606e28a01b1bc8c1351026a0fa6e616bf6c4 parity.
Immutable Web asset manifest: 8d61f453753ac7a7d839ac12318b70986a814748d86029993118c19454293aa9.
Update native LiteRT-LM to v0.17.0-6 with Apple companion 0.0.11,
Qwen3 tokenizer compatibility, corrected Linux loading, and explicit Linux
and Windows GPU selection while retaining CPU defaults. The Windows x64
runtime bundles dxil.dll and dxcompiler.dll, which D3D12 GPU engine
creation requires
(litert-lm-native#47).
Web LiteRT-LM stays at @litert-lm/core@0.15.0.
Preserve required iOS LiteRT-LM provider and Metal plugins, handle dependency
ordering in companion libraries, and exclude metadata/import archives from
runtime library inventories.
Respect greedy sampling for zero-temperature native LiteRT-LM generation.
Restore native Qwen3 chat text when thinking is disabled and preserve plain
system instructions when seeding LiteRT-LM conversation history.
Settle pending LiteRT-LM requests when a worker stops, close response ports,
and report unverified native cleanup as an error.
Preserve Unicode when detokenizing native GGUF tokens and suppress caller
stop markers across chunk boundaries and speculative decoding.
Restore native Qwen3-ASR file/encoded-byte transcription parity by keeping
encoded audio out of string chat-template prompts.
Render typed tool results as JSON text while preserving string results,
validate Qwen XML argument types against schemas, and reject malformed or
undeclared tool calls without exposing executable tool deltas.
Prevent split MiniMax M3 thinking delimiters from leaking into reasoning.
Discover Windows backend libraries in compiled CLI bundles and provide
bounded diagnostics for unavailable native thinking-budget helpers.
Fix fresh macOS Flutter dependency scanning while retaining Apple ABI and
local-override guards.
Add locked Gemma 4/Qwen3.5 validation profiles, explicit NPU coverage, and
opt-in speech and voice-pipeline diagnostics.
Preserve physical iOS speech-test failure-phase diagnostics and keep
repository writer checks independent of generated website output.
Retain open LiteRT-LM qualification gaps: macOS Qwen3.5 GPU reload latency
(#521), Gemma exact-history
behavior (#513), and
Qwen3-0.6B arithmetic on Android CPU and iOS CPU/GPU
(#509). The affected cases
remain unqualified; these changes do not resolve the failures or establish
their remaining owning layer.