Repository navigation
v0.11.0
-
Fix a Flutter macOS app aborting in ggml-metal when it quits while a
llama.cpp model loads or after a hot restart during a load, and a Dart
program aborting when it ends or kills an isolate during a load: the
llama.cpp runtime now frees the models, contexts, projectors and decision
heads still allocated at exit
(#813,
llamadart-native#96).
A native host's Cexit()while a llama.cpp call is running is not covered. -
Update the default llama.cpp runtime to
leehack/llamadart-native@v0.5.0-2,
a rebuild of llama.cppv0.5.0whose Apple XCFramework carries a privacy
manifest (File Timestamp, reasonC617.1). A runtime older thanv0.5.0-1
keeps the previous exit behavior and logs a warning at the first model load. -
Fix a process aborting when a Dart program dies of an error, or a Flutter
macOS app quits or hot restarts, while an image model loads or generates,
and a macOS Metal process aborting in ggml-metal when a native host exits
with an image model loaded: the stable_diffusion runtime now records
progress instead of calling back into Dart, and frees the image models
still allocated at exit
(stable-diffusion-native#10).
Image progress events now arrive up to about 50 ms after the runtime
reports them. A quit during an image generation waits for the generation
to finish; dispose the engine first to quit at once. -
Update the stable_diffusion runtime to
leehack/stable-diffusion-native@v0.2.0-1, a rebuild of the same
stable-diffusion.cpp commit whose Apple XCFramework carries a privacy
manifest (File Timestamp, reasonsC617.1and3B52.1). Image generation
needs this runtime or a later one. -
Documented that LiteRT-LM on iOS has run on a device only on iOS 18.3.2;
iOS 16.4 remains the declared, untested deployment floor
(#831). -
Documented Apple privacy manifests: one reaches an app inside a companion
package's XCFramework, never through the hook path, and an app whose
llama.cpp or stable_diffusion runtime has none declares File Timestamp
reasonC617.1(plus3B52.1for stable_diffusion) itself. -
Flutter iOS and macOS apps pair core
0.11.0with
llamadart_llama_cpp_flutter0.0.21,llamadart_litert_lm_flutter
0.0.13andllamadart_stable_diffusion_flutter0.0.2, which link the
runtimes this release pins; an older llama.cpp or stable_diffusion companion
fails the Apple build. -
Adopt LiteRT-LM
v0.17.0-7with provider-free iOS artifacts; Gemma FST
constrained decoding remains unavailable on iOS. -
Update the default LiteRT-LM runtime to
leehack/litert-lm-native@v0.17.0-8.
Its iOS SwiftPM frameworks carry Apple privacy manifests (File Timestamp
C617.1and3B52.1, System Boot Time35F9.1, User DefaultsCA92.1);
an app that runs LiteRT-LM on the hook path declares them itself. -
Fix Android LiteRT-LM GPU engines keeping their graphics memory after
deletion, which got an app killed after one or two model reloads on a
Galaxy S24
(litert-lm-native#59). -
Documented that LiteRT-LM's default GPU selection on Android generates wrong
text for Qwen3 0.6B on Adreno 750; load it withComputeDevice.cpu
(#553). -
Documented that llama.cpp Vulkan on Android is experimental and
device-dependent, with known failures on Pixel 9 Pro, Galaxy A53 and Galaxy
S24;autostays on the CPU
(#948). -
LiteRT-LM release sync can remove an obsolete iOS provider target from modern
Swift packages while preserving required macOS runtime libraries. -
Reentrant engine disposal now shares one teardown and immediately reports
disposed state, including calls from logging or backend cancellation hooks. -
Native requests, generation, and speech synthesis fail promptly when their
worker exits unexpectedly, allowing disposal to finish. -
Native recurrent and hybrid models reject nonzero speculative rollback
capacity before context creation, avoiding runtime graph-budget aborts. -
Native embeddings reject inputs that exceed a required one-pass micro-batch,
avoiding non-causal attention aborts and incorrect MEAN/CLS pooled vectors. -
Disposal now cancels model and LoRA downloads started by deprecated
loadModelSource, without waiting for resolution that ignores cancellation
(#895). -
Prevented execution of incomplete parallel tool-call replies when generation
reaches a reported limit; the tool loop rolls back the whole turn. -
Automatic tool loops now reject pinned LiteRT-LM native and Web runtimes
before generation because they cannot report token-limit truncation reliably;
manually managed completion remains available (#919). -
Web GGUF completions preserve runtime limits as
length, so tool loops
reporttruncatedand roll back cut-off turns; typed Web speech recognition
raises a runtime truncation error with the partial transcript. -
Redacted known URL secrets repeated in model filenames from download
cancellation messages. -
Model unload, replacement, and disposal now report interrupted tool loops as
cancelled, preserving partial answers and rolling back unfinished turns. -
Projector loads now reject model changes during loading instead of reporting
success after the model has been unloaded. -
Fixed: URL-valued
ModelSource.pathdiagnostics redact credentials and
signed queries while preserving the loading path and cache identity
(#846). -
Fixed LiteRT-LM speech recognition startup failures hanging indefinitely;
failed workers now release their resources and allow another attempt. -
Kept cancelled replies out of reset chats and rolled back turns cancelled
before any output; structured JSON cancellation now reports a state error.
A model change during draft-model resolution also rolls back the turn;
resets during context preparation preserve the replacement conversation. -
Redacted URL credentials and signed query strings from redirected download
failures, download snapshots and invalid model-source errors, including
slashless URLs; LiteRT-LM Web model names reject decoded URL delimiters. -
LlamaEngine.load(LlamaModel(source, projector:), params:, download:, onProgress:, store:)creates an engine and loads a model with its
projector in one atomic call, andsetModelloads or replaces the model of
an engine: the loaded model keeps serving until every new file has
downloaded, except on the Web, where the runtime unloads it before it
fetches the new one. A class thatimplements LlamaEnginemust add
setModel, and an override ofloadMultimodalProjectorSourceits new
downloadparameter
(#846). -
Deprecated:
LlamaEngine.loadModel,loadModelSource,
loadModelFromUrlandloadMultimodalProjector; useLlamaEngine.loador
setModel.loadMultimodalProjectorSource(options:)is nowdownload:.
They still work until 1.0
(#846). -
Fixed: a load checks that it can proceed before it downloads:
LlamaEngine.loadandsetModelreject a projector for a LiteRT-LM model
andComputeDevice.npufor a GGUF first, and the deprecated
loadModelSourcethrows for an already loaded engine first
(#846). -
Fixed:
LlamaEngine.dispose()stops the downloads of a running
setModelat once, andunloadModel()anddispose()stop a
loadMultimodalProjectorSourcedownload instead of waiting for it; the
load throwsLlamaStateException
(#895,
#896). -
Behavior change: on the Web, a
ModelSource.pathloads as a URL
relative to the document, or as ablob:URL, for models, projectors, LoRA
adapters and draft models; it used to throwLlamaUnsupportedException
(#846). -
Fixed: the llama.cpp WebGPU bridge resolves a relative model,
projector, LoRA or draft model URL against the document, as LiteRT-LM Web
does; its worker resolved one againstwebgpu_bridge/, so the fetch
failed (#846). -
A local
ModelSource.pathwhose file name holds%2For%5C, or whose
directory is named%2eor%2e%2e, loads throughLlamaEngine.load,
setModel,loadMultimodalProjectorSource,setLoraSource,
ModelParams.loras, draft models and the image, decision and speech
engines'load(#846). -
Breaking (Preview):
ImageGenerationEngine.generatereturns
Future<ImageGenerationTask>; await it before readingeventsor calling
cancel(#850). -
Breaking:
ImageGenerationTask,SpeechToTextTaskand
TextToSpeechTaskno longer report a failure as an error onevents; read
it fromdone, or from the newresult, which returns the result or
throws the failure, orLlamaStateExceptionwhen the task is cancelled
(#850). -
Fixed:
SpeechToTextTask.cancel()stops only that recognition; it no
longer cancels chat and other requests on the sameLlamaEngine
(#850). -
ModelParams.device(ComputeDevice) selects the device for every
runtime:autokeeps each runtime's default, and an explicitcpu,gpu
ornpuruns there or throwsLlamaUnsupportedExceptioninstead of
falling back to another device
(#849). -
Deprecated:
ModelParams.liteRtLmBackend,LiteRtLmBackendPreference
andLiteRtLmBackend(preferredBackend:); useModelParams.device. They
still work until 1.0
(#849). -
Behavior change:
DecisionModelParams(device: ComputeDevice.gpu)runs
the encoder on a GPU, Vulkan on Android, or throws
LlamaUnsupportedExceptionfrom the encoder load; it ran on the CPU on
Android and threw before loading where GPU modules load with the model
(#849). -
Behavior change:
LlamaEngineloads runModelParams.validate()before
any download or native call, so an invalid combination throws
LlamaArgumentExceptioninstead ofLlamaModelException
(#849). -
Breaking:
ModelParams.validate()throwsLlamaArgumentException,
ModelDownloadControllerthrowsLlamaArgumentExceptionor
LlamaStateException, and Web backend calls before a model load throw
LlamaStateException, instead ofArgumentErrororStateError; a failed
load'sdetailsis now the cause's message instead of a{type, message}
map, and the not-ready error namesLlamaEngine.loadandsetModel
(#843). -
Deprecated, behavior change:
sourceLangCodeandtargetLangCodeon
LlamaEngine.create,createStructuredJson,chatTemplateand
BackendNativeChatGeneration.generateChat; pass
chatTemplateKwargs: {'source_lang_code': 'en', 'target_lang_code': 'ko'},
which TranslateGemma templates now read, as llama.cpp does. A custom
BackendNativeChatGenerationnow gets the codes fromLlamaEngineonly in
chatTemplateKwargs
(#853). -
Deprecated:
LlamaLogging.configure(level:, nativeLevel:, handler:)
replacesLlamaEngine.configureLoggingand the engine'ssetLogLevel,
setDartLogLevelandsetNativeLogLevel; levels are now library-wide, so
the last call wins and reaches every running engine's worker, including the
default native backend's (#845). -
Deprecated: speech engines follow the shared engine pattern:
SpeechToTextEngine.load(SpeechToTextModel(...))and
TextToSpeechEngine.load(TextToSpeechModel(...))download every
ModelSourceand own what they load,attach(engine, adapter:)borrows a
loadedLlamaEngine,dispose()cancels the running task, and
transcribeOnceandsynthesizeOncereturn the final result. Adapters
(Qwen3AsrAdapter,LiteRtLmAsrAdapter,Qwen3TtsAdapter, or your own
SpeechToTextPromptAdapter) replaceSpeechToTextModelProfile,
TextToSpeechModelProfile, themodelProfileconstructors and
SpeechToTextEngine.liteRtLm, which keep working until 1.0
(#848). -
Breaking:
SpeechToTextEnginegainsdispose(),isDisposed,
adapterandtranscribeOnce, andTextToSpeechEnginegainsdispose(),
isDisposed,adapterandsynthesizeOnce, so a class thatimplements
either must add them
(#848). -
Breaking:
DecisionEngine.load(DecisionModel(encoder:, head:, config:), params:, download:, onProgress:)loads a decision model from
ModelSources into an engine it owns, atomically, anddispose()frees it
all;DecisionEngine.attach(engine, head:, config:)adds a head to a loaded
LlamaEngine; decision engines report an instancecapabilities. The
String-pathload(engine, headPath:, configPath:)is removed; use
attach(#847). -
Breaking (Preview): image generation follows the shared engine
pattern:ImageGenerationEngine.load(ImageGenerationModel(source, components: [...]), params:, download:, onProgress:)downloads every
ModelSourceinto the model cache, with combined progress, cancellation
and cache reuse, and detects each file's role from its header. Generation
settings move toImageGenerationRequest.download's bearer token and
headers never reach more than one origin: such remote files throw
LlamaArgumentException. The model presets,ImageGenerationModelFamily,
ImageGenerationModelFiles,ImageGenerationDefaults,
ImageGenerationOptions(nowImageModelParams, and the engine'soptions
getterparams),ImageGenerationDevice(nowComputeDevice) andString
paths are removed;MIGRATION.mdmaps each former preset to its files and
request settings
(#883). -
Breaking:
LlamaEngine.dispose()is idempotent and terminal, like
every other engine's: each call returns the same future, and afterwards
loads, requests andDecisionEngine.attachthrowLlamaStateException
whilecapabilitiesreports the engine as disposed.getBackendName,
getAvailableBackends,isGpuSupported,getVramInfo,listGpuDevices
andgetResolvedGpuLayersused to answer afterdispose()and now throw
LlamaStateExceptiontoo. A load running when it is called throws
LlamaStateException, and its model is unloaded.LlamaEnginegains
isDisposed, so a class thatimplementsit must add it
(#851). -
Breaking (Preview):
ImageGenerationEngine.capabilitiesis async, and
ImageGenerationEngine.runtimeCapabilities()is removed; use
checkRuntime().ImageGenerationCapabilitiesandDecisionCapabilities
implementEngineCapabilities
(#851). -
Behavior change:
LlamaEngine.supportsVisionandsupportsAudioreport
whatcapabilitiesreports: true for a LiteRT-LM bundle that takes media
directly, and false instead of throwing when the runtime cannot probe the
projector (#851). -
Behavior change:
responseFormatmaps with an unknowntypeor
key, such asjson_shemaor a misspelledschma, now throw
LlamaUnsupportedExceptionbefore generation instead of generating
unconstrained output; anull-valued key counts as absent
(#836,
#864). -
ChatSession.createtakesresponseFormat, and the new
ChatSession.createStructuredJsondecodes the reply. A turn that fails or
is cancelled before its first chunk, such as a strict format on LiteRT-LM,
removes its user message from the history, and one cancelled or failing
mid-stream keeps the partial reply, so alternating-role templates keep
working
(#836,
#864). -
LlamaCompletionChunk.model, observer model names and load logs report
llama_model, and web LiteRT-LM omitsgeneral.name, when a URL's last
path segment repeats its userinfo credential; web LiteRT-LM
litert_lm.model_urlshows a relative URL as given and ablob:ordata:
URL as its scheme (#822). -
loadModelSourcethrowsLlamaUnsupportedExceptioninstead of
ArgumentErrorfor a local path whose file name holds%2For%5Cor
whose directory is named%2eor%2e%2e;loadModelstill loads it
(#822). -
Behavior change: on Android and iOS, the default model cache is now
llamadart/modelsin the app's cache directory instead of the temporary
directory, which Android empties on every app update and iOS purges; add
DefaultModelDownloadManager.globalCacheDirectoryto move every default
download (#838). -
Native
LlamaBackend()picks llama.cpp or LiteRT-LM from the model file's
header, not its extension, so extensionless downloads load in the right
runtime and mislabelled files throwLlamaModelFormatException; name a Web
URL's format withModelSource.url(..., format: ModelFormat.liteRtLm)
(#837). -
Add
LlamaEngine.runtimeandLlamaEngine.capabilities, one snapshot of
what the loaded model's runtime supports: image and audio input,
embeddings, multi-turn chat, tools, structured output, grammars, every
sampling control includingpenalty, stream batching and speculative
strategies. Native LiteRT-LM reports the image, audio and speculative
decoding support the bundle declares, best-effort (litert-lm-native#60),
and a request the runtime then fails for lack of one throws
LlamaUnsupportedExceptionnaming it instead of an opaque error.
backendGenerationCapabilitiesis deprecated; on native LiteRT-LM it now
reportsstreamBatchingand, for bundles without a declared drafter, no
speculative strategy (#841). -
Read completions without
choices.first.delta:chunk.text,
chunk.thinking,chunk.toolCallsand a typedchunk.finishReason
(LlamaFinishReason);stream.text(),stream.textDeltas()and
stream.collect()(aLlamaCompletionwith assembled tool calls and an
assistantmessage); and the one-shotengine.complete(messages)and
session.send('...')
(#840). -
session.sendWithTools(text, tools: ...)andcompleteWithTools(parts, ...)run the model's tool calls with eachToolDefinition.handler,
concurrently for parallel calls, until it answers, and return a
LlamaToolLoopResultwhosestopReasonalso reportsmaxRounds,
unhandled calls, context overflow, a reply cut off atmaxTokens
(truncated, rolled back) and cancellation
(#842). -
Breaking:
ChatSessiongainscreateStructuredJson, andLlamaEngine
gainsruntime,capabilities,setLoraSourceandremoveLoraSource, so
a class thatimplementseither must add them. -
Breaking:
ToolDefinition.handleris nullable, so tools the app runs
itself can leave it out; code that callstool.handler(params)must check
it for null first (#842). -
GenerationGrammarTrigger.typed(type: GrammarTriggerType.word, ...)
replaces the raw-intconstructor, now deprecated; an unknown raw trigger
type throwsLlamaUnsupportedExceptionon llama.cpp instead of being
ignored (#844). -
LoRA adapters (
setLoraSource,removeLoraSource,
LoraAdapterConfig.source), speculative draft models
(SpeculativeDecodingConfig.draftModel,withDraftModel,
withDraftModelDownload) take aModelSource, so they download and cache
like models, andLiteRtLmAsrRuntimeConfig.sourcetakes local
ModelSourcefiles
(#852). -
Deprecated: the
Stringpath forms of LoRA adapters, speculative draft
models and LiteRT-LM ASR files
(#852). -
Breaking:
package:llamadart/llamadart.dartis the app API. The raw
ffigen bindings move topackage:llamadart/llama_cpp_bindings.dart(native
only, outside semantic versioning), and the custom-backend SPI moves to the
newpackage:llamadart/backend.dart: everyBackend*type except
BackendPerfContextDataandBackendTextToSpeechModel,LiteRtLmBackend,
LiteRtLmRuntimeClient,LiteRtLmRuntimeMetrics,LiteRtLmRuntimeResult,
LiteRtLmAsrRuntimeSession,LiteRtLmAsrPushResult,
LiteRtLmAsrProcessResultandLiteRtLmAsrProcessState. The app API now
exportsTemplateToolCallSerialization
(#355). -
Breaking: the
LlamaEnginetext-to-speech and decision hooks,
modelHandleandcontextHandlemove to theLlamaEngineBackendHooks
extension inpackage:llamadart/backend.dart, so neither a subclass nor an
implements LlamaEnginefake can override them; fake a backend that
implementsBackendTextToSpeechorBackendDecisioninstead
(#355). -
Breaking: the deprecated
LiteRtLmBenchmarkClient,
LiteRtLmBenchmarkMetrics,LiteRtLmBenchmarkResult,
LiteRtLmRuntimeClient.conversationTokenCountand
LiteRtLmRuntimeClient.replaceConversationWithCloneare removed
(#355).
- Aligned the default WebGPU bridge assets to
v0.1.54, unchanged from 0.10.0:
they embed llama.cppv0.5.0, are qualified against nativev0.5.0, and keep
Web/native llama.cppv0.5.0@7fe450e19305b828c199d602c23a8337aaa1f03bparity
and Web@litert-lm/core@0.15.0. Immutable Web asset manifest:
8a9f83c15035eeb034a6563e6f753382d7d7f9be81503ef76902138da7841176.