Skip to content

v0.11.0

Choose a tag to compare

@github-actions github-actions released this 07 Oct 17:07
· 27 commits to main since this release
827b56e
  • Fix a Flutter macOS app aborting in ggml-metal when it quits while a
    llama.cpp model loads or after a hot restart during a load, and a Dart
    program aborting when it ends or kills an isolate during a load: the
    llama.cpp runtime now frees the models, contexts, projectors and decision
    heads still allocated at exit
    (#813,
    llamadart-native#96).
    A native host's C exit() while a llama.cpp call is running is not covered.

  • Update the default llama.cpp runtime to leehack/llamadart-native@v0.5.0-2,
    a rebuild of llama.cpp v0.5.0 whose Apple XCFramework carries a privacy
    manifest (File Timestamp, reason C617.1). A runtime older than v0.5.0-1
    keeps the previous exit behavior and logs a warning at the first model load.

  • Fix a process aborting when a Dart program dies of an error, or a Flutter
    macOS app quits or hot restarts, while an image model loads or generates,
    and a macOS Metal process aborting in ggml-metal when a native host exits
    with an image model loaded: the stable_diffusion runtime now records
    progress instead of calling back into Dart, and frees the image models
    still allocated at exit
    (stable-diffusion-native#10).
    Image progress events now arrive up to about 50 ms after the runtime
    reports them. A quit during an image generation waits for the generation
    to finish; dispose the engine first to quit at once.

  • Update the stable_diffusion runtime to
    leehack/stable-diffusion-native@v0.2.0-1, a rebuild of the same
    stable-diffusion.cpp commit whose Apple XCFramework carries a privacy
    manifest (File Timestamp, reasons C617.1 and 3B52.1). Image generation
    needs this runtime or a later one.

  • Documented that LiteRT-LM on iOS has run on a device only on iOS 18.3.2;
    iOS 16.4 remains the declared, untested deployment floor
    (#831).

  • Documented Apple privacy manifests: one reaches an app inside a companion
    package's XCFramework, never through the hook path, and an app whose
    llama.cpp or stable_diffusion runtime has none declares File Timestamp
    reason C617.1 (plus 3B52.1 for stable_diffusion) itself.

  • Flutter iOS and macOS apps pair core 0.11.0 with
    llamadart_llama_cpp_flutter 0.0.21, llamadart_litert_lm_flutter
    0.0.13 and llamadart_stable_diffusion_flutter 0.0.2, which link the
    runtimes this release pins; an older llama.cpp or stable_diffusion companion
    fails the Apple build.

  • Adopt LiteRT-LM v0.17.0-7 with provider-free iOS artifacts; Gemma FST
    constrained decoding remains unavailable on iOS.

  • Update the default LiteRT-LM runtime to leehack/litert-lm-native@v0.17.0-8.
    Its iOS SwiftPM frameworks carry Apple privacy manifests (File Timestamp
    C617.1 and 3B52.1, System Boot Time 35F9.1, User Defaults CA92.1);
    an app that runs LiteRT-LM on the hook path declares them itself.

  • Fix Android LiteRT-LM GPU engines keeping their graphics memory after
    deletion, which got an app killed after one or two model reloads on a
    Galaxy S24
    (litert-lm-native#59).

  • Documented that LiteRT-LM's default GPU selection on Android generates wrong
    text for Qwen3 0.6B on Adreno 750; load it with ComputeDevice.cpu
    (#553).

  • Documented that llama.cpp Vulkan on Android is experimental and
    device-dependent, with known failures on Pixel 9 Pro, Galaxy A53 and Galaxy
    S24; auto stays on the CPU
    (#948).

  • LiteRT-LM release sync can remove an obsolete iOS provider target from modern
    Swift packages while preserving required macOS runtime libraries.

  • Reentrant engine disposal now shares one teardown and immediately reports
    disposed state, including calls from logging or backend cancellation hooks.

  • Native requests, generation, and speech synthesis fail promptly when their
    worker exits unexpectedly, allowing disposal to finish.

  • Native recurrent and hybrid models reject nonzero speculative rollback
    capacity before context creation, avoiding runtime graph-budget aborts.

  • Native embeddings reject inputs that exceed a required one-pass micro-batch,
    avoiding non-causal attention aborts and incorrect MEAN/CLS pooled vectors.

  • Disposal now cancels model and LoRA downloads started by deprecated
    loadModelSource, without waiting for resolution that ignores cancellation
    (#895).

  • Prevented execution of incomplete parallel tool-call replies when generation
    reaches a reported limit; the tool loop rolls back the whole turn.

  • Automatic tool loops now reject pinned LiteRT-LM native and Web runtimes
    before generation because they cannot report token-limit truncation reliably;
    manually managed completion remains available (#919).

  • Web GGUF completions preserve runtime limits as length, so tool loops
    report truncated and roll back cut-off turns; typed Web speech recognition
    raises a runtime truncation error with the partial transcript.

  • Redacted known URL secrets repeated in model filenames from download
    cancellation messages.

  • Model unload, replacement, and disposal now report interrupted tool loops as
    cancelled, preserving partial answers and rolling back unfinished turns.

  • Projector loads now reject model changes during loading instead of reporting
    success after the model has been unloaded.

  • Fixed: URL-valued ModelSource.path diagnostics redact credentials and
    signed queries while preserving the loading path and cache identity
    (#846).

  • Fixed LiteRT-LM speech recognition startup failures hanging indefinitely;
    failed workers now release their resources and allow another attempt.

  • Kept cancelled replies out of reset chats and rolled back turns cancelled
    before any output; structured JSON cancellation now reports a state error.
    A model change during draft-model resolution also rolls back the turn;
    resets during context preparation preserve the replacement conversation.

  • Redacted URL credentials and signed query strings from redirected download
    failures, download snapshots and invalid model-source errors, including
    slashless URLs; LiteRT-LM Web model names reject decoded URL delimiters.

  • LlamaEngine.load(LlamaModel(source, projector:), params:, download:, onProgress:, store:) creates an engine and loads a model with its
    projector in one atomic call, and setModel loads or replaces the model of
    an engine: the loaded model keeps serving until every new file has
    downloaded, except on the Web, where the runtime unloads it before it
    fetches the new one. A class that implements LlamaEngine must add
    setModel, and an override of loadMultimodalProjectorSource its new
    download parameter
    (#846).

  • Deprecated: LlamaEngine.loadModel, loadModelSource,
    loadModelFromUrl and loadMultimodalProjector; use LlamaEngine.load or
    setModel. loadMultimodalProjectorSource(options:) is now download:.
    They still work until 1.0
    (#846).

  • Fixed: a load checks that it can proceed before it downloads:
    LlamaEngine.load and setModel reject a projector for a LiteRT-LM model
    and ComputeDevice.npu for a GGUF first, and the deprecated
    loadModelSource throws for an already loaded engine first
    (#846).

  • Fixed: LlamaEngine.dispose() stops the downloads of a running
    setModel at once, and unloadModel() and dispose() stop a
    loadMultimodalProjectorSource download instead of waiting for it; the
    load throws LlamaStateException
    (#895,
    #896).

  • Behavior change: on the Web, a ModelSource.path loads as a URL
    relative to the document, or as a blob: URL, for models, projectors, LoRA
    adapters and draft models; it used to throw LlamaUnsupportedException
    (#846).

  • Fixed: the llama.cpp WebGPU bridge resolves a relative model,
    projector, LoRA or draft model URL against the document, as LiteRT-LM Web
    does; its worker resolved one against webgpu_bridge/, so the fetch
    failed (#846).

  • A local ModelSource.path whose file name holds %2F or %5C, or whose
    directory is named %2e or %2e%2e, loads through LlamaEngine.load,
    setModel, loadMultimodalProjectorSource, setLoraSource,
    ModelParams.loras, draft models and the image, decision and speech
    engines' load (#846).

  • Breaking (Preview): ImageGenerationEngine.generate returns
    Future<ImageGenerationTask>; await it before reading events or calling
    cancel (#850).

  • Breaking: ImageGenerationTask, SpeechToTextTask and
    TextToSpeechTask no longer report a failure as an error on events; read
    it from done, or from the new result, which returns the result or
    throws the failure, or LlamaStateException when the task is cancelled
    (#850).

  • Fixed: SpeechToTextTask.cancel() stops only that recognition; it no
    longer cancels chat and other requests on the same LlamaEngine
    (#850).

  • ModelParams.device (ComputeDevice) selects the device for every
    runtime: auto keeps each runtime's default, and an explicit cpu, gpu
    or npu runs there or throws LlamaUnsupportedException instead of
    falling back to another device
    (#849).

  • Deprecated: ModelParams.liteRtLmBackend, LiteRtLmBackendPreference
    and LiteRtLmBackend(preferredBackend:); use ModelParams.device. They
    still work until 1.0
    (#849).

  • Behavior change: DecisionModelParams(device: ComputeDevice.gpu) runs
    the encoder on a GPU, Vulkan on Android, or throws
    LlamaUnsupportedException from the encoder load; it ran on the CPU on
    Android and threw before loading where GPU modules load with the model
    (#849).

  • Behavior change: LlamaEngine loads run ModelParams.validate() before
    any download or native call, so an invalid combination throws
    LlamaArgumentException instead of LlamaModelException
    (#849).

  • Breaking: ModelParams.validate() throws LlamaArgumentException,
    ModelDownloadController throws LlamaArgumentException or
    LlamaStateException, and Web backend calls before a model load throw
    LlamaStateException, instead of ArgumentError or StateError; a failed
    load's details is now the cause's message instead of a {type, message}
    map, and the not-ready error names LlamaEngine.load and setModel
    (#843).

  • Deprecated, behavior change: sourceLangCode and targetLangCode on
    LlamaEngine.create, createStructuredJson, chatTemplate and
    BackendNativeChatGeneration.generateChat; pass
    chatTemplateKwargs: {'source_lang_code': 'en', 'target_lang_code': 'ko'},
    which TranslateGemma templates now read, as llama.cpp does. A custom
    BackendNativeChatGeneration now gets the codes from LlamaEngine only in
    chatTemplateKwargs
    (#853).

  • Deprecated: LlamaLogging.configure(level:, nativeLevel:, handler:)
    replaces LlamaEngine.configureLogging and the engine's setLogLevel,
    setDartLogLevel and setNativeLogLevel; levels are now library-wide, so
    the last call wins and reaches every running engine's worker, including the
    default native backend's (#845).

  • Deprecated: speech engines follow the shared engine pattern:
    SpeechToTextEngine.load(SpeechToTextModel(...)) and
    TextToSpeechEngine.load(TextToSpeechModel(...)) download every
    ModelSource and own what they load, attach(engine, adapter:) borrows a
    loaded LlamaEngine, dispose() cancels the running task, and
    transcribeOnce and synthesizeOnce return the final result. Adapters
    (Qwen3AsrAdapter, LiteRtLmAsrAdapter, Qwen3TtsAdapter, or your own
    SpeechToTextPromptAdapter) replace SpeechToTextModelProfile,
    TextToSpeechModelProfile, the modelProfile constructors and
    SpeechToTextEngine.liteRtLm, which keep working until 1.0
    (#848).

  • Breaking: SpeechToTextEngine gains dispose(), isDisposed,
    adapter and transcribeOnce, and TextToSpeechEngine gains dispose(),
    isDisposed, adapter and synthesizeOnce, so a class that implements
    either must add them
    (#848).

  • Breaking: DecisionEngine.load(DecisionModel(encoder:, head:, config:), params:, download:, onProgress:) loads a decision model from
    ModelSources into an engine it owns, atomically, and dispose() frees it
    all; DecisionEngine.attach(engine, head:, config:) adds a head to a loaded
    LlamaEngine; decision engines report an instance capabilities. The
    String-path load(engine, headPath:, configPath:) is removed; use
    attach (#847).

  • Breaking (Preview): image generation follows the shared engine
    pattern: ImageGenerationEngine.load(ImageGenerationModel(source, components: [...]), params:, download:, onProgress:) downloads every
    ModelSource into the model cache, with combined progress, cancellation
    and cache reuse, and detects each file's role from its header. Generation
    settings move to ImageGenerationRequest. download's bearer token and
    headers never reach more than one origin: such remote files throw
    LlamaArgumentException. The model presets, ImageGenerationModelFamily,
    ImageGenerationModelFiles, ImageGenerationDefaults,
    ImageGenerationOptions (now ImageModelParams, and the engine's options
    getter params), ImageGenerationDevice (now ComputeDevice) and String
    paths are removed; MIGRATION.md maps each former preset to its files and
    request settings
    (#883).

  • Breaking: LlamaEngine.dispose() is idempotent and terminal, like
    every other engine's: each call returns the same future, and afterwards
    loads, requests and DecisionEngine.attach throw LlamaStateException
    while capabilities reports the engine as disposed. getBackendName,
    getAvailableBackends, isGpuSupported, getVramInfo, listGpuDevices
    and getResolvedGpuLayers used to answer after dispose() and now throw
    LlamaStateException too. A load running when it is called throws
    LlamaStateException, and its model is unloaded. LlamaEngine gains
    isDisposed, so a class that implements it must add it
    (#851).

  • Breaking (Preview): ImageGenerationEngine.capabilities is async, and
    ImageGenerationEngine.runtimeCapabilities() is removed; use
    checkRuntime(). ImageGenerationCapabilities and DecisionCapabilities
    implement EngineCapabilities
    (#851).

  • Behavior change: LlamaEngine.supportsVision and supportsAudio report
    what capabilities reports: true for a LiteRT-LM bundle that takes media
    directly, and false instead of throwing when the runtime cannot probe the
    projector (#851).

  • Behavior change: responseFormat maps with an unknown type or
    key, such as json_shema or a misspelled schma, now throw
    LlamaUnsupportedException before generation instead of generating
    unconstrained output; a null-valued key counts as absent
    (#836,
    #864).

  • ChatSession.create takes responseFormat, and the new
    ChatSession.createStructuredJson decodes the reply. A turn that fails or
    is cancelled before its first chunk, such as a strict format on LiteRT-LM,
    removes its user message from the history, and one cancelled or failing
    mid-stream keeps the partial reply, so alternating-role templates keep
    working
    (#836,
    #864).

  • LlamaCompletionChunk.model, observer model names and load logs report
    llama_model, and web LiteRT-LM omits general.name, when a URL's last
    path segment repeats its userinfo credential; web LiteRT-LM
    litert_lm.model_url shows a relative URL as given and a blob: or data:
    URL as its scheme (#822).

  • loadModelSource throws LlamaUnsupportedException instead of
    ArgumentError for a local path whose file name holds %2F or %5C or
    whose directory is named %2e or %2e%2e; loadModel still loads it
    (#822).

  • Behavior change: on Android and iOS, the default model cache is now
    llamadart/models in the app's cache directory instead of the temporary
    directory, which Android empties on every app update and iOS purges; add
    DefaultModelDownloadManager.globalCacheDirectory to move every default
    download (#838).

  • Native LlamaBackend() picks llama.cpp or LiteRT-LM from the model file's
    header, not its extension, so extensionless downloads load in the right
    runtime and mislabelled files throw LlamaModelFormatException; name a Web
    URL's format with ModelSource.url(..., format: ModelFormat.liteRtLm)
    (#837).

  • Add LlamaEngine.runtime and LlamaEngine.capabilities, one snapshot of
    what the loaded model's runtime supports: image and audio input,
    embeddings, multi-turn chat, tools, structured output, grammars, every
    sampling control including penalty, stream batching and speculative
    strategies. Native LiteRT-LM reports the image, audio and speculative
    decoding support the bundle declares, best-effort (litert-lm-native#60),
    and a request the runtime then fails for lack of one throws
    LlamaUnsupportedException naming it instead of an opaque error.
    backendGenerationCapabilities is deprecated; on native LiteRT-LM it now
    reports streamBatching and, for bundles without a declared drafter, no
    speculative strategy (#841).

  • Read completions without choices.first.delta: chunk.text,
    chunk.thinking, chunk.toolCalls and a typed chunk.finishReason
    (LlamaFinishReason); stream.text(), stream.textDeltas() and
    stream.collect() (a LlamaCompletion with assembled tool calls and an
    assistant message); and the one-shot engine.complete(messages) and
    session.send('...')
    (#840).

  • session.sendWithTools(text, tools: ...) and completeWithTools(parts, ...) run the model's tool calls with each ToolDefinition.handler,
    concurrently for parallel calls, until it answers, and return a
    LlamaToolLoopResult whose stopReason also reports maxRounds,
    unhandled calls, context overflow, a reply cut off at maxTokens
    (truncated, rolled back) and cancellation
    (#842).

  • Breaking: ChatSession gains createStructuredJson, and LlamaEngine
    gains runtime, capabilities, setLoraSource and removeLoraSource, so
    a class that implements either must add them.

  • Breaking: ToolDefinition.handler is nullable, so tools the app runs
    itself can leave it out; code that calls tool.handler(params) must check
    it for null first (#842).

  • GenerationGrammarTrigger.typed(type: GrammarTriggerType.word, ...)
    replaces the raw-int constructor, now deprecated; an unknown raw trigger
    type throws LlamaUnsupportedException on llama.cpp instead of being
    ignored (#844).

  • LoRA adapters (setLoraSource, removeLoraSource,
    LoraAdapterConfig.source), speculative draft models
    (SpeculativeDecodingConfig.draftModel, withDraftModel,
    withDraftModelDownload) take a ModelSource, so they download and cache
    like models, and LiteRtLmAsrRuntimeConfig.source takes local
    ModelSource files
    (#852).

  • Deprecated: the String path forms of LoRA adapters, speculative draft
    models and LiteRT-LM ASR files
    (#852).

  • Breaking: package:llamadart/llamadart.dart is the app API. The raw
    ffigen bindings move to package:llamadart/llama_cpp_bindings.dart (native
    only, outside semantic versioning), and the custom-backend SPI moves to the
    new package:llamadart/backend.dart: every Backend* type except
    BackendPerfContextData and BackendTextToSpeechModel, LiteRtLmBackend,
    LiteRtLmRuntimeClient, LiteRtLmRuntimeMetrics, LiteRtLmRuntimeResult,
    LiteRtLmAsrRuntimeSession, LiteRtLmAsrPushResult,
    LiteRtLmAsrProcessResult and LiteRtLmAsrProcessState. The app API now
    exports TemplateToolCallSerialization
    (#355).

  • Breaking: the LlamaEngine text-to-speech and decision hooks,
    modelHandle and contextHandle move to the LlamaEngineBackendHooks
    extension in package:llamadart/backend.dart, so neither a subclass nor an
    implements LlamaEngine fake can override them; fake a backend that
    implements BackendTextToSpeech or BackendDecision instead
    (#355).

  • Breaking: the deprecated LiteRtLmBenchmarkClient,
    LiteRtLmBenchmarkMetrics, LiteRtLmBenchmarkResult,
    LiteRtLmRuntimeClient.conversationTokenCount and
    LiteRtLmRuntimeClient.replaceConversationWithClone are removed
    (#355).

  • Aligned the default WebGPU bridge assets to v0.1.54, unchanged from 0.10.0:
    they embed llama.cpp v0.5.0, are qualified against native v0.5.0, and keep
    Web/native llama.cpp v0.5.0@7fe450e19305b828c199d602c23a8337aaa1f03b parity
    and Web @litert-lm/core@0.15.0. Immutable Web asset manifest:
    8a9f83c15035eeb034a6563e6f753382d7d7f9be81503ef76902138da7841176.