Skip to content

v0.10.0

Choose a tag to compare

@github-actions github-actions released this 02 Oct 10:08
· 57 commits to main since this release
9070538
  • Add the llamadart_stable_diffusion_flutter 0.0.1 companion package:
    Flutter iOS and macOS apps that add it link the image generation runtime
    through Swift Package Manager, since App Store Connect rejects the iOS
    framework the build hook bundles. Without it the hook keeps bundling the
    runtime, and an iOS build reports an Xcode build warning (shown by Xcode and
    xcodebuild, not by plain flutter build or flutter run output).
  • Image generation now runs on the x86_64 iOS simulator, with the runtime
    updated to leehack/stable-diffusion-native@v0.2.0.
  • Label image generation a Preview in the README and docs, and list each
    image preset's model license.
  • Behavior change: render and parse llama.cpp chat with
    ModelParams.chatTemplate, which engine.create and
    engine.chatTemplate ignored in favor of the GGUF template
    (#710).
  • Breaking: WebGPU LlamaBackend.applyChatTemplate now throws
    LlamaUnsupportedException for a template override it cannot render
    (#710).
  • Behavior change: report only the last path segment of the model
    source, such as qwen.gguf, in LlamaCompletionChunk.model instead of
    the full local path or redacted URL; compare it with the file name, not
    the path. Like LlamaOperation.model, it now reads local paths as paths
    and leaves out data: and blob: URLs
    (#718).
  • Strip URL userinfo, query and fragment from the litert_lm.model_url
    metadata that web LiteRT-LM reports.
  • Behavior change: load models whose local path contains %, such as
    C:\models\qwen 100%.gguf, or whose URL file name decodes to one, instead
    of throwing ArgumentError; loadModelSource and ModelCacheEntry keep
    % in local and cache paths literal instead of percent-decoding them into
    another file (#819).
  • Fix Dart programs aborting on macOS Metal when main returns or throws
    with a model, decision head or image model still loaded; llamadart now
    frees them as the program ends, and an undisposed ImageGenerationEngine
    no longer keeps the program running
    (#613).
  • Fix the Flutter chat example aborting on macOS when quit while an image or
    chat model was still being freed, such as right after leaving the image
    screen (#796).
  • Document disposing engines from AppLifecycleListener.onExitRequested,
    since quitting a macOS app with a Metal model loaded aborts the process.
  • Ship agent skills for coding agents covering setup, chat and streaming,
    tool calling, web, Flutter apps, multimodal input, embeddings, speech, LoRA
    adapters, decision models and image generation; install them with
    dart run skills@ get.
  • Add an observability guide and tested optional OpenTelemetry example with Langfuse and Grafana recipes.
  • Behavior change: apply ModelParams.loras at model load on native
    llama.cpp and WebGPU, where they were silently ignored; an adapter that
    cannot be applied, or WebGPU bridge assets before v0.1.54, fail the load
    instead (#709).
  • Breaking: throw LlamaUnsupportedException instead of
    LlamaModelException when native or web LiteRT-LM rejects a ModelParams
    field, including more than one LoRA adapter or a non-default adapter
    scale.
  • Add an experimental opt-in stable_diffusion native runtime
    (stable-diffusion.cpp) to llamadart_native_runtimes for image
    generation; it is never bundled by default or by all, and
    llamadart_stable_diffusion_backends picks its CPU or Vulkan build on Linux
    and Windows (#777).
  • Add experimental on-device image generation, ImageGenerationEngine, on
    the opt-in stable_diffusion runtime: SDXS and SD-Turbo presets,
    phase-labelled progress, cancellation, PNG output, a memory check before
    loading, and an output size that defaults to the model's
    (ImageGenerationDefaults.width and height); not available on the web
    (#778).
  • Add an experimental image-generation screen to the Flutter chat example,
    with SDXS and SD-Turbo downloads
    (#776).
  • Add ImageGenerationEngine.warmUp, which compiles the GPU pipelines of
    the first image ahead of time, and document first-image latency per
    platform (#790).
  • Add ImageGenerationEngine.checkRuntime, which probes the image runtime
    without blocking the calling isolate; load now probes the same way, and
    the chat example uses it, so the first probe's Metal library compile (about
    16 s with an empty shader cache) no longer freezes the UI
    (#798).
  • Breaking: throw LlamaUnsupportedException for image or audio parts
    sent to a GGUF model with no projector loaded: native llama.cpp answered
    from the text alone, and WebGPU threw an untyped error surfaced as
    LlamaInferenceException. Native llama.cpp and LiteRT-LM also throw it
    for LlamaImageContent.url.
  • Add LlamaEngine.supportsEmbeddings. Native llama.cpp embeddings of an
    encoder-decoder model now throw LlamaUnsupportedException, and input
    longer than the context throws LlamaInferenceException, instead of a
    plain Exception.
  • Fix the example chat app failing every later turn after its projector is
    cleared or a text-only model is loaded in a conversation that already has
    images or audio; earlier media is now left out of the prompt, and
    unsupported-input errors are shown instead of generic reload advice.
  • Fix the example chat app garbling a streaming reply when mmproj is loaded
    manually mid-reply; the reply is now stopped first, as with Stop.
  • Name the missing Visual C++ runtime DLLs when native llama.cpp fails to
    load on Windows, and the missing Vulkan loader or NVIDIA driver when a
    Windows GPU module does not load; the docs now list the latest Visual C++
    v14 Redistributable as a Windows requirement
    (#788).
  • Add desktop-model settings to experimental image generation: an llm
    text-encoder file for Z-Image and Qwen-Image, sampler, scheduler and flow
    shift on the request and model defaults, and flash attention and direct
    VAE convolutions, which now turn on automatically where measured faster or
    smaller, such as a 1024x1024 Vulkan decode in about 1 s instead of up to
    56 s in stable-diffusion.cpp's native CLI
    (#802).
  • Name the missing VAE or text-encoder role when stable-diffusion.cpp
    rejects a split image checkpoint, instead of a generic load error
    (#802).
  • Check image models against Metal's recommended GPU working set on macOS,
    skip the check on Vulkan instead of comparing with host memory, and raise
    the estimate's fixed allowance to 512 MiB so it covers measured CPU peaks
    (#802).
  • Check image models on Android against the larger of MemAvailable and
    half of physical memory less the app's own memory, so SD-Turbo loads on
    8 GB phones whose MemAvailable leaves out memory the system frees on
    demand; 6 GB phones still refuse it in most cases
    (#792).
  • Add experimental desktop image presets: SDXL-Lightning, FLUX.1-schnell,
    SD 3.5 Large Turbo and Z-Image-Turbo, which generate and warm up at
    1024x1024 when a request or warmUp leaves the size unset, and the basic
    example's image CLI downloads them
    (#802).
  • Aligned the default WebGPU bridge assets to v0.1.54, unchanged from 0.9.0:
    they embed llama.cpp v0.5.0, are qualified against native v0.5.0, and keep
    Web/native llama.cpp v0.5.0@7fe450e19305b828c199d602c23a8337aaa1f03b parity
    and Web @litert-lm/core@0.15.0. Immutable Web asset manifest:
    8a9f83c15035eeb034a6563e6f753382d7d7f9be81503ef76902138da7841176.