Skip to content

v0.9.0

Choose a tag to compare

@github-actions github-actions released this 27 Sep 19:23
· 95 commits to main since this release
e1424dc
  • Document generic JSON tool calling as an intentional fallback, including
    its prompt and model-reliability limits; runtime behavior is unchanged
    (#755).
  • Load Qwen3.5-0.8B on the Web CPU (WebAssembly) backend at the default
    contextSize using smaller processing batches, while preserving
    full-context defaults for unknown models, including embedding models
    (#752).
  • Report a failed Web model load on a page without cross-origin isolation as
    LlamaModelException with its real cause, not as a COOP/COEP
    worker-thread error; only a real worker-thread failure still names COOP/COEP
    (#753).
  • Log a Dart warning when an explicit preferredBackend GPU module is not
    bundled and the model loads on CPU instead, as happens for cuda with the
    default Windows bundle; the native runtime docs now say when CUDA is bundled
    (#756).
  • Fix the llamadart_server example exiting at startup on Windows; it stops
    on Ctrl+C there, and on SIGINT or SIGTERM elsewhere
    (#757).
  • Accept MP3 and FLAC bytes, as well as WAV, for Qwen3-ASR speech to text on
    Web (#723).
  • Apply presencePenalty, minP and thinkingBudget, and runtime LoRA
    adapters (setLora, removeLora, clearLoras), on WebGPU with bridge
    assets whose capability probes report them; other assets still reject them
    (#722).
  • Run speculative decoding on WebGPU with bridge assets whose capability
    probe reports the strategy, and report each runtime's strategies in
    backendGenerationCapabilities.speculativeDecodingStrategies; other assets
    still reject it (#722).
  • Reject a non-zero GenerationParams.minP on WebGPU when the bridge lacks
    Min-P, with LlamaUnsupportedException instead of ignoring it, and ignore a
    stop sequence equal to a preservedTokens entry there, as native llama.cpp
    does (#661).
  • Add LlamaEngine.backendGenerationCapabilities, which reports whether the
    loaded runtime applies presencePenalty, minP and thinkingBudget; the
    example chat app uses it to send Min-P and enable its slider only where
    supported (#661).
  • Return a DeepSeek V3 forced-open thought that never closes as reasoning, as
    llama.cpp does with the DeepSeek V3.1 template, instead of as content with
    its tool calls (#743).
  • Keep escaped \n and \r in Qwen3-Coder XML reasoning, as llama.cpp does
    with the Qwen3.5 template; before, they became line breaks unless a tool
    call ended the thought
    (#743).
  • Stream content and reasoning with the whitespace the non-streamed parse
    keeps, with or without tools, so streamed answers and ChatSession history
    no longer start with the blank lines after </think>. Only whitespace at
    either end, and text that may be a tag or tool-call opening, waits for more
    output, so reasoning still streams token by token. Without tools, Hermes,
    DeepSeek R1, Qwen3-Coder XML and the other formats the template engine
    guide lists also drop a start tag repeated at the start of a forced-open
    thought, as the parse does. The guide lists the exceptions
    (#754).
  • Stream content that equals the non-streamed parse for Qwen3-Coder XML,
    Mistral Nemo and 15 more tool-call formats, and for output parsed with a PEG
    parser, so text before a tool call no longer carries the tool-call envelope
    into streamed content or ChatSession history. Content and reasoning are
    trimmed as for Hermes, and text after a call arrives at the end of the
    stream. See the template engine guide for the formats and exceptions
    (#732).
  • Stream Hermes-format content that equals the non-streamed parse, so text
    before a tool call no longer carries the <tool_call> envelope into
    streamed content or ChatSession history; only a possible envelope opening
    and trailing whitespace wait for more output. With tools, streamed content
    is now trimmed as the parse trims it, and text the parse keeps after a tool
    call, including a malformed envelope, arrives at the end of the stream.
    Streamed reasoning, and so ChatSession thinking, is trimmed per thought as
    the parse trims it. The exception is a forced-open thought that never
    closes and starts with whitespace: the parse keeps it untrimmed, but it
    streams without its leading and trailing whitespace, so
    " \n Hello there. \n\n" streams as "Hello there.". Before, it streamed
    as the parse gives it, except for some thoughts containing a backslash,
    depending on chunking. After a forced-open
    thought, text after a tool call arrives at the end
    (#701).
  • Throw LlamaModelException when a WebGPU model load fails with a bridge
    error that has no specific mapping, and LlamaInferenceException or
    LlamaStateException for such Web embedding, next-token scoring and state
    errors, with URL credentials and signed query values redacted from the
    details and the load-failure console log
    (#704).
  • Keep URL credentials, signed query values and fragments out of
    LlamaEngine model and projector load errors and logs and the model field
    of completion chunks for every URL form, including scheme-relative
    //user:pass@host/... URLs and relative paths with a query, and out of
    native model download errors. A projector load error that is not a
    LlamaException now throws LlamaModelException. The details of a
    model or projector load failure is now a {type, message} map instead of
    the original error, and a native download that fails with a network error
    carries the error text as a String in details
    (#704).
  • Keep the text after a U+0000 in native llama.cpp tokenization, embeddings
    and generation prompts instead of dropping it
    (#608).
  • Make DecisionEngine.load throw LlamaStateException when another model is
    loaded while it runs, even under the same backend handle
    (#626).
  • Bound speech validation pack memory by a footprint counter instead of the
    resident set: phys_footprint on macOS and iOS, RssAnon plus RssShmem
    plus VmSwap on Linux and Android, read after malloc_trim(0) where the C
    library provides it (glibc, not Android), and PrivateUsage plus
    SharedCommitUsage on Windows (PrivateUsage alone on builds without it).
    Evicting file-backed pages, such as the
    mmapped weights, compressing memory under pressure, or glibc keeping freed
    memory across reloads no longer fails peak_memory_bound without memory
    growth, and each report names its counter
    (#633,
    #762).
  • Fail speech validation leak_slope_bound when the least-squares footprint
    slope over cleanup cycles 1-8 exceeds 7 MiB per cycle; it failed only when
    every cycle grew by more than 7 MiB, and passed leaks of 16 MiB per reload
    (#762).
  • Detect chat template capabilities with llama.cpp's probes, and give
    templates that read only typed content text parts, as llama.cpp does:
    SmolVLM prompts keep the message text, Ministral 3 renders an image
    followed by a reasoning-only turn as llama-server does, TranslateGemma 2B
    keeps the text next to an image, and Kimi-K2 tool results after an image
    are plain text
    (#720).
  • Use the LFM2 format for LFM2.5 templates that list tools without
    <|tool_list_start|>, as llama.cpp does, so LFM2.5-1.2B-Instruct and
    LFM2.5-1.2B-Thinking tool prompts drop the stray "Respond in JSON format"
    instruction (#716).
  • Report per-request usage on WebGPU with the newly pinned bridge assets
    v0.1.54, on the final create chunk and to observers
    (#696).
  • Give assistant turns that hold only tool calls or only reasoning empty
    content instead of null in chat templates, as llama.cpp does: QwQ-32B
    renders them instead of throwing, and LFM2 and Devstral prompts drop a stray
    null or <function text>
    (#715).
  • Pass Map and List tool results as compact JSON text to LFM2, gpt-oss, Solar
    Open, Ministral, DeepSeek V3 and TranslateGemma templates too, instead of
    Python-style or spaced text
    (#717).
  • Pass earlier tool-call arguments as JSON objects to templates that read them
    as objects, as llama.cpp does, so Qwen3, Ministral, Devstral, gpt-oss and
    similar prompts format them with the template's own JSON spacing
    (#702).
  • Require dinja 1.2.0, so more chat prompts match llama.cpp: tojson
    output such as tool declarations uses llama.cpp's spacing, number format
    and non-ASCII text; Qwen3-Coder, GLM-4.6, GLM-4.7-Flash, MiniMax-M2,
    Nemotron-3-Nano, Command R7B, Cohere2 MoE and Ling 3.0 prompts lose stray
    indentation; Functionary v3.1 adds no tool instructions without tools;
    Granite 3.3 spells out the month in its date; Hunyuan Hy3 keeps the system
    prompt first instead of merging it into the user turn; and Bielik 11B v3
    tool-call turns without text render instead of throwing.
  • Count generated tokens with an empty text piece in llama.cpp
    getPerformanceContext() evalTokens and sampleCount without speculative
    decoding, as the speculative path already did
    (#706).
  • Report per-request token usage and timings on the final create chunk as
    LlamaCompletionChunk.usage on native llama.cpp
    (#696).
  • Add LlamaEngine(observers: ...), which reports chat and text completions,
    embeddings and model loads, with their usage, to tracing and metrics code
    (#696).
  • Updated the default llama.cpp native runtime pin to
    leehack/llamadart-native@v0.5.0 (llama.cpp v0.5.0) with Apple companion
    0.0.20, regenerated matching Dart FFI bindings, refreshed the llamadart_llama_cpp_flutter Apple SwiftPM
    checksum, and aligned current README/website native override docs.
  • Add LlamaEngine.scoreNextToken(...) for next-token log-probabilities on
    native llama.cpp and WebGPU bridge assets v0.1.52+, matching llama-server
    n_probs; check
    supportsNextTokenScoring first
    (#694).

  • Add example/laya_command_bar, a Flutter text field that reshapes into a
    reminder, message, calculation or other command as you type, read by a
    Laya decision model, by EmbeddingGemma and labelled examples, or by small
    LLMs' next-token scores, including the decision model decider-2b.

  • Throw LlamaModelException when native llama.cpp cannot find or load a
    multimodal projector, and LlamaUnsupportedException when the runtime lacks
    the mtmd functions; LlamaEngine.supportsAudio also throws the latter.
    Speech-to-text capabilities now say when no projector is loaded, using the
    new LlamaEngine.hasMultimodalProjector
    (#325).

  • Honour LlamaEngine.cancelGeneration() issued right after listening to a
    create, generate or ChatSession.create stream, before it reaches the
    backend, instead of running the whole generation
    (#602).

  • Cancel an active text-to-speech synthesis on LlamaEngine.unloadModel() and
    dispose() instead of waiting for it to finish
    (#628).

  • Cancel an active Qwen3-ASR transcription on LlamaEngine.unloadModel() and
    dispose() instead of completing it with the transcript cut at the unload
    (#670).

  • Send LiteRT-LM tool calls and tool results in the runtime's own message
    format, so Gemma 4 reads tool output and Qwen3 tool histories no longer
    fail
    (#681).

  • Start a native llama.cpp generation requested right after a cancel once the
    cancelled run stops, instead of failing with generation is already in progress. An overlap with a running generation that was not cancelled now
    throws LlamaStateException
    (#655).

  • Render Qwen3 prompts as llama.cpp does: an earlier assistant tool-call
    turn without reasoning no longer gets an empty <think> block
    (#691).

  • Require dinja 1.1.0. Its Jinja string comparison makes three more chat
    templates render as llama.cpp does: MiniMax-M1 adds no empty
    system block for an empty or whitespace-only system message; NVIDIA
    Nemotron Nano v2 drops the blank line before a tool call, the blank lines
    before its tool instructions when tools come with an empty or
    whitespace-only system message, and an empty final assistant turn without
    a generation prompt; and Functionary v3.2 tool declarations drop stray
    // Format=<|NONE|> lines and spell out nested object parameters
    (#351).

  • Cancel a generation's backend run as soon as its stream subscription is
    cancelled, instead of at its next token, which during prompt evaluation
    meant after the whole prompt. A native llama.cpp generation requested right
    after such a cancel now waits for it instead of throwing
    LlamaStateException, and native llama.cpp sees a cancel between text
    prompt micro-batches (ModelParams.microBatchSize, 512 tokens by default)
    or, with speculative decoding, between batches (ModelParams.batchSize)
    (#663,
    #660).

  • Render every result of a tool message holding several
    LlamaToolResultContent parts, as one tool message per result like
    llama.cpp, instead of only the first; LlamaChatMessage.toJson lists them
    all (#683).

  • Render Gemma 4 tool calls and tool results as llama.cpp does, so Gemma 4
    GGUF models can read tool output
    (#669).

  • Stop a Qwen3-TTS audio decode at its next chunk boundary when native
    text-to-speech is cancelled, instead of finishing the native step in
    progress first. This needs llamadart-native v0.4.1-1 or later; older
    runtimes keep the previous behaviour
    (llamadart-native#86,
    #322).

  • Add an experimental DecisionEngine for Laya-style decision models (a
    ModernBERT encoder GGUF plus a safetensors head) on native llama.cpp, with
    typed ChoiceKey, ScoreKey and NoulKey questions
    (#604).

  • Add example/basic_app/bin/llamadart_decision_example.dart, a console demo
    that triages a support ticket with DecisionEngine
    (#604).

  • Add example/laya_tetris, a Flutter app in which a Laya decision model
    plays real-time Tetris through DecisionEngine
    (#604).

  • Run example/laya_tetris on Web through the WebGPU bridge, with a live demo
    at https://leehack-flutter-laya-tetris.static.hf.space.

  • Add a notebook in example/laya_tetris/training/ that fine-tunes a Laya
    decision head for the Tetris example and exports it for DecisionEngine
    (#604).

  • Run DecisionEngine on WebGPU through the bridge decision API
    (apiVersion 1), which bridge assets v0.1.47+ include
    (#604).

  • Stop native image and audio requests from seeding the repeat penalty with
    leftover memory, which made output depend on the previous request
    (#603).

  • After a failed native prompt decode, the next reusePromptPrefix request no
    longer runs on the wrong KV cache or keeps failing
    (#601).

  • Reject embed() and embedBatch() on rank-pooled reranker GGUFs with
    LlamaUnsupportedException on native, instead of returning memory read
    past llama.cpp's classifier-score buffer
    (#583).

  • Throw LlamaInferenceException from native embed() and embedBatch()
    when input to an encoder-only model or a model without a KV cache (such as
    BERT-family and ModernBERT GGUFs) does not fit one microBatchSize pass,
    instead of aborting the process or embedding only the last chunk
    (#607).

  • Aligned the default WebGPU bridge assets to v0.1.54 for the decision API,
    next-token scoring, presence penalty, Min-P, thinking budgets, runtime LoRA
    adapters, speculative decoding and the Web runtime fixes below; the bridge
    also adds its supportsCompletionUsage flag, which llamadart does not use yet
    (#729). The assets embed llama.cpp v0.5.0, are
    qualified against native v0.5.0, and keep Web/native llama.cpp
    v0.5.0@7fe450e19305b828c199d602c23a8337aaa1f03b parity and Web
    @litert-lm/core@0.15.0. Immutable Web asset manifest:
    8a9f83c15035eeb034a6563e6f753382d7d7f9be81503ef76902138da7841176.
  • On Web, an invalid GBNF grammar now fails generation with a
    LlamaInferenceException whose details contain (invalid grammar), and the
    loaded model stays usable, instead of aborting the WebGPU bridge runtime
    (llama-web-bridge#125).
  • On Web, when a bridge worker fails and the main-thread reload fails too, or
    a replacement worker cannot start, the bridge now forgets the model instead
    of being left broken with a TypeError. Later calls fail with
    No model loaded. Call loadModelFromUrl first.; call unloadModel(), then
    loadModel() and any projector again to recover
    (llama-web-bridge#123,
    llama-web-bridge#127).
  • On Web, a replacement bridge worker reloads the current model before its next
    request
    (llama-web-bridge#126).
  • On Web, grammar-constrained generation no longer aborts the bridge runtime
    when top-k or top-p keeps only tokens the grammar rejects; it resamples as
    llama.cpp does, and fails with Grammar rejected every candidate token only
    when the grammar cannot continue
    (llama-web-bridge#118).
  • On Web, an ordinary error from a healthy bridge worker, such as a prompt that
    overflows the context or empty embedding input, is now rethrown with the
    worker kept, instead of moving the session to the main thread for good and
    re-running the request
    (llama-web-bridge#119,
    llama-web-bridge#120).
  • Extend the GGUF speech-to-text validation pack with four synthetic edge
    fixtures built in-process, so no extra audio is stored: generated digital
    silence, plus a truncated RIFF, a stereo 44.1 kHz re-encode and a 33-second
    concatenation, the last three derived from the locked jfk.wav. Every GGUF
    STT pack run executes the four checks and each one gates functional_pass;
    the LiteRT-ASR and TTS packs pass no edge fixtures
    (#325).
  • Raise every stt, tts and litert-asr speech validation pack run from 8
    to 15 lifecycle checks: an immediate cancel, three
    cancel/dispose/load/generate cycles, and bounds that fail the run when a
    cancellation takes over 500 ms to end its task or the peak resident set
    exceeds 1.10x the one sampled after the first generation
    (#594).
  • Add cross-platform validation cases for a cancel issued right after
    listening, a generation requested right after a cancel, an overlapping
    generation, an invalid GBNF grammar and ToolChoice.auto on a prompt that
    needs no tool, and run the tool cases on the GGUF chat profiles
    (#602,
    #655,
    #654).
  • Add decision-gguf-{cpu,metal,vulkan,cuda,webgpu} validation profiles that
    check DecisionEngine token ids, raw logits and answers against the Laya
    0.3.5 reference, plus batching, reload and typed rejections, on desktop,
    mobile, Web WebGPU and GCE CUDA
    (#604).
  • Add a Web-only chat-gguf-webgpu validation profile, hash Web validation
    models while they stream so GGUFs over 2 GiB pass preparation, list every
    bundled profile in the validation app, and verify iOS GGUF GPU placement
    from the XCTest console log.
  • Run eight cleanup cycles instead of three in every speech validation pack,
    and fail a run whose resident set grows by more than 7 MiB in each of seven
    warm cycles. The 1.10x peak ratio no longer applies on Linux CUDA, where reload
    overhead that levels off failed it without a leak
    (#686).
  • Add GGUF speech validation pack checks: tts unloads and disposes the engine
    during a synthesis, cancels one during its audio decode and bounds the
    resident set those checks add, and stt must fail with
    LlamaSpeechTranscriptTruncatedException at maxOutputTokens and at the
    context size. stt runs now execute 28 checks and tts runs 25
    (#628,
    #636,
    #322).
  • Force greedy topK: 1 for zero-temperature LiteRT-LM Web generation, matching
    the native clamp
    (#548).
  • Log Model … loaded from …; native engine creation is deferred until the first generation or tokenizer call instead of loaded successfully when
    the native LiteRT-LM backend finishes loadModel, since it creates the
    engine lazily
    (#569).
  • Require dxcompiler.dll and dxil.dll in the Windows x64 LiteRT-LM runtime
    cache and desktop validation bundle checks, matching the hook's v0.17.0-6
    inventory. The runtime does not preload them: Dawn's D3D12 backend loads the
    pair at GPU engine creation
    (#570).
  • Document that Linux llama.cpp loads need the OpenMP runtime
    (libgomp.so.1; libgomp1 on Ubuntu/Debian, libgomp on Fedora and Arch)
    and that Linux LiteRT-LM GPU needs a hardware Vulkan ICD: with only Mesa
    llvmpipe the runtime segfaults after model load instead of failing cleanly
    (llamadart-native#82,
    #572).
  • Record one startup diagnostic when every candidate of a native backend
    module family fails to load, or when a ggml/wrapper symbol is missing from
    both the primary FFI asset and every fallback library. Candidates are named
    by asset URI or file name only and loader errors are classified, never
    quoted, so no directory or loader search path reaches the diagnostic
    (#416).
  • Forward llama.cpp and LiteRT-LM worker-isolate log records to the
    LlamaEngine.configureLogging handler. A worker takes the Dart logger level
    when it starts and LlamaEngine.setDartLogLevel/setLogLevel update a
    running worker; the default none sends nothing and debug records are
    capped at 1000 per worker. The LiteRT-LM program-cache pruning warnings are
    now ordinary warn records gated by that level instead of the native log
    level. Adds LlamaLogger.level and the BackendDartLogLevel capability
    (#567).
  • Replace the token in Bearer <token> and the value in token=, key=,
    secret=, password=, api_key= and apikey=<value> outside HTTP URLs in
    native startup diagnostics with <redacted-secret>; URL and
    control-character handling is unchanged
    (#551).
  • Move native release pins (llama.cpp tag, LiteRT-LM tag and per-bundle
    checksums) from hook/build.dart into
    lib/src/hook/native_release_pins.dart, the only file the pin sync now
    rewrites.
  • Add ModelParams.liteRtLmCacheDir to choose the native LiteRT-LM runtime
    cache directory and opt-in ModelParams.liteRtLmMaxProgramCacheBytes, which
    deletes *_mldrift_program_cache.bin files above the cap before each engine
    create and logs a warning per deleted file. Defaults are unchanged: the same
    per-platform directory and no pruning
    (#552).
  • Force greedy topK: 1 for zero-temperature
    LiteRtLmRuntimeClient.createConversation calls, which returned incoherent
    text on the LiteRT WebGPU sampler with the default top-k.
  • Keep root-cause native startup diagnostics when the buffer or the rendered
    startupDiagnostics=[...] suffix overflows: teardown entries, now prefixed
    teardown: , are dropped first, duplicates are recorded once, entries are
    capped at 2048 characters, and each omitted run renders as ...
    (#415).
  • Skip the Windows altered-search-path preload for wrapper library candidates
    whose absolute path does not exist, so lazy wrapper API lookups no longer
    record a Failed to preload Windows backend module startup diagnostic per
    missing candidate
    (#550).
  • Accept promptTemplate on the non-native
    LiteRtLmRuntimeClient.createConversation placeholder, so callers passing
    it compile for Web as they do on native
    (#549).
  • Cache TemplateCaps.detect results in a per-isolate LRU keyed by exact
    template source and bounded at 16 entries, so repeated chat-template renders
    skip both Jinja parses and all four capability probes. Detections in which
    any analysis step failed are not cached and keep logging on every call
    (#448).
  • Detect supportsTools and supportsToolCalls for chat templates that
    reject two tool calls in one assistant message (Llama 3.2) or a user turn
    directly after a tool call (Ministral 3). The tools capability probe now
    renders a single tool call, and a separate parallel probe clears only
    supportsParallelToolCalls when it throws
    (#557).
  • Limit the Ministral tool-call grammar to a single [TOOL_CALLS] block unless
    parallel tool calls are enabled; it previously always allowed repeats while
    the parser kept only the first call
    (#559).
  • Limit the Nemotron v3 tool-call grammar (Qwen3-Coder XML format) to a single
    tool call unless parallel tool calls are enabled; it previously always
    allowed repeats while the parser kept only one call
    (#562).
  • Pin the WebGPU model-load retry ladder with browser tests for the ladder
    advance, the wasm64 BigInt restart on wasm32, the restart without the remote
    fetch backend, and forced remote-fetch chunk halving stopping on both its
    ten-restart cap and its 4 KiB minimum chunk, then collapse the duplicated
    attempt thread-count switch into one helper
    (#361).
  • Apply the JSON Schema pattern keyword when generating GBNF, for anchored
    patterns built from literals, positive character classes, (...) groups
    nested at most 32 deep, grouped alternation and */+/?/{m,n}
    repetition; any other pattern, including deeper nesting, falls back to the
    rule the schema would have produced without it, so minLength and
    maxLength still apply there. An applied pattern replaces
    minLength/maxLength as it does in llama.cpp, so a schema carrying both
    pattern and maxLength is no longer length-bounded. A schema carrying
    pattern but no explicit type now yields a string rule instead of
    throwing Unrecognized schema. Mistral Nemo and Magistral tool-call ids are
    now grammar-constrained to exactly nine alphanumerics
    (#582).
  • Stop a cancelled llama.cpp image or audio prompt at the next prompt chunk,
    or between a media chunk's encode and its embedding decode, instead of after
    the whole prompt is ingested. The native call already running still
    finishes
    (#599).
  • SpeechToTextEngine now fails a native Qwen3-ASR transcript that reaches
    the context size or maxOutputTokens with
    LlamaSpeechTranscriptTruncatedException instead of completing with
    truncated text, and create() reports finishReason: 'length' when native
    llama.cpp stops at either limit. The chat app caps Qwen3-ASR recordings at
    the validated 30 seconds
    (#636).
  • Stop ToolChoice.auto on WebGPU from forcing a tool call: it now skips
    the lazy tool-call grammar and parses tool calls best-effort. WebGPU
    rejects GenerationParams.grammarLazy and a non-root grammarRoot with
    LlamaUnsupportedException; backends report this through the new
    BackendLazyGrammarSupport
    (#654).
  • Name the CUDA 12 runtime libraries (libcudart.so.12, libcublas.so.12)
    that the Linux cuda backend needs on the default loader path; llamadart
    does not ship them, and validation bundles refuse LD_LIBRARY_PATH
    (#587).
  • Remote validation runs resolve packages/llamadart_validation before
    building the report, instead of reporting FAILED with a null error on a
    fresh checkout. A failed report step is now the run's error, with its exit
    code and a redacted stderr tail
    (#688).
  • Select the devices of an explicit GpuBackend.metal or GpuBackend.hip
    on llama.cpp: they looked up ggml registries named Metal and HIP, but
    ggml names them MTL and ROCm, so loading fell back to automatic device
    selection. A HIP load on a ROCm build now reports its backend as HIP
    instead of CPU
    (#611).
  • Report a WebGPU model load that fails with error 138 as the documented
    cross-origin isolation (COOP/COEP) UnsupportedError, as
    thread constructor failed already was, instead of rethrowing the raw
    bridge error
    (#598).
  • Throw LlamaModelException when WebGPU cannot fetch or load a multimodal
    projector, instead of the raw JavaScript error. Its details drop these
    parts of the projector URL the app passed, as written, JSON-escaped,
    percent-encoded or percent-decoded: the userinfo and password, as whole
    tokens of any length; the ?query and #fragment, where they directly
    follow a non-space character; the query and each &-separated part that
    contain =, as whole tokens; and bare query values and the fragment of 10
    or more characters, as whole tokens. A whole token has no ASCII letter or
    digit directly before or after it. Shorter bare values printed on their
    own, such as the 1 of ?v=1, stay. Other URLs in the details lose
    userinfo, query and fragment on a best-effort basis
    (#642).
  • Leave no envelope text in the parsed content when Qwen2.5 wraps a Hermes
    tool call in double braces (<tool_call>{{"name": ...}}</tool_call>, with
    any number of extra closing braces) without a grammar. Calls are extracted as
    before; a double-brace call with other malformed envelope text keeps that
    text. This deliberately differs from upstream llama.cpp, which rejects that
    output and extracts no call
    (#662).