Repository navigation
v0.9.0
- Document generic JSON tool calling as an intentional fallback, including
its prompt and model-reliability limits; runtime behavior is unchanged
(#755). - Load Qwen3.5-0.8B on the Web CPU (WebAssembly) backend at the default
contextSizeusing smaller processing batches, while preserving
full-context defaults for unknown models, including embedding models
(#752). - Report a failed Web model load on a page without cross-origin isolation as
LlamaModelExceptionwith its real cause, not as a COOP/COEP
worker-thread error; only a real worker-thread failure still names COOP/COEP
(#753). - Log a Dart warning when an explicit
preferredBackendGPU module is not
bundled and the model loads on CPU instead, as happens forcudawith the
default Windows bundle; the native runtime docs now say when CUDA is bundled
(#756). - Fix the
llamadart_serverexample exiting at startup on Windows; it stops
on Ctrl+C there, and on SIGINT or SIGTERM elsewhere
(#757). - Accept MP3 and FLAC bytes, as well as WAV, for Qwen3-ASR speech to text on
Web (#723). - Apply
presencePenalty,minPandthinkingBudget, and runtime LoRA
adapters (setLora,removeLora,clearLoras), on WebGPU with bridge
assets whose capability probes report them; other assets still reject them
(#722). - Run speculative decoding on WebGPU with bridge assets whose capability
probe reports the strategy, and report each runtime's strategies in
backendGenerationCapabilities.speculativeDecodingStrategies; other assets
still reject it (#722). - Reject a non-zero
GenerationParams.minPon WebGPU when the bridge lacks
Min-P, withLlamaUnsupportedExceptioninstead of ignoring it, and ignore a
stop sequence equal to apreservedTokensentry there, as native llama.cpp
does (#661). - Add
LlamaEngine.backendGenerationCapabilities, which reports whether the
loaded runtime appliespresencePenalty,minPandthinkingBudget; the
example chat app uses it to send Min-P and enable its slider only where
supported (#661). - Return a DeepSeek V3 forced-open thought that never closes as reasoning, as
llama.cpp does with the DeepSeek V3.1 template, instead of as content with
its tool calls (#743). - Keep escaped
\nand\rin Qwen3-Coder XML reasoning, as llama.cpp does
with the Qwen3.5 template; before, they became line breaks unless a tool
call ended the thought
(#743). - Stream content and reasoning with the whitespace the non-streamed parse
keeps, with or without tools, so streamed answers andChatSessionhistory
no longer start with the blank lines after</think>. Only whitespace at
either end, and text that may be a tag or tool-call opening, waits for more
output, so reasoning still streams token by token. Without tools, Hermes,
DeepSeek R1, Qwen3-Coder XML and the other formats the template engine
guide lists also drop a start tag repeated at the start of a forced-open
thought, as the parse does. The guide lists the exceptions
(#754). - Stream content that equals the non-streamed parse for Qwen3-Coder XML,
Mistral Nemo and 15 more tool-call formats, and for output parsed with a PEG
parser, so text before a tool call no longer carries the tool-call envelope
into streamed content orChatSessionhistory. Content and reasoning are
trimmed as for Hermes, and text after a call arrives at the end of the
stream. See the template engine guide for the formats and exceptions
(#732). - Stream Hermes-format content that equals the non-streamed parse, so text
before a tool call no longer carries the<tool_call>envelope into
streamed content orChatSessionhistory; only a possible envelope opening
and trailing whitespace wait for more output. With tools, streamed content
is now trimmed as the parse trims it, and text the parse keeps after a tool
call, including a malformed envelope, arrives at the end of the stream.
Streamed reasoning, and soChatSessionthinking, is trimmed per thought as
the parse trims it. The exception is a forced-open thought that never
closes and starts with whitespace: the parse keeps it untrimmed, but it
streams without its leading and trailing whitespace, so
" \n Hello there. \n\n"streams as"Hello there.". Before, it streamed
as the parse gives it, except for some thoughts containing a backslash,
depending on chunking. After a forced-open
thought, text after a tool call arrives at the end
(#701). - Throw
LlamaModelExceptionwhen a WebGPU model load fails with a bridge
error that has no specific mapping, andLlamaInferenceExceptionor
LlamaStateExceptionfor such Web embedding, next-token scoring and state
errors, with URL credentials and signed query values redacted from the
details and the load-failure console log
(#704). - Keep URL credentials, signed query values and fragments out of
LlamaEnginemodel and projector load errors and logs and themodelfield
of completion chunks for every URL form, including scheme-relative
//user:pass@host/...URLs and relative paths with a query, and out of
native model download errors. A projector load error that is not a
LlamaExceptionnow throwsLlamaModelException. Thedetailsof a
model or projector load failure is now a{type, message}map instead of
the original error, and a native download that fails with a network error
carries the error text as aStringindetails
(#704). - Keep the text after a U+0000 in native llama.cpp tokenization, embeddings
and generation prompts instead of dropping it
(#608). - Make
DecisionEngine.loadthrowLlamaStateExceptionwhen another model is
loaded while it runs, even under the same backend handle
(#626). - Bound speech validation pack memory by a footprint counter instead of the
resident set:phys_footprinton macOS and iOS,RssAnonplusRssShmem
plusVmSwapon Linux and Android, read aftermalloc_trim(0)where the C
library provides it (glibc, not Android), andPrivateUsageplus
SharedCommitUsageon Windows (PrivateUsagealone on builds without it).
Evicting file-backed pages, such as the
mmapped weights, compressing memory under pressure, or glibc keeping freed
memory across reloads no longer failspeak_memory_boundwithout memory
growth, and each report names its counter
(#633,
#762). - Fail speech validation
leak_slope_boundwhen the least-squares footprint
slope over cleanup cycles 1-8 exceeds 7 MiB per cycle; it failed only when
every cycle grew by more than 7 MiB, and passed leaks of 16 MiB per reload
(#762). - Detect chat template capabilities with llama.cpp's probes, and give
templates that read only typed content text parts, as llama.cpp does:
SmolVLM prompts keep the message text, Ministral 3 renders an image
followed by a reasoning-only turn as llama-server does, TranslateGemma 2B
keeps the text next to an image, and Kimi-K2 tool results after an image
are plain text
(#720). - Use the LFM2 format for LFM2.5 templates that list tools without
<|tool_list_start|>, as llama.cpp does, so LFM2.5-1.2B-Instruct and
LFM2.5-1.2B-Thinking tool prompts drop the stray "Respond in JSON format"
instruction (#716). - Report per-request usage on WebGPU with the newly pinned bridge assets
v0.1.54, on the finalcreatechunk and to observers
(#696). - Give assistant turns that hold only tool calls or only reasoning empty
content instead ofnullin chat templates, as llama.cpp does: QwQ-32B
renders them instead of throwing, and LFM2 and Devstral prompts drop a stray
nullor<function text>
(#715). - Pass Map and List tool results as compact JSON text to LFM2, gpt-oss, Solar
Open, Ministral, DeepSeek V3 and TranslateGemma templates too, instead of
Python-style or spaced text
(#717). - Pass earlier tool-call arguments as JSON objects to templates that read them
as objects, as llama.cpp does, so Qwen3, Ministral, Devstral, gpt-oss and
similar prompts format them with the template's own JSON spacing
(#702). - Require
dinja1.2.0, so more chat prompts match llama.cpp:tojson
output such as tool declarations uses llama.cpp's spacing, number format
and non-ASCII text; Qwen3-Coder, GLM-4.6, GLM-4.7-Flash, MiniMax-M2,
Nemotron-3-Nano, Command R7B, Cohere2 MoE and Ling 3.0 prompts lose stray
indentation; Functionary v3.1 adds no tool instructions without tools;
Granite 3.3 spells out the month in its date; Hunyuan Hy3 keeps the system
prompt first instead of merging it into the user turn; and Bielik 11B v3
tool-call turns without text render instead of throwing. - Count generated tokens with an empty text piece in llama.cpp
getPerformanceContext()evalTokensandsampleCountwithout speculative
decoding, as the speculative path already did
(#706). - Report per-request token usage and timings on the final
createchunk as
LlamaCompletionChunk.usageon native llama.cpp
(#696). - Add
LlamaEngine(observers: ...), which reports chat and text completions,
embeddings and model loads, with their usage, to tracing and metrics code
(#696).
- Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@v0.5.0(llama.cppv0.5.0) with Apple companion
0.0.20, regenerated matching Dart FFI bindings, refreshed thellamadart_llama_cpp_flutterApple SwiftPM
checksum, and aligned current README/website native override docs.
-
Add
LlamaEngine.scoreNextToken(...)for next-token log-probabilities on
native llama.cpp and WebGPU bridge assetsv0.1.52+, matching llama-server
n_probs; check
supportsNextTokenScoringfirst
(#694). -
Add
example/laya_command_bar, a Flutter text field that reshapes into a
reminder, message, calculation or other command as you type, read by a
Laya decision model, by EmbeddingGemma and labelled examples, or by small
LLMs' next-token scores, including the decision model decider-2b. -
Throw
LlamaModelExceptionwhen native llama.cpp cannot find or load a
multimodal projector, andLlamaUnsupportedExceptionwhen the runtime lacks
the mtmd functions;LlamaEngine.supportsAudioalso throws the latter.
Speech-to-text capabilities now say when no projector is loaded, using the
newLlamaEngine.hasMultimodalProjector
(#325). -
Honour
LlamaEngine.cancelGeneration()issued right after listening to a
create,generateorChatSession.createstream, before it reaches the
backend, instead of running the whole generation
(#602). -
Cancel an active text-to-speech synthesis on
LlamaEngine.unloadModel()and
dispose()instead of waiting for it to finish
(#628). -
Cancel an active Qwen3-ASR transcription on
LlamaEngine.unloadModel()and
dispose()instead of completing it with the transcript cut at the unload
(#670). -
Send LiteRT-LM tool calls and tool results in the runtime's own message
format, so Gemma 4 reads tool output and Qwen3 tool histories no longer
fail
(#681). -
Start a native llama.cpp generation requested right after a cancel once the
cancelled run stops, instead of failing withgeneration is already in progress. An overlap with a running generation that was not cancelled now
throwsLlamaStateException
(#655). -
Render Qwen3 prompts as llama.cpp does: an earlier assistant tool-call
turn without reasoning no longer gets an empty<think>block
(#691). -
Require
dinja1.1.0. Its Jinja string comparison makes three more chat
templates render as llama.cpp does: MiniMax-M1 adds no empty
system block for an empty or whitespace-only system message; NVIDIA
Nemotron Nano v2 drops the blank line before a tool call, the blank lines
before its tool instructions when tools come with an empty or
whitespace-only system message, and an empty final assistant turn without
a generation prompt; and Functionary v3.2 tool declarations drop stray
// Format=<|NONE|>lines and spell out nested object parameters
(#351). -
Cancel a generation's backend run as soon as its stream subscription is
cancelled, instead of at its next token, which during prompt evaluation
meant after the whole prompt. A native llama.cpp generation requested right
after such a cancel now waits for it instead of throwing
LlamaStateException, and native llama.cpp sees a cancel between text
prompt micro-batches (ModelParams.microBatchSize, 512 tokens by default)
or, with speculative decoding, between batches (ModelParams.batchSize)
(#663,
#660). -
Render every result of a tool message holding several
LlamaToolResultContentparts, as onetoolmessage per result like
llama.cpp, instead of only the first;LlamaChatMessage.toJsonlists them
all (#683). -
Render Gemma 4 tool calls and tool results as llama.cpp does, so Gemma 4
GGUF models can read tool output
(#669). -
Stop a Qwen3-TTS audio decode at its next chunk boundary when native
text-to-speech is cancelled, instead of finishing the native step in
progress first. This needs llamadart-native v0.4.1-1 or later; older
runtimes keep the previous behaviour
(llamadart-native#86,
#322). -
Add an experimental
DecisionEnginefor Laya-style decision models (a
ModernBERT encoder GGUF plus a safetensors head) on native llama.cpp, with
typedChoiceKey,ScoreKeyandNoulKeyquestions
(#604). -
Add
example/basic_app/bin/llamadart_decision_example.dart, a console demo
that triages a support ticket withDecisionEngine
(#604). -
Add
example/laya_tetris, a Flutter app in which a Laya decision model
plays real-time Tetris throughDecisionEngine
(#604). -
Run
example/laya_tetrison Web through the WebGPU bridge, with a live demo
at https://leehack-flutter-laya-tetris.static.hf.space. -
Add a notebook in
example/laya_tetris/training/that fine-tunes a Laya
decision head for the Tetris example and exports it forDecisionEngine
(#604). -
Run
DecisionEngineon WebGPU through the bridge decision API
(apiVersion 1), which bridge assetsv0.1.47+include
(#604). -
Stop native image and audio requests from seeding the repeat penalty with
leftover memory, which made output depend on the previous request
(#603). -
After a failed native prompt decode, the next
reusePromptPrefixrequest no
longer runs on the wrong KV cache or keeps failing
(#601). -
Reject
embed()andembedBatch()on rank-pooled reranker GGUFs with
LlamaUnsupportedExceptionon native, instead of returning memory read
past llama.cpp's classifier-score buffer
(#583). -
Throw
LlamaInferenceExceptionfrom nativeembed()andembedBatch()
when input to an encoder-only model or a model without a KV cache (such as
BERT-family and ModernBERT GGUFs) does not fit onemicroBatchSizepass,
instead of aborting the process or embedding only the last chunk
(#607).
- Aligned the default WebGPU bridge assets to
v0.1.54for the decision API,
next-token scoring, presence penalty, Min-P, thinking budgets, runtime LoRA
adapters, speculative decoding and the Web runtime fixes below; the bridge
also adds itssupportsCompletionUsageflag, which llamadart does not use yet
(#729). The assets embed llama.cppv0.5.0, are
qualified against nativev0.5.0, and keep Web/native llama.cpp
v0.5.0@7fe450e19305b828c199d602c23a8337aaa1f03bparity and Web
@litert-lm/core@0.15.0. Immutable Web asset manifest:
8a9f83c15035eeb034a6563e6f753382d7d7f9be81503ef76902138da7841176.
- On Web, an invalid GBNF grammar now fails generation with a
LlamaInferenceExceptionwhose details contain(invalid grammar), and the
loaded model stays usable, instead of aborting the WebGPU bridge runtime
(llama-web-bridge#125). - On Web, when a bridge worker fails and the main-thread reload fails too, or
a replacement worker cannot start, the bridge now forgets the model instead
of being left broken with aTypeError. Later calls fail with
No model loaded. Call loadModelFromUrl first.; callunloadModel(), then
loadModel()and any projector again to recover
(llama-web-bridge#123,
llama-web-bridge#127). - On Web, a replacement bridge worker reloads the current model before its next
request
(llama-web-bridge#126). - On Web, grammar-constrained generation no longer aborts the bridge runtime
when top-k or top-p keeps only tokens the grammar rejects; it resamples as
llama.cpp does, and fails withGrammar rejected every candidate tokenonly
when the grammar cannot continue
(llama-web-bridge#118). - On Web, an ordinary error from a healthy bridge worker, such as a prompt that
overflows the context or empty embedding input, is now rethrown with the
worker kept, instead of moving the session to the main thread for good and
re-running the request
(llama-web-bridge#119,
llama-web-bridge#120). - Extend the GGUF speech-to-text validation pack with four synthetic edge
fixtures built in-process, so no extra audio is stored: generated digital
silence, plus a truncated RIFF, a stereo 44.1 kHz re-encode and a 33-second
concatenation, the last three derived from the lockedjfk.wav. Every GGUF
STT pack run executes the four checks and each one gatesfunctional_pass;
the LiteRT-ASR and TTS packs pass no edge fixtures
(#325). - Raise every
stt,ttsandlitert-asrspeech validation pack run from 8
to 15 lifecycle checks: an immediate cancel, three
cancel/dispose/load/generate cycles, and bounds that fail the run when a
cancellation takes over 500 ms to end its task or the peak resident set
exceeds 1.10x the one sampled after the first generation
(#594). - Add cross-platform validation cases for a cancel issued right after
listening, a generation requested right after a cancel, an overlapping
generation, an invalid GBNF grammar andToolChoice.autoon a prompt that
needs no tool, and run the tool cases on the GGUF chat profiles
(#602,
#655,
#654). - Add
decision-gguf-{cpu,metal,vulkan,cuda,webgpu}validation profiles that
checkDecisionEnginetoken ids, raw logits and answers against the Laya
0.3.5 reference, plus batching, reload and typed rejections, on desktop,
mobile, Web WebGPU and GCE CUDA
(#604). - Add a Web-only
chat-gguf-webgpuvalidation profile, hash Web validation
models while they stream so GGUFs over 2 GiB pass preparation, list every
bundled profile in the validation app, and verify iOS GGUF GPU placement
from the XCTest console log. - Run eight cleanup cycles instead of three in every speech validation pack,
and fail a run whose resident set grows by more than 7 MiB in each of seven
warm cycles. The 1.10x peak ratio no longer applies on Linux CUDA, where reload
overhead that levels off failed it without a leak
(#686). - Add GGUF speech validation pack checks:
ttsunloads and disposes the engine
during a synthesis, cancels one during its audio decode and bounds the
resident set those checks add, andsttmust fail with
LlamaSpeechTranscriptTruncatedExceptionatmaxOutputTokensand at the
context size.sttruns now execute 28 checks andttsruns 25
(#628,
#636,
#322). - Force greedy
topK: 1for zero-temperature LiteRT-LM Web generation, matching
the native clamp
(#548). - Log
Model … loaded from …; native engine creation is deferred until the first generation or tokenizer callinstead ofloaded successfullywhen
the native LiteRT-LM backend finishesloadModel, since it creates the
engine lazily
(#569). - Require
dxcompiler.dllanddxil.dllin the Windows x64 LiteRT-LM runtime
cache and desktop validation bundle checks, matching the hook's v0.17.0-6
inventory. The runtime does not preload them: Dawn's D3D12 backend loads the
pair at GPU engine creation
(#570). - Document that Linux llama.cpp loads need the OpenMP runtime
(libgomp.so.1;libgomp1on Ubuntu/Debian,libgompon Fedora and Arch)
and that Linux LiteRT-LM GPU needs a hardware Vulkan ICD: with only Mesa
llvmpipe the runtime segfaults after model load instead of failing cleanly
(llamadart-native#82,
#572). - Record one startup diagnostic when every candidate of a native backend
module family fails to load, or when aggml/wrapper symbol is missing from
both the primary FFI asset and every fallback library. Candidates are named
by asset URI or file name only and loader errors are classified, never
quoted, so no directory or loader search path reaches the diagnostic
(#416). - Forward llama.cpp and LiteRT-LM worker-isolate log records to the
LlamaEngine.configureLogginghandler. A worker takes the Dart logger level
when it starts andLlamaEngine.setDartLogLevel/setLogLevelupdate a
running worker; the defaultnonesends nothing anddebugrecords are
capped at 1000 per worker. The LiteRT-LM program-cache pruning warnings are
now ordinarywarnrecords gated by that level instead of the native log
level. AddsLlamaLogger.leveland theBackendDartLogLevelcapability
(#567). - Replace the token in
Bearer <token>and the value intoken=,key=,
secret=,password=,api_key=andapikey=<value>outside HTTP URLs in
native startup diagnostics with<redacted-secret>; URL and
control-character handling is unchanged
(#551). - Move native release pins (llama.cpp tag, LiteRT-LM tag and per-bundle
checksums) fromhook/build.dartinto
lib/src/hook/native_release_pins.dart, the only file the pin sync now
rewrites. - Add
ModelParams.liteRtLmCacheDirto choose the native LiteRT-LM runtime
cache directory and opt-inModelParams.liteRtLmMaxProgramCacheBytes, which
deletes*_mldrift_program_cache.binfiles above the cap before each engine
create and logs a warning per deleted file. Defaults are unchanged: the same
per-platform directory and no pruning
(#552). - Force greedy
topK: 1for zero-temperature
LiteRtLmRuntimeClient.createConversationcalls, which returned incoherent
text on the LiteRT WebGPU sampler with the default top-k. - Keep root-cause native startup diagnostics when the buffer or the rendered
startupDiagnostics=[...]suffix overflows: teardown entries, now prefixed
teardown:, are dropped first, duplicates are recorded once, entries are
capped at 2048 characters, and each omitted run renders as...
(#415). - Skip the Windows altered-search-path preload for wrapper library candidates
whose absolute path does not exist, so lazy wrapper API lookups no longer
record aFailed to preload Windows backend modulestartup diagnostic per
missing candidate
(#550). - Accept
promptTemplateon the non-native
LiteRtLmRuntimeClient.createConversationplaceholder, so callers passing
it compile for Web as they do on native
(#549). - Cache
TemplateCaps.detectresults in a per-isolate LRU keyed by exact
template source and bounded at 16 entries, so repeated chat-template renders
skip both Jinja parses and all four capability probes. Detections in which
any analysis step failed are not cached and keep logging on every call
(#448). - Detect
supportsToolsandsupportsToolCallsfor chat templates that
reject two tool calls in one assistant message (Llama 3.2) or a user turn
directly after a tool call (Ministral 3). The tools capability probe now
renders a single tool call, and a separate parallel probe clears only
supportsParallelToolCallswhen it throws
(#557). - Limit the Ministral tool-call grammar to a single
[TOOL_CALLS]block unless
parallel tool calls are enabled; it previously always allowed repeats while
the parser kept only the first call
(#559). - Limit the Nemotron v3 tool-call grammar (Qwen3-Coder XML format) to a single
tool call unless parallel tool calls are enabled; it previously always
allowed repeats while the parser kept only one call
(#562). - Pin the WebGPU model-load retry ladder with browser tests for the ladder
advance, the wasm64 BigInt restart on wasm32, the restart without the remote
fetch backend, and forced remote-fetch chunk halving stopping on both its
ten-restart cap and its 4 KiB minimum chunk, then collapse the duplicated
attempt thread-count switch into one helper
(#361). - Apply the JSON Schema
patternkeyword when generating GBNF, for anchored
patterns built from literals, positive character classes,(...)groups
nested at most 32 deep, grouped alternation and*/+/?/{m,n}
repetition; any other pattern, including deeper nesting, falls back to the
rule the schema would have produced without it, sominLengthand
maxLengthstill apply there. An applied pattern replaces
minLength/maxLengthas it does in llama.cpp, so a schema carrying both
patternandmaxLengthis no longer length-bounded. A schema carrying
patternbut no explicittypenow yields a string rule instead of
throwingUnrecognized schema. Mistral Nemo and Magistral tool-call ids are
now grammar-constrained to exactly nine alphanumerics
(#582). - Stop a cancelled llama.cpp image or audio prompt at the next prompt chunk,
or between a media chunk's encode and its embedding decode, instead of after
the whole prompt is ingested. The native call already running still
finishes
(#599). SpeechToTextEnginenow fails a native Qwen3-ASR transcript that reaches
the context size ormaxOutputTokenswith
LlamaSpeechTranscriptTruncatedExceptioninstead of completing with
truncated text, andcreate()reportsfinishReason: 'length'when native
llama.cpp stops at either limit. The chat app caps Qwen3-ASR recordings at
the validated 30 seconds
(#636).- Stop
ToolChoice.autoon WebGPU from forcing a tool call: it now skips
the lazy tool-call grammar and parses tool calls best-effort. WebGPU
rejectsGenerationParams.grammarLazyand a non-rootgrammarRootwith
LlamaUnsupportedException; backends report this through the new
BackendLazyGrammarSupport
(#654). - Name the CUDA 12 runtime libraries (
libcudart.so.12,libcublas.so.12)
that the Linuxcudabackend needs on the default loader path; llamadart
does not ship them, and validation bundles refuseLD_LIBRARY_PATH
(#587). - Remote validation runs resolve
packages/llamadart_validationbefore
building the report, instead of reportingFAILEDwith a null error on a
fresh checkout. A failed report step is now the run's error, with its exit
code and a redacted stderr tail
(#688). - Select the devices of an explicit
GpuBackend.metalorGpuBackend.hip
on llama.cpp: they looked up ggml registries namedMetalandHIP, but
ggml names themMTLandROCm, so loading fell back to automatic device
selection. A HIP load on a ROCm build now reports its backend asHIP
instead ofCPU
(#611). - Report a WebGPU model load that fails with
error 138as the documented
cross-origin isolation (COOP/COEP)UnsupportedError, as
thread constructor failedalready was, instead of rethrowing the raw
bridge error
(#598). - Throw
LlamaModelExceptionwhen WebGPU cannot fetch or load a multimodal
projector, instead of the raw JavaScript error. Its details drop these
parts of the projector URL the app passed, as written, JSON-escaped,
percent-encoded or percent-decoded: the userinfo and password, as whole
tokens of any length; the?queryand#fragment, where they directly
follow a non-space character; the query and each&-separated part that
contain=, as whole tokens; and bare query values and the fragment of 10
or more characters, as whole tokens. A whole token has no ASCII letter or
digit directly before or after it. Shorter bare values printed on their
own, such as the1of?v=1, stay. Other URLs in the details lose
userinfo, query and fragment on a best-effort basis
(#642). - Leave no envelope text in the parsed
contentwhen Qwen2.5 wraps a Hermes
tool call in double braces (<tool_call>{{"name": ...}}</tool_call>, with
any number of extra closing braces) without a grammar. Calls are extracted as
before; a double-brace call with other malformed envelope text keeps that
text. This deliberately differs from upstream llama.cpp, which rejects that
output and extracts no call
(#662).