b10342
model : Granite-Switch Architecture (#25107)
- granite-switch: add llama.cpp backend (POC, CPU)
New "granite-switch" architecture: a dense, all-attention Granite-4.1
model with N embedded LoRA adapters selected per-token by control tokens.
- gguf-py schema (arch, KV keys, stacked LoRA tensor names) + writer helpers
- conversion/granite.py: GraniteSwitchModel converter (stacks N adapters +
zero base slot into per-projection A/B tensors; emits switch metadata) - C++ arch registration (llama-arch.{h,cpp}, llama-model.{h,cpp})
- src/models/granite_switch.cpp: load + per-token switched-LoRA graph via
ggml_mul_mat_id over stacked tensors; sticky per-token index + control-token
substitution in llm_graph_input_switch::set_input - llm_graph_input_switch in src/models/models.h
Runs end-to-end on CPU: convert 3b checkpoint (842 tensors, stacked dim 13)
and generate on both base and control-token paths. Sticky switch state is
single-sequence (POC); full multi-sequence machinery is a follow-up.
- granite-switch: add Mac (Metal) build + mid-sequence switch demo script
Self-contained script to build llama.cpp on Apple Silicon (Metal),
convert the composed 3b checkpoint, and run the crisp mid-sequence
adapter-switch demos verified on Vela:
- answerability: <|answerability|> mid-seq -> "unanswerable"
- query_rewrite: <|query_rewrite|> mid-seq -> {"rewritten_question": ...}
Each demo runs the same prompt twice, differing only by a control token
placed before the assistant turn, so the per-token switch is visible.
- granite-switch mac demo: add -no-cnv so each run is one-shot
The composed model ships a chat template, so llama-completion auto-enables
interactive conversation mode and halts at a > prompt after generating,
stalling the script. -no-cnv disables conversation mode: generate once from
the raw prompt and exit (also prints special tokens, making the switch visible).
- granite-switch: replace global sticky index with in-graph router attention
The POC computed the per-token adapter index on the CPU and carried it
across ubatches in ONE global mutable int32_t poc_sticky_index, reset
only when a ubatch contained sequence position 0. That global had two
problems:
- Concurrency: with multiple sequences in a batch it was last-writer-
wins — one sequence's adapter leaked into the others. - Multi-turn: an interactive
ollama runchat continues one KV cache,
so turn 2 never saw position 0 and the index never reset — the
adapter stayed stuck on across turns.
Port the vLLM/HF backend mechanism faithfully: a single-head causal
"router" attention recovers the adapter index in-graph. Per token, only
dim 0 carries signal — Q[0]=1, K[0]=+gain for a control token / -gain
otherwise, V[0]=adapter slot / 0 — and the causal softmax over the single
visible control token recovers that adapter's slot (readback =
clamp(round(V[0]), 0, n_adapters)). gain=15 matches config.py and is
F16-safe (no F32 cache).
The router's K/V live in the model KV cache at an extra layer
R == hparams.router_layer (== n_layer). We bump n_layer_all to n_real+1
so the cache allocator gives the router its own per-sequence slot, and
set n_layer_nextn=1 so n_layer() stays n_real — the decoder loop and
tensor loading are untouched and never reference layer R. The router K is
exempted from the k-shift RoPE loop (its dim-0 value is a literal
magnitude, not a rotation).
Because the selection now lives in the per-sequence KV cache, CONCURRENT
requests are isolated for free (problem 1 fixed; verified by
scratch/concurrent_switch_test.cpp). set_input becomes stateless pure
per-token maps; the global is gone.
Single-switch contract / known limitation, identical to vLLM & HF: the
gain is flat (no recency), so within one sequence there is no mechanism to
revert to base mid-sequence — once an adapter fires it stays on until that
sequence ends (problem 2 is therefore NOT fixed by a faithful copy; vLLM/HF
avoid it only because each served request is a fresh sequence). A client
continuing one KV cache across turns must start a fresh sequence per turn,
or opt into a recency-biased router (a deliberate divergence, not done
here). Documented in granite_switch.cpp and asserted by
scratch/multiturn_leak_test.cpp.
Verified (CPU): both demos unchanged (answerability -> "unanswerable",
query_rewrite -> rewritten query); concurrent two-sequence isolation
passes; multi-turn carry-over matches the vLLM/HF contract.
- granite-switch: drop scratch tests and mac demo for upstream PR
Remove the local-only development artifacts that should not ship in the
upstream PR:
- granite-switch-mac-demo.sh (local Metal build + demo driver)
- scratch/concurrent_switch_test.cpp
- scratch/multiturn_leak_test.cpp
Also drop the now-dangling reference to the scratch tests from the
granite_switch.cpp header comment. Leaves only the core architecture
support (conversion, gguf constants, llama-arch/model/kv-cache, and the
granite_switch graph).
-
granite-switch: trim comments to match native llama.cpp style
-
granite-switch: trim conversion comments to match native style
-
granite-switch: drop unused adapter_ranks metadata
-
granite-switch: rename arch to graniteswitch and drop obid alias
-
granite-switch: fix non-ASCII comments and document router gain assumption
-
granite-switch: drop section comments from constants.py to match native style
-
granite-switch: add functional tensor block comments matching Granite4 Vision style
-
granite-switch: clarify n_expert_used comment
State the actual constraint: mul_mat_id needs n_expert_used == 1, and
since the GGUF carries expert_count = 0 the generic loader's
n_expert == 0 => n_expert_used == 0 assertion has already passed by the
time load_arch_hparams runs, so it is forced to 1 here.
- granite-switch: note n_layer_nextn reuse has no MTP
The router carving reuses n_layer_nextn, normally the MTP/next-token
count. Clarify in the comment that it is borrowed here purely as the
trailing-layers lever and that there is no MTP head, to spare readers
the double-take.
-
granite-switch: rename source file and apply review nits
-
granite-switch: don't force LoRA tensors to F16, follow --outtype instead
-
granite-switch: drop redundant _permute_qk wrapper, call LlamaModel.permute directly
-
granite-switch: read router gain from GGUF (control_token_gain) instead of hardcoding 15.0
-
granite-switch: derive n_slots()
-
granite-switch: move llm_graph_input_switch into granite-switch.cpp
-
granite-switch: cut AI-style narration comments
-
granite-switch: collapse multi-line comments
-
granite-switch: rename control_token_* maps to adapter_token_*
-
granite-switch: cut noise comments
-
granite-switch: rename embedded LoRA tensors to .lora_a/lora_b
-
granite-switch: GGML_ASSERT token input to avoid UB on embeddings
-
granite-switch: TODO for raw embedding input support
-
granite-switch: collapse LoRA tensor constants to .lora_a/.lora_b suffix
-
granite-switch: drop n_expert_used hack, guard mul_mat_id buft probe
-
granite-switch: stop forcing dense expert counts, read from config
-
granite-switch: renamed control_token_gain metadata key to router_gain
-
granite-switch: trim header comments to match native style
-
granite-switch: collapse LoRA tensors to base name + suffix
-
granite-switch: inline suffix checks in tensor op resolution
-
granite-switch: drop switch-lora struct comment
-
granite-switch: guard router layer index and inline n_slots
-
granite-switch: group adapter metadata under {arch}.adapters.* namespace
-
granite-switch: add hparams.has_rope(il) for KV-shift rope skipping
-
granite-switch: skip arch in test-llama-archs (adapter fixture missing, TODO)
-
granite-switch: Keys.Adapters namespace + simplify n_slots
-
granite-switch: validate substitute token ids against n_vocab
-
granite-switch: bound adapter count and lora rank from GGUF
-
granite-switch: reject MTP context type when router_layer is set
-
granite-switch: throw on bad adapter metadata instead of GGML_ASSERT
-
granite-switch: use ASCII +/- in router K signal comment
-
granite-switch: document n_layer_nextn repurpose and its leak points
-
granite-switch: gate lora_a/lora_b op mapping on router_layer
-
granite-switch: label all three preview model sizes
Website:
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
- iOS XCFramework
Linux:
- Ubuntu x64 (CPU)
- Ubuntu arm64 (CPU)
- Ubuntu s390x (CPU)
- Ubuntu x64 (Vulkan)
- Ubuntu arm64 (Vulkan)
- Ubuntu x64 (ROCm 7.2)
- Ubuntu x64 (OpenVINO)
- Ubuntu x64 (SYCL FP32)
- Ubuntu x64 (SYCL FP16)
Android:
Windows:
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows arm64 (OpenCL Adreno)
- Windows x64 (CUDA 12) - CUDA 12.4 DLLs
- Windows x64 (CUDA 13) - CUDA 13.3 DLLs
- Windows x64 (Vulkan)
- Windows x64 (OpenVINO)
- Windows x64 (SYCL)
- Windows x64 (HIP)
openEuler:
- DISABLED
- openEuler x86 (310p)
- openEuler x86 (910b, ACL Graph)
- openEuler aarch64 (310p)
- openEuler aarch64 (910b, ACL Graph)
UI: