feat(ggml-metal): Add template specialization for mul_mm_id w/ ne20 == 10 #15799

gabe-l-hart · 2025-09-04T15:39:23Z

Description

When testing an internal checkpoint, I hit this abort, so this PR adds the missing template specialization.

…= 10 Branch: GGMLMetalNE20 Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

gabe-l-hart · 2025-09-04T15:40:26Z

@ggerganov Assigning you since I'm not sure the right person and this is small, but feel free to reassign.

…upport * origin/master: (72 commits) metal : Add template specialization for mul_mm_id w/ ne20 == 10 (ggml-org#15799) llama : set n_outputs to 1 to avoid 0 outputs mean-pooling (ggml-org#15791) CANN: Refactor ND to NZ workspace to be per-device (ggml-org#15763) server: add exceed_context_size_error type (ggml-org#15780) Document the new max GPU layers default in help (ggml-org#15771) ggml: add ops for WAN video model (cuda && cpu) (ggml-org#15669) CANN: Fix precision issue on 310I DUO multi-devices (ggml-org#15784) opencl: add hs=40 to FA (ggml-org#15758) CANN: fix acl_rstd allocation size in ggml_cann_rms_norm (ggml-org#15760) vulkan: fix mmv subgroup16 selection (ggml-org#15775) vulkan: don't use std::string in load_shaders, to improve compile time (ggml-org#15724) vulkan : update ggml_vk_instance_validation_ext_available (ggml-org#15666) ggml vulkan: add hardsigmoid and hardswish operations (ggml-org#15762) CUDA: Optimize `rms_norm_f32` kernel and its fused variants, giving 1-6% perf E2E (ggml-org#15715) model-conversion : fix pyright errors (ggml-org#15770) sampling : optimize dist sampler (ggml-org#15704) llama : fix incorrect model type for Gemma 270M (ggml-org#15764) model-conversion : remove hardcoded /bin/bash shebangs [no ci] (ggml-org#15765) CANN: Add RoPE contiguous check for 310I DUP device (ggml-org#15735) ggml-cpu : optimize RVV kernels (ggml-org#15720) ...

…g-model-disabled-agent-prefill * origin/master: (84 commits) CUDA: fastdiv, launch bounds for mmvq + q8_1 quant (ggml-org#15802) tests : add --list-ops and --show-coverage options (ggml-org#15745) gguf: gguf_writer refactor (ggml-org#15691) kv-cache : fix SWA checks + disable cacheless iSWA (ggml-org#15811) model-conversion : add --embeddings flag to modelcard.template [no ci] (ggml-org#15801) chat : fixed crash when Hermes 2 <tool_call> had a newline before it (ggml-org#15639) chat : nemotron thinking & toolcalling support (ggml-org#15676) scripts : add Jinja tester PySide6 simple app (ggml-org#15756) llama : add support for EmbeddingGemma 300m (ggml-org#15798) metal : Add template specialization for mul_mm_id w/ ne20 == 10 (ggml-org#15799) llama : set n_outputs to 1 to avoid 0 outputs mean-pooling (ggml-org#15791) CANN: Refactor ND to NZ workspace to be per-device (ggml-org#15763) server: add exceed_context_size_error type (ggml-org#15780) Document the new max GPU layers default in help (ggml-org#15771) ggml: add ops for WAN video model (cuda && cpu) (ggml-org#15669) CANN: Fix precision issue on 310I DUO multi-devices (ggml-org#15784) opencl: add hs=40 to FA (ggml-org#15758) CANN: fix acl_rstd allocation size in ggml_cann_rms_norm (ggml-org#15760) vulkan: fix mmv subgroup16 selection (ggml-org#15775) vulkan: don't use std::string in load_shaders, to improve compile time (ggml-org#15724) ...

…-org#15799) Branch: GGMLMetalNE20 Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

feat(ggml-metal): Add template specialization for mul_mm_id w/ ne20 =…

2432b4d

…= 10 Branch: GGMLMetalNE20 Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

gabe-l-hart requested a review from ggerganov September 4, 2025 15:40

ggerganov approved these changes Sep 4, 2025

View reviewed changes

ggerganov merged commit 856ed09 into master Sep 4, 2025
48 checks passed

ggerganov deleted the gabe-l-hart/GGMLMetalNE20 branch September 4, 2025 15:53

github-actions bot added ggml changes relating to the ggml tensor library for machine learning Apple Metal https://en.wikipedia.org/wiki/Metal_(API) labels Sep 4, 2025

walidbr pushed a commit to walidbr/llama.cpp that referenced this pull request Sep 7, 2025

metal : Add template specialization for mul_mm_id w/ ne20 == 10 (ggml…

881904a

…-org#15799) Branch: GGMLMetalNE20 Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

feat(ggml-metal): Add template specialization for mul_mm_id w/ ne20 == 10 #15799

feat(ggml-metal): Add template specialization for mul_mm_id w/ ne20 == 10 #15799

Uh oh!

gabe-l-hart commented Sep 4, 2025

Uh oh!

gabe-l-hart commented Sep 4, 2025

Uh oh!

Uh oh!

Uh oh!

feat(ggml-metal): Add template specialization for mul_mm_id w/ ne20 == 10 #15799

feat(ggml-metal): Add template specialization for mul_mm_id w/ ne20 == 10 #15799

Uh oh!

Conversation

gabe-l-hart commented Sep 4, 2025

Description

Uh oh!

gabe-l-hart commented Sep 4, 2025

Uh oh!

Uh oh!

Uh oh!