0.19.918.561 I srv ensure_model: waiting until model name=GLM-Air-4.5-106B-Animus-V12.1-Q5_K_S is fully loaded...
[57663] 0.00.197.681 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
[57663] 0.00.197.690 I device_info:
[57663] 0.00.416.420 I - CUDA0 : NVIDIA GeForce RTX 3090 (24124 MiB, 23859 MiB free)
[57663] 0.00.682.625 I - CUDA1 : NVIDIA GeForce RTX 3090 (24124 MiB, 23859 MiB free)
[57663] 0.00.856.019 I - CUDA2 : NVIDIA GeForce RTX 3090 (24124 MiB, 23859 MiB free)
[57663] 0.01.004.510 I - CUDA3 : NVIDIA GeForce RTX 3090 (24124 MiB, 23859 MiB free)
[57663] 0.01.004.531 I - CPU : AMD EPYC 7F72 24-Core Processor (128664 MiB, 128664 MiB free)
[57663] 0.01.004.639 I system_info: n_threads = 48 (n_threads_batch = 48) / 48 | CUDA : ARCHS = 860 | USE_GRAPHS = 1 | PEER_MAX_BATCH_SIZE = 128 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
[57663] 0.01.004.692 I srv init: running without SSL
[57663] 0.01.004.745 I srv init: using 47 threads for HTTP server
[57663] 0.01.004.923 I srv start: binding port with default address family
[57663] 0.01.006.264 I srv llama_server: loading model
[57663] 0.01.006.273 I srv load_model: loading model '/mnt/oiseauxai1data/quanted_models/gguf/GLM-Air-4.5-106B-Animus-V12.1-Q5_K_S.gguf'
[57663] 0.01.611.709 W load: special_eot_id is not in special_eog_ids - the tokenizer config may be incorrect
[57663] 0.01.611.716 W load: special_eom_id is not in special_eog_ids - the tokenizer config may be incorrect
[57663] 0.01.688.074 W model has unused tensor blk.46.attn_norm.weight (size = 16384 bytes) -- ignoring
[57663] 0.01.688.087 W model has unused tensor blk.46.attn_q.weight (size = 34603008 bytes) -- ignoring
[57663] 0.01.688.091 W model has unused tensor blk.46.attn_k.weight (size = 2883584 bytes) -- ignoring
[57663] 0.01.688.096 W model has unused tensor blk.46.attn_v.weight (size = 2883584 bytes) -- ignoring
[57663] 0.01.688.125 W model has unused tensor blk.46.attn_output.weight (size = 34603008 bytes) -- ignoring
[57663] 0.01.688.137 W model has unused tensor blk.46.post_attention_norm.weight (size = 16384 bytes) -- ignoring
[57663] 0.01.688.141 W model has unused tensor blk.46.ffn_gate_inp.weight (size = 2097152 bytes) -- ignoring
[57663] 0.01.688.146 W model has unused tensor blk.46.exp_probs_b.bias (size = 512 bytes) -- ignoring
[57663] 0.01.688.151 W model has unused tensor blk.46.ffn_gate_exps.weight (size = 507510784 bytes) -- ignoring
[57663] 0.01.688.155 W model has unused tensor blk.46.ffn_down_exps.weight (size = 553648128 bytes) -- ignoring
[57663] 0.01.688.160 W model has unused tensor blk.46.ffn_up_exps.weight (size = 507510784 bytes) -- ignoring
[57663] 0.01.688.164 W model has unused tensor blk.46.ffn_gate_shexp.weight (size = 3964928 bytes) -- ignoring
[57663] 0.01.688.169 W model has unused tensor blk.46.ffn_down_shexp.weight (size = 4325376 bytes) -- ignoring
[57663] 0.01.688.173 W model has unused tensor blk.46.ffn_up_shexp.weight (size = 3964928 bytes) -- ignoring
[57663] 0.01.688.179 W model has unused tensor blk.46.nextn.eh_proj.weight (size = 23068672 bytes) -- ignoring
[57663] 0.01.688.183 W model has unused tensor blk.46.nextn.enorm.weight (size = 16384 bytes) -- ignoring
[57663] 0.01.688.188 W model has unused tensor blk.46.nextn.hnorm.weight (size = 16384 bytes) -- ignoring
[57663] 0.01.688.193 W model has unused tensor blk.46.nextn.embed_tokens.weight (size = 426770432 bytes) -- ignoring
[57663] 0.01.688.198 W model has unused tensor blk.46.nextn.shared_head_head.weight (size = 426770432 bytes) -- ignoring
[57663] 0.01.688.203 W model has unused tensor blk.46.nextn.shared_head_norm.weight (size = 16384 bytes) -- ignoring
[57663] /home/som1tokmynam/LLM/llama.cpp/src/llama-model.cpp:410: GGML_ASSERT(tensor_axis_0 != nullptr) failed
[57663] [New LWP 397069]
[57663] [New LWP 397068]
[57663] [New LWP 397067]
[57663] [New LWP 397066]
[57663] [New LWP 397065]
[57663] [New LWP 397064]
[57663] [New LWP 397063]
[57663] [New LWP 397062]
[57663] [New LWP 397061]
[57663] [New LWP 397060]
[57663] [New LWP 397059]
[57663] [New LWP 397058]
[57663] [New LWP 397057]
[57663] [New LWP 397056]
[57663] [New LWP 397055]
[57663] [New LWP 397054]
[57663] [New LWP 397053]
[57663] [New LWP 397052]
[57663] [New LWP 397050]
[57663] [New LWP 397048]
[57663] [New LWP 397046]
[57663] [New LWP 397045]
[57663] [New LWP 397044]
[57663] [New LWP 397042]
[57663] [New LWP 397040]
[57663] [New LWP 397038]
[57663] [New LWP 397036]
[57663] [New LWP 397034]
[57663] [New LWP 397032]
[57663] [New LWP 397030]
[57663] [New LWP 397028]
[57663] [New LWP 397026]
[57663] [New LWP 397024]
[57663] [New LWP 397023]
[57663] [New LWP 397022]
[57663] [New LWP 397021]
[57663] [New LWP 397020]
[57663] [New LWP 397018]
[57663] [New LWP 397016]
[57663] [New LWP 397014]
[57663] [New LWP 397012]
[57663] [New LWP 397010]
[57663] [New LWP 397008]
[57663] [New LWP 397006]
[57663] [New LWP 397004]
[57663] [New LWP 397002]
[57663] [New LWP 397000]
[57663] [New LWP 396997]
[57663] [New LWP 396927]
[57663] [New LWP 396926]
[57663] [New LWP 396925]
[57663] [New LWP 396924]
[57663] [New LWP 396828]
[57663] [New LWP 396827]
[57663] [New LWP 396820]
[57663] [New LWP 396819]
[57663] [New LWP 396818]
[57663] [New LWP 396817]
[57663]
[57663] This GDB supports auto-downloading debuginfo from the following URLs:
[57663] <https://debuginfod.ubuntu.com>
[57663] Enable debuginfod for this session? (y or [n]) [answered N; input not from terminal]
[57663] Debuginfod has been disabled.
[57663] To make this setting permanent, add 'set debuginfod enabled off' to .gdbinit.
[57663] [Thread debugging using libthread_db enabled]
[57663] Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".
[57663] 0x00007d2dbf5fd813 in __GI___wait4 (pid=397361, stat_loc=0x0, options=0, usage=0x0) at ../sysdeps/unix/sysv/linux/wait4.c:30
[57663] warning: 30 ../sysdeps/unix/sysv/linux/wait4.c: No such file or directory
[57663] #0 0x00007d2dbf5fd813 in __GI___wait4 (pid=397361, stat_loc=0x0, options=0, usage=0x0) at ../sysdeps/unix/sysv/linux/wait4.c:30
[57663] 30 in ../sysdeps/unix/sysv/linux/wait4.c
[57663] #1 0x00007d2dbeb71663 in ggml_print_backtrace () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libggml-base.so.0
[57663] #2 0x00007d2dbeb7180b in ggml_abort () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libggml-base.so.0
[57663] #3 0x00007d2dbeda4a97 in llama_meta_device_get_split_state(ggml_tensor const*, void*)::{lambda(ggml_backend_meta_split_axis, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&)#1}::operator()(ggml_backend_meta_split_axis, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&) const () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libllama.so.0
[57663] #4 0x00007d2dbedb1d85 in llama_meta_device_get_split_state(ggml_tensor const*, void*)::{lambda()#1}::operator()() const () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libllama.so.0
[57663] #5 0x00007d2dbedb3505 in llama_meta_device_get_split_state(ggml_tensor const*, void*) () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libllama.so.0
[57663] #6 0x00007d2dbeb97825 in ggml_backend_meta_get_split_state(ggml_backend_meta_simple_tensor_container&, ggml_tensor const*, bool)::{lambda()#1}::operator()() const () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libggml-base.so.0
[57663] #7 0x00007d2dbeb922ca in ggml_backend_meta_get_split_state(ggml_backend_meta_simple_tensor_container&, ggml_tensor const*, bool) () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libggml-base.so.0
[57663] #8 0x00007d2dbeb99f7f in ggml_backend_meta_buffer_init_tensor_impl(ggml_backend_meta_simple_tensor_container&, ggml_tensor*) () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libggml-base.so.0
[57663] #9 0x00007d2dbeb9b91f in ggml_backend_meta_alloc_ctx_tensors_from_buft () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libggml-base.so.0
[57663] #10 0x00007d2dbedae388 in llama_model_base::load_tensors(llama_model_loader&) () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libllama.so.0
[57663] #11 0x00007d2dbeced9d0 in llama_model_load(gguf_context*, void (*)(ggml_tensor*, void*), void*, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > >&, _IO_FILE*, llama_model_params&) () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libllama.so.0
[57663] #12 0x00007d2dbeceea72 in llama_model_load_from_file_impl(gguf_context*, void (*)(ggml_tensor*, void*), void*, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::vector<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::allocator<std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > > >&, _IO_FILE*, llama_model_params) () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libllama.so.0
[57663] #13 0x00007d2dbeceee0c in llama_model_load_from_file () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libllama.so.0
[57663] #14 0x00007d2dbf285808 in common_init_result::common_init_result(common_params&, bool) () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libllama-common.so.0
[57663] #15 0x00007d2dbf2869f3 in common_init_from_params(common_params&, bool) () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libllama-common.so.0
[57663] #16 0x00007d2dbfb4469c in server_context_impl::load_model(common_params&) () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libllama-server-impl.so
[57663] #17 0x00007d2dbfa95cf6 in llama_server(int, char**) () from /home/som1tokmynam/LLM/llama.cpp/build/bin/libllama-server-impl.so
[57663] #18 0x00007d2dbf5171ca in __libc_start_call_main (main=main@entry=0x5b87d356f270 <main>, argc=argc@entry=41, argv=argv@entry=0x7fffe0f1fd08) at ../sysdeps/nptl/libc_start_call_main.h:58
[57663] warning: 58 ../sysdeps/nptl/libc_start_call_main.h: No such file or directory
[57663] #19 0x00007d2dbf51728b in __libc_start_main_impl (main=0x5b87d356f270 <main>, argc=41, argv=0x7fffe0f1fd08, init=<optimized out>, fini=<optimized out>, rtld_fini=<optimized out>, stack_end=0x7fffe0f1fcf8) at ../csu/libc-start.c:360
[57663] warning: 360 ../csu/libc-start.c: No such file or directory
[57663] #20 0x00005b87d356f2a5 in _start ()
[57663] [Inferior 1 (process 396812) detached]
Name and Version
version: 9553 (9e3b928)
built with GNU 13.3.0 for Linux x86_64
Operating systems
Linux
GGML backends
CUDA
Hardware
Epyc 7F72 + 4x 3090 + 1x3060 (unused for llama.cpp)
Models
Darkhn/GLM-Air-4.5-106B-Animus-V12.1 (finetune of zai-org/GLM-4.5-Air, same issue with base model)
Problem description & steps to reproduce
Problem description:
When loading a GLM-4.5 MoE model (or fine-tune) with split mode tensor
-sm tensor, llama.cpp crashes during model loading with:GGML_ASSERT(tensor_axis_0 != nullptr) failed
The crash originates in llama-model.cpp:410 inside llama_meta_device_get_split_state, triggered by the nextn speculative draft head tensors at blk.46 (the last layer). These tensors are already reported as unused by llama.cpp itself, but the tensor-split allocator still attempts to compute a split axis for them and hits an assertion failure on their non-standard shape.
Workaround:
Strip the nextn layer and patch the model metadata using llama-quantize:
The --override-kv flags are essential — without them, --prune-layers renumbers layers and causes a different loading failure (missing tensor 'blk.45.nextn.eh_proj.weight').
Expected behavior: tensors already marked unused should be skipped entirely by the tensor-split allocator, not cause an assertion failure.
Run command for inference:
CUDA_VISIBLE_DEVICES=0,1,2,3 /home/som1tokmynam/LLM/llama.cpp/build/bin/llama-server
--host 10.0.0.8
--n-gpu-layers 999
--threads -1
--cache-ram 65536
--mlock
--no-mmap
--fit off
--batch-size 1024
--ubatch-size 1024
-dev CUDA0,CUDA1,CUDA2,CUDA3
-ctv bf16
-ctk bf16
-sm tensor
--special
--jinja
--port 5002
--no-warmup
--model /mnt/oiseauxai1data/quanted_models/gguf/GLM-Air-4.5-106B-Animus-V12.1-Q5_K_S.gguf
--tensor-split 24,24,24,24
--ctx-size 82000
First Bad Commit
No response
Relevant log output
Logs