curl 'http://127.0.0.1:11011/v1/chat/completions' \
-m 1200 \
--data-raw $'{"messages":[{"role":"user","content":[{"type":"text","text":"\\n\\n--- File: Pasted ---\\nwrite a Lua program that manage the command line options shown bellow it\'s important to preserve the order when showing then and include all of then (don\'t tell me to fill the rest).\\n```\\n----- common params -----\\n\\n-h, --help, --usage print usage and exit\\n--version show version and build info\\n-cl, --cache-list show list of models in cache\\n--completion-bash print source-able bash completion script for llama.cpp\\n--verbose-prompt print a verbose prompt before generation (default: false)\\n-t, --threads N number of CPU threads to use during generation (default: -1)\\n (env: LLAMA_ARG_THREADS)\\n-tb, --threads-batch N number of threads to use during batch and prompt processing (default:\\n same as --threads)\\n-C, --cpu-mask M CPU affinity mask: arbitrarily long hex. Complements cpu-range\\n (default: \\"\\")\\n-Cr, --cpu-range lo-hi range of CPUs for affinity. Complements --cpu-mask\\n--cpu-strict <0|1> use strict CPU placement (default: 0)\\n--prio N set process/thread priority : low(-1), normal(0), medium(1), high(2),\\n realtime(3) (default: 0)\\n--poll <0...100> use polling level to wait for work (0 - no polling, default: 50)\\n-Cb, --cpu-mask-batch M CPU affinity mask: arbitrarily long hex. Complements cpu-range-batch\\n (default: same as --cpu-mask)\\n-Crb, --cpu-range-batch lo-hi ranges of CPUs for affinity. Complements --cpu-mask-batch\\n--cpu-strict-batch <0|1> use strict CPU placement (default: same as --cpu-strict)\\n--prio-batch N set process/thread priority : 0-normal, 1-medium, 2-high, 3-realtime\\n (default: 0)\\n--poll-batch <0|1> use polling to wait for work (default: same as --poll)\\n-c, --ctx-size N size of the prompt context (default: 0, 0 = loaded from model)\\n (env: LLAMA_ARG_CTX_SIZE)\\n-n, --predict, --n-predict N number of tokens to predict (default: -1, -1 = infinity)\\n (env: LLAMA_ARG_N_PREDICT)\\n-b, --batch-size N logical maximum batch size (default: 2048)\\n (env: LLAMA_ARG_BATCH)\\n-ub, --ubatch-size N physical maximum batch size (default: 512)\\n (env: LLAMA_ARG_UBATCH)\\n--keep N number of tokens to keep from the initial prompt (default: 0, -1 =\\n all)\\n--swa-full use full-size SWA cache (default: false)\\n [(more\\n info)](https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)\\n (env: LLAMA_ARG_SWA_FULL)\\n--kv-unified, -kvu use single unified KV buffer for the KV cache of all sequences\\n (default: false)\\n [(more info)](https://github.com/ggml-org/llama.cpp/pull/14363)\\n (env: LLAMA_ARG_KV_UNIFIED)\\n-fa, --flash-attn [on|off|auto] set Flash Attention use (\'on\', \'off\', or \'auto\', default: \'auto\')\\n (env: LLAMA_ARG_FLASH_ATTN)\\n-p, --prompt PROMPT prompt to start generation with; for system message, use -sys\\n--perf, --no-perf whether to enable internal libllama performance timings (default:\\n false)\\n (env: LLAMA_ARG_PERF)\\n-f, --file FNAME a file containing the prompt (default: none)\\n-bf, --binary-file FNAME binary file containing the prompt (default: none)\\n-e, --escape, --no-escape whether to process escapes sequences (\\\\n, \\\\r, \\\\t, \\\\\', \\\\\\", \\\\\\\\)\\n (default: true)\\n--rope-scaling {none,linear,yarn} RoPE frequency scaling method, defaults to linear unless specified by\\n the model\\n (env: LLAMA_ARG_ROPE_SCALING_TYPE)\\n--rope-scale N RoPE context scaling factor, expands context by a factor of N\\n (env: LLAMA_ARG_ROPE_SCALE)\\n--rope-freq-base N RoPE base frequency, used by NTK-aware scaling (default: loaded from\\n model)\\n (env: LLAMA_ARG_ROPE_FREQ_BASE)\\n--rope-freq-scale N RoPE frequency scaling factor, expands context by a factor of 1/N\\n (env: LLAMA_ARG_ROPE_FREQ_SCALE)\\n--yarn-orig-ctx N YaRN: original context size of model (default: 0 = model training\\n context size)\\n (env: LLAMA_ARG_YARN_ORIG_CTX)\\n--yarn-ext-factor N YaRN: extrapolation mix factor (default: -1.0, 0.0 = full\\n interpolation)\\n (env: LLAMA_ARG_YARN_EXT_FACTOR)\\n--yarn-attn-factor N YaRN: scale sqrt(t) or attention magnitude (default: -1.0)\\n (env: LLAMA_ARG_YARN_ATTN_FACTOR)\\n--yarn-beta-slow N YaRN: high correction dim or alpha (default: -1.0)\\n (env: LLAMA_ARG_YARN_BETA_SLOW)\\n--yarn-beta-fast N YaRN: low correction dim or beta (default: -1.0)\\n (env: LLAMA_ARG_YARN_BETA_FAST)\\n-kvo, --kv-offload, -nkvo, --no-kv-offload\\n whether to enable KV cache offloading (default: enabled)\\n (env: LLAMA_ARG_KV_OFFLOAD)\\n--repack, -nr, --no-repack whether to enable weight repacking (default: enabled)\\n (env: LLAMA_ARG_REPACK)\\n--no-host bypass host buffer allowing extra buffers to be used\\n (env: LLAMA_ARG_NO_HOST)\\n-ctk, --cache-type-k TYPE KV cache data type for K\\n allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1\\n (default: f16)\\n (env: LLAMA_ARG_CACHE_TYPE_K)\\n-ctv, --cache-type-v TYPE KV cache data type for V\\n allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1\\n (default: f16)\\n (env: LLAMA_ARG_CACHE_TYPE_V)\\n-dt, --defrag-thold N KV cache defragmentation threshold (DEPRECATED)\\n (env: LLAMA_ARG_DEFRAG_THOLD)\\n-np, --parallel N number of parallel sequences to decode (default: 1)\\n (env: LLAMA_ARG_N_PARALLEL)\\n--mlock force system to keep model in RAM rather than swapping or compressing\\n (env: LLAMA_ARG_MLOCK)\\n--mmap, --no-mmap whether to memory-map model (if disabled, slower load but may reduce\\n pageouts if not using mlock) (default: enabled)\\n (env: LLAMA_ARG_MMAP)\\n--numa TYPE attempt optimizations that help on some NUMA systems\\n - distribute: spread execution evenly over all nodes\\n - isolate: only spawn threads on CPUs on the node that execution\\n started on\\n - numactl: use the CPU map provided by numactl\\n if run without this previously, it is recommended to drop the system\\n page cache before using this\\n see https://github.com/ggml-org/llama.cpp/issues/1437\\n (env: LLAMA_ARG_NUMA)\\n-dev, --device <dev1,dev2,..> comma-separated list of devices to use for offloading (none = don\'t\\n offload)\\n use --list-devices to see a list of available devices\\n (env: LLAMA_ARG_DEVICE)\\n--list-devices print list of available devices and exit\\n--override-tensor, -ot <tensor name pattern>=<buffer type>,...\\n override tensor buffer type\\n--cpu-moe, -cmoe keep all Mixture of Experts (MoE) weights in the CPU\\n (env: LLAMA_ARG_CPU_MOE)\\n--n-cpu-moe, -ncmoe N keep the Mixture of Experts (MoE) weights of the first N layers in the\\n CPU\\n (env: LLAMA_ARG_N_CPU_MOE)\\n-ngl, --gpu-layers, --n-gpu-layers N max. number of layers to store in VRAM (default: -1)\\n (env: LLAMA_ARG_N_GPU_LAYERS)\\n-sm, --split-mode {none,layer,row} how to split the model across multiple GPUs, one of:\\n - none: use one GPU only\\n - layer (default): split layers and KV across GPUs\\n - row: split rows across GPUs\\n (env: LLAMA_ARG_SPLIT_MODE)\\n-ts, --tensor-split N0,N1,N2,... fraction of the model to offload to each GPU, comma-separated list of\\n proportions, e.g. 3,1\\n (env: LLAMA_ARG_TENSOR_SPLIT)\\n-mg, --main-gpu INDEX the GPU to use for the model (with split-mode = none), or for\\n intermediate results and KV (with split-mode = row) (default: 0)\\n (env: LLAMA_ARG_MAIN_GPU)\\n-fit, --fit [on|off] whether to adjust unset arguments to fit in device memory (\'on\' or\\n \'off\', default: \'on\')\\n (env: LLAMA_ARG_FIT)\\n-fitt, --fit-target MiB target margin per device for --fit option, default: 1024\\n (env: LLAMA_ARG_FIT_TARGET)\\n-fitc, --fit-ctx N minimum ctx size that can be set by --fit option, default: 4096\\n (env: LLAMA_ARG_FIT_CTX)\\n--check-tensors check model tensor data for invalid values (default: false)\\n--override-kv KEY=TYPE:VALUE advanced option to override model metadata by key. may be specified\\n multiple times.\\n types: int, float, bool, str. example: --override-kv\\n tokenizer.ggml.add_bos_token=bool:false\\n--op-offload, --no-op-offload whether to offload host tensor operations to device (default: true)\\n--lora FNAME path to LoRA adapter (can be repeated to use multiple adapters)\\n--lora-scaled FNAME SCALE path to LoRA adapter with user defined scaling (can be repeated to use\\n multiple adapters)\\n--control-vector FNAME add a control vector\\n note: this argument can be repeated to add multiple control vectors\\n--control-vector-scaled FNAME SCALE add a control vector with user defined scaling SCALE\\n note: this argument can be repeated to add multiple scaled control\\n vectors\\n--control-vector-layer-range START END\\n layer range to apply the control vector(s) to, start and end inclusive\\n-m, --model FNAME model path to load\\n (env: LLAMA_ARG_MODEL)\\n-mu, --model-url MODEL_URL model download url (default: unused)\\n (env: LLAMA_ARG_MODEL_URL)\\n-dr, --docker-repo [<repo>/]<model>[:quant]\\n Docker Hub model repository. repo is optional, default to ai/. quant\\n is optional, default to :latest.\\n example: gemma3\\n (default: unused)\\n (env: LLAMA_ARG_DOCKER_REPO)\\n-hf, -hfr, --hf-repo <user>/<model>[:quant]\\n Hugging Face model repository; quant is optional, case-insensitive,\\n default to Q4_K_M, or falls back to the first file in the repo if\\n Q4_K_M doesn\'t exist.\\n mmproj is also downloaded automatically if available. to disable, add\\n --no-mmproj\\n example: unsloth/phi-4-GGUF:q4_k_m\\n (default: unused)\\n (env: LLAMA_ARG_HF_REPO)\\n-hfd, -hfrd, --hf-repo-draft <user>/<model>[:quant]\\n Same as --hf-repo, but for the draft model (default: unused)\\n (env: LLAMA_ARG_HFD_REPO)\\n-hff, --hf-file FILE Hugging Face model file. If specified, it will override the quant in\\n --hf-repo (default: unused)\\n (env: LLAMA_ARG_HF_FILE)\\n-hfv, -hfrv, --hf-repo-v <user>/<model>[:quant]\\n Hugging Face model repository for the vocoder model (default: unused)\\n (env: LLAMA_ARG_HF_REPO_V)\\n-hffv, --hf-file-v FILE Hugging Face model file for the vocoder model (default: unused)\\n (env: LLAMA_ARG_HF_FILE_V)\\n-hft, --hf-token TOKEN Hugging Face access token (default: value from HF_TOKEN environment\\n variable)\\n (env: HF_TOKEN)\\n--log-disable Log disable\\n--log-file FNAME Log to file\\n (env: LLAMA_LOG_FILE)\\n--log-colors [on|off|auto] Set colored logging (\'on\', \'off\', or \'auto\', default: \'auto\')\\n \'auto\' enables colors when output is to a terminal\\n (env: LLAMA_LOG_COLORS)\\n-v, --verbose, --log-verbose Set verbosity level to infinity (i.e. log all messages, useful for\\n debugging)\\n--offline Offline mode: forces use of cache, prevents network access\\n (env: LLAMA_OFFLINE)\\n-lv, --verbosity, --log-verbosity N Set the verbosity threshold. Messages with a higher verbosity will be\\n ignored. Values:\\n - 0: generic output\\n - 1: error\\n - 2: warning\\n - 3: info\\n - 4: debug\\n (default: 1)\\n\\n (env: LLAMA_LOG_VERBOSITY)\\n--log-prefix Enable prefix in log messages\\n (env: LLAMA_LOG_PREFIX)\\n--log-timestamps Enable timestamps in log messages\\n (env: LLAMA_LOG_TIMESTAMPS)\\n-ctkd, --cache-type-k-draft TYPE KV cache data type for K for the draft model\\n allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1\\n (default: f16)\\n (env: LLAMA_ARG_CACHE_TYPE_K_DRAFT)\\n-ctvd, --cache-type-v-draft TYPE KV cache data type for V for the draft model\\n allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1\\n (default: f16)\\n (env: LLAMA_ARG_CACHE_TYPE_V_DRAFT)\\n\\n\\n----- sampling params -----\\n\\n--samplers SAMPLERS samplers that will be used for generation in the order, separated by\\n \';\'\\n (default:\\n penalties;dry;top_n_sigma;top_k;typ_p;top_p;min_p;xtc;temperature)\\n-s, --seed SEED RNG seed (default: -1, use random seed for -1)\\n--sampling-seq, --sampler-seq SEQUENCE\\n simplified sequence for samplers that will be used (default:\\n edskypmxt)\\n--ignore-eos ignore end of stream token and continue generating (implies\\n --logit-bias EOS-inf)\\n--temp N temperature (default: 0.8)\\n--top-k N top-k sampling (default: 40, 0 = disabled)\\n (env: LLAMA_ARG_TOP_K)\\n--top-p N top-p sampling (default: 0.9, 1.0 = disabled)\\n--min-p N min-p sampling (default: 0.1, 0.0 = disabled)\\n--top-nsigma N top-n-sigma sampling (default: -1.0, -1.0 = disabled)\\n--xtc-probability N xtc probability (default: 0.0, 0.0 = disabled)\\n--xtc-threshold N xtc threshold (default: 0.1, 1.0 = disabled)\\n--typical N locally typical sampling, parameter p (default: 1.0, 1.0 = disabled)\\n--repeat-last-n N last n tokens to consider for penalize (default: 64, 0 = disabled, -1\\n = ctx_size)\\n--repeat-penalty N penalize repeat sequence of tokens (default: 1.0, 1.0 = disabled)\\n--presence-penalty N repeat alpha presence penalty (default: 0.0, 0.0 = disabled)\\n--frequency-penalty N repeat alpha frequency penalty (default: 0.0, 0.0 = disabled)\\n--dry-multiplier N set DRY sampling multiplier (default: 0.0, 0.0 = disabled)\\n--dry-base N set DRY sampling base value (default: 1.75)\\n--dry-allowed-length N set allowed length for DRY sampling (default: 2)\\n--dry-penalty-last-n N set DRY penalty for the last n tokens (default: -1, 0 = disable, -1 =\\n context size)\\n--dry-sequence-breaker STRING add sequence breaker for DRY sampling, clearing out default breakers\\n (\'\\\\n\', \':\', \'\\"\', \'*\') in the process; use \\"none\\" to not use any\\n sequence breakers\\n--dynatemp-range N dynamic temperature range (default: 0.0, 0.0 = disabled)\\n--dynatemp-exp N dynamic temperature exponent (default: 1.0)\\n--mirostat N use Mirostat sampling.\\n Top K, Nucleus and Locally Typical samplers are ignored if used.\\n (default: 0, 0 = disabled, 1 = Mirostat, 2 = Mirostat 2.0)\\n--mirostat-lr N Mirostat learning rate, parameter eta (default: 0.1)\\n--mirostat-ent N Mirostat target entropy, parameter tau (default: 5.0)\\n-l, --logit-bias TOKEN_ID(+/-)BIAS modifies the likelihood of token appearing in the completion,\\n i.e. `--logit-bias 15043+1` to increase likelihood of token \' Hello\',\\n or `--logit-bias 15043-1` to decrease likelihood of token \' Hello\'\\n--grammar GRAMMAR BNF-like grammar to constrain generations (see samples in grammars/\\n dir) (default: \'\')\\n--grammar-file FNAME file to read grammar from\\n-j, --json-schema SCHEMA JSON schema to constrain generations (https://json-schema.org/), e.g.\\n `{}` for any JSON object\\n For schemas w/ external $refs, use --grammar +\\n example/json_schema_to_grammar.py instead\\n-jf, --json-schema-file FILE File containing a JSON schema to constrain generations\\n (https://json-schema.org/), e.g. `{}` for any JSON object\\n For schemas w/ external $refs, use --grammar +\\n example/json_schema_to_grammar.py instead\\n\\n\\n----- example-specific params -----\\n\\n--display-prompt, --no-display-prompt whether to print prompt at generation (default: true)\\n-co, --color [on|off|auto] Colorize output to distinguish prompt and user input from generations\\n (\'on\', \'off\', or \'auto\', default: \'auto\')\\n \'auto\' enables colors when output is to a terminal\\n--ctx-checkpoints, --swa-checkpoints N\\n max number of context checkpoints to create per slot (default: 8)\\n [(more info)](https://github.com/ggml-org/llama.cpp/pull/15293)\\n (env: LLAMA_ARG_CTX_CHECKPOINTS)\\n--cache-ram, -cram N set the maximum cache size in MiB (default: 8192, -1 - no limit, 0 -\\n disable)\\n [(more info)](https://github.com/ggml-org/llama.cpp/pull/16391)\\n (env: LLAMA_ARG_CACHE_RAM)\\n--context-shift, --no-context-shift whether to use context shift on infinite text generation (default:\\n disabled)\\n (env: LLAMA_ARG_CONTEXT_SHIFT)\\n-sys, --system-prompt PROMPT system prompt to use with model (if applicable, depending on chat\\n template)\\n--show-timings, --no-show-timings whether to show timing information after each response (default: true)\\n (env: LLAMA_ARG_SHOW_TIMINGS)\\n-sysf, --system-prompt-file FNAME a file containing the system prompt (default: none)\\n-r, --reverse-prompt PROMPT halt generation at PROMPT, return control in interactive mode\\n-sp, --special special tokens output enabled (default: false)\\n-cnv, --conversation, -no-cnv, --no-conversation\\n whether to run in conversation mode:\\n - does not print special tokens and suffix/prefix\\n - interactive mode is also enabled\\n (default: auto enabled if chat template is available)\\n-st, --single-turn run conversation for a single turn only, then exit when done\\n will not be interactive if first turn is predefined with --prompt\\n (default: false)\\n-mli, --multiline-input allows you to write or paste multiple lines without ending each in \'\\\\\'\\n--warmup, --no-warmup whether to perform warmup with an empty run (default: enabled)\\n-mm, --mmproj FILE path to a multimodal projector file. see tools/mtmd/README.md\\n note: if -hf is used, this argument can be omitted\\n (env: LLAMA_ARG_MMPROJ)\\n-mmu, --mmproj-url URL URL to a multimodal projector file. see tools/mtmd/README.md\\n (env: LLAMA_ARG_MMPROJ_URL)\\n--mmproj-auto, --no-mmproj, --no-mmproj-auto\\n whether to use multimodal projector file (if available), useful when\\n using -hf (default: enabled)\\n (env: LLAMA_ARG_MMPROJ_AUTO)\\n--mmproj-offload, --no-mmproj-offload whether to enable GPU offloading for multimodal projector (default:\\n enabled)\\n (env: LLAMA_ARG_MMPROJ_OFFLOAD)\\n--image, --audio FILE path to an image or audio file. use with multimodal models, can be\\n repeated if you have multiple files\\n--image-min-tokens N minimum number of tokens each image can take, only used by vision\\n models with dynamic resolution (default: read from model)\\n (env: LLAMA_ARG_IMAGE_MIN_TOKENS)\\n--image-max-tokens N maximum number of tokens each image can take, only used by vision\\n models with dynamic resolution (default: read from model)\\n (env: LLAMA_ARG_IMAGE_MAX_TOKENS)\\n--override-tensor-draft, -otd <tensor name pattern>=<buffer type>,...\\n override tensor buffer type for draft model\\n--cpu-moe-draft, -cmoed keep all Mixture of Experts (MoE) weights in the CPU for the draft\\n model\\n (env: LLAMA_ARG_CPU_MOE_DRAFT)\\n--n-cpu-moe-draft, -ncmoed N keep the Mixture of Experts (MoE) weights of the first N layers in the\\n CPU for the draft model\\n (env: LLAMA_ARG_N_CPU_MOE_DRAFT)\\n--chat-template-kwargs STRING sets additional params for the json template parser\\n (env: LLAMA_CHAT_TEMPLATE_KWARGS)\\n--jinja, --no-jinja whether to use jinja template engine for chat (default: enabled)\\n (env: LLAMA_ARG_JINJA)\\n--reasoning-format FORMAT controls whether thought tags are allowed and/or extracted from the\\n response, and in which format they\'re returned; one of:\\n - none: leaves thoughts unparsed in `message.content`\\n - deepseek: puts thoughts in `message.reasoning_content`\\n - deepseek-legacy: keeps `<think>` tags in `message.content` while\\n also populating `message.reasoning_content`\\n (default: auto)\\n (env: LLAMA_ARG_THINK)\\n--reasoning-budget N controls the amount of thinking allowed; currently only one of: -1 for\\n unrestricted thinking budget, or 0 to disable thinking (default: -1)\\n (env: LLAMA_ARG_THINK_BUDGET)\\n--chat-template JINJA_TEMPLATE set custom jinja chat template (default: template taken from model\'s\\n metadata)\\n if suffix/prefix are specified, template will be disabled\\n only commonly used templates are accepted (unless --jinja is set\\n before this flag):\\n list of built-in templates:\\n bailing, bailing-think, bailing2, chatglm3, chatglm4, chatml,\\n command-r, deepseek, deepseek2, deepseek3, exaone3, exaone4, falcon3,\\n gemma, gigachat, glmedge, gpt-oss, granite, grok-2, hunyuan-dense,\\n hunyuan-moe, kimi-k2, llama2, llama2-sys, llama2-sys-bos,\\n llama2-sys-strip, llama3, llama4, megrez, minicpm, mistral-v1,\\n mistral-v3, mistral-v3-tekken, mistral-v7, mistral-v7-tekken, monarch,\\n openchat, orion, pangu-embedded, phi3, phi4, rwkv-world, seed_oss,\\n smolvlm, vicuna, vicuna-orca, yandex, zephyr\\n (env: LLAMA_ARG_CHAT_TEMPLATE)\\n--chat-template-file JINJA_TEMPLATE_FILE\\n set custom jinja chat template file (default: template taken from\\n model\'s metadata)\\n if suffix/prefix are specified, template will be disabled\\n only commonly used templates are accepted (unless --jinja is set\\n before this flag):\\n list of built-in templates:\\n bailing, bailing-think, bailing2, chatglm3, chatglm4, chatml,\\n command-r, deepseek, deepseek2, deepseek3, exaone3, exaone4, falcon3,\\n gemma, gigachat, glmedge, gpt-oss, granite, grok-2, hunyuan-dense,\\n hunyuan-moe, kimi-k2, llama2, llama2-sys, llama2-sys-bos,\\n llama2-sys-strip, llama3, llama4, megrez, minicpm, mistral-v1,\\n mistral-v3, mistral-v3-tekken, mistral-v7, mistral-v7-tekken, monarch,\\n openchat, orion, pangu-embedded, phi3, phi4, rwkv-world, seed_oss,\\n smolvlm, vicuna, vicuna-orca, yandex, zephyr\\n (env: LLAMA_ARG_CHAT_TEMPLATE_FILE)\\n--simple-io use basic IO for better compatibility in subprocesses and limited\\n consoles\\n--draft-max, --draft, --draft-n N number of tokens to draft for speculative decoding (default: 16)\\n (env: LLAMA_ARG_DRAFT_MAX)\\n--draft-min, --draft-n-min N minimum number of draft tokens to use for speculative decoding\\n (default: 0)\\n (env: LLAMA_ARG_DRAFT_MIN)\\n--draft-p-min P minimum speculative decoding probability (greedy) (default: 0.8)\\n (env: LLAMA_ARG_DRAFT_P_MIN)\\n-cd, --ctx-size-draft N size of the prompt context for the draft model (default: 0, 0 = loaded\\n from model)\\n (env: LLAMA_ARG_CTX_SIZE_DRAFT)\\n-devd, --device-draft <dev1,dev2,..> comma-separated list of devices to use for offloading the draft model\\n (none = don\'t offload)\\n use --list-devices to see a list of available devices\\n-ngld, --gpu-layers-draft, --n-gpu-layers-draft N\\n number of layers to store in VRAM for the draft model\\n (env: LLAMA_ARG_N_GPU_LAYERS_DRAFT)\\n-md, --model-draft FNAME draft model for speculative decoding (default: unused)\\n (env: LLAMA_ARG_MODEL_DRAFT)\\n--spec-replace TARGET DRAFT translate the string in TARGET into DRAFT if the draft model and main\\n model are not compatible\\n--gpt-oss-20b-default use gpt-oss-20b (note: can download weights from the internet)\\n--gpt-oss-120b-default use gpt-oss-120b (note: can download weights from the internet)\\n--vision-gemma-4b-default use Gemma 3 4B QAT (note: can download weights from the internet)\\n--vision-gemma-12b-default use Gemma 3 12B QAT (note: can download weights from the internet)\\n```\\n"}]}],"stream":false,"model":"ggml-org/gpt-oss-20b-GGUF","reasoning_format":"auto","temperature":0.3,"max_tokens":-1,"dynatemp_range":0,"dynatemp_exponent":1,"top_k":30,"top_p":0.95,"min_p":0.05,"xtc_probability":0,"xtc_threshold":0.1,"typ_p":1,"repeat_last_n":64,"repeat_penalty":1.1,"presence_penalty":0,"frequency_penalty":0,"dry_multiplier":0,"dry_base":1.75,"dry_allowed_length":2,"dry_penalty_last_n":8192,"samplers":["penalties","dry","top_n_sigma","top_k","typ_p","top_p","min_p","xtc","temperature"],"timings_per_token":true}'
Name and Version
llama-server --version
version: 7412 (0f4f35e)
built with AppleClang 17.0.0.17000013 for Darwin arm64
Operating systems
Mac
Which llama.cpp modules do you know to be affected?
llama-server
Command line
Problem description & steps to reproduce
When trying to call with curl with
"stream": false(see command bellow) we get this output (also on llama-cc logssrv operator(): http client error: Failed to read connection):Curl call:
Here #17636 (comment) also there is the same message.
First Bad Commit
No response
Relevant log output