Skip to content

Misc. bug: llama-server fail with 'Failed to read connection' when 'stream: false' #18063

Description

@mingodad

Name and Version

llama-server --version
version: 7412 (0f4f35e)
built with AppleClang 17.0.0.17000013 for Darwin arm64

Operating systems

Mac

Which llama.cpp modules do you know to be affected?

llama-server

Command line

llama-server --host 0.0.0.0 --port 11011 --offline --models-dir /Users/mingo/dev/llms --models-preset /Users/mingo/dev/llms/server-preset.ini --models-max 1
llama-server --chat-template chatml --host 127.0.0.1 --jinja --no-mmap -ncmoe 12 --offline --port 49247 --alias ggml-org/gpt-oss-20b-GGUF --ctx-size 32768 --flash-attn on --hf-repo ggml-org/gpt-oss-20b-GGUF

Problem description & steps to reproduce

When trying to call with curl with "stream": false (see command bellow) we get this output (also on llama-cc logs srv operator(): http client error: Failed to read connection):

proxy error: Failed to read connection

Curl call:

curl 'http://127.0.0.1:11011/v1/chat/completions' \
  -m 1200 \
  --data-raw $'{"messages":[{"role":"user","content":[{"type":"text","text":"\\n\\n--- File: Pasted ---\\nwrite a Lua program that manage the command line options shown bellow it\'s important to preserve the order when showing then and include all of then (don\'t tell me to fill the rest).\\n```\\n----- common params -----\\n\\n-h,    --help, --usage                  print usage and exit\\n--version                               show version and build info\\n-cl,   --cache-list                     show list of models in cache\\n--completion-bash                       print source-able bash completion script for llama.cpp\\n--verbose-prompt                        print a verbose prompt before generation (default: false)\\n-t,    --threads N                      number of CPU threads to use during generation (default: -1)\\n                                        (env: LLAMA_ARG_THREADS)\\n-tb,   --threads-batch N                number of threads to use during batch and prompt processing (default:\\n                                        same as --threads)\\n-C,    --cpu-mask M                     CPU affinity mask: arbitrarily long hex. Complements cpu-range\\n                                        (default: \\"\\")\\n-Cr,   --cpu-range lo-hi                range of CPUs for affinity. Complements --cpu-mask\\n--cpu-strict <0|1>                      use strict CPU placement (default: 0)\\n--prio N                                set process/thread priority : low(-1), normal(0), medium(1), high(2),\\n                                        realtime(3) (default: 0)\\n--poll <0...100>                        use polling level to wait for work (0 - no polling, default: 50)\\n-Cb,   --cpu-mask-batch M               CPU affinity mask: arbitrarily long hex. Complements cpu-range-batch\\n                                        (default: same as --cpu-mask)\\n-Crb,  --cpu-range-batch lo-hi          ranges of CPUs for affinity. Complements --cpu-mask-batch\\n--cpu-strict-batch <0|1>                use strict CPU placement (default: same as --cpu-strict)\\n--prio-batch N                          set process/thread priority : 0-normal, 1-medium, 2-high, 3-realtime\\n                                        (default: 0)\\n--poll-batch <0|1>                      use polling to wait for work (default: same as --poll)\\n-c,    --ctx-size N                     size of the prompt context (default: 0, 0 = loaded from model)\\n                                        (env: LLAMA_ARG_CTX_SIZE)\\n-n,    --predict, --n-predict N         number of tokens to predict (default: -1, -1 = infinity)\\n                                        (env: LLAMA_ARG_N_PREDICT)\\n-b,    --batch-size N                   logical maximum batch size (default: 2048)\\n                                        (env: LLAMA_ARG_BATCH)\\n-ub,   --ubatch-size N                  physical maximum batch size (default: 512)\\n                                        (env: LLAMA_ARG_UBATCH)\\n--keep N                                number of tokens to keep from the initial prompt (default: 0, -1 =\\n                                        all)\\n--swa-full                              use full-size SWA cache (default: false)\\n                                        [(more\\n                                        info)](https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)\\n                                        (env: LLAMA_ARG_SWA_FULL)\\n--kv-unified, -kvu                      use single unified KV buffer for the KV cache of all sequences\\n                                        (default: false)\\n                                        [(more info)](https://github.com/ggml-org/llama.cpp/pull/14363)\\n                                        (env: LLAMA_ARG_KV_UNIFIED)\\n-fa,   --flash-attn [on|off|auto]       set Flash Attention use (\'on\', \'off\', or \'auto\', default: \'auto\')\\n                                        (env: LLAMA_ARG_FLASH_ATTN)\\n-p,    --prompt PROMPT                  prompt to start generation with; for system message, use -sys\\n--perf, --no-perf                       whether to enable internal libllama performance timings (default:\\n                                        false)\\n                                        (env: LLAMA_ARG_PERF)\\n-f,    --file FNAME                     a file containing the prompt (default: none)\\n-bf,   --binary-file FNAME              binary file containing the prompt (default: none)\\n-e,    --escape, --no-escape            whether to process escapes sequences (\\\\n, \\\\r, \\\\t, \\\\\', \\\\\\", \\\\\\\\)\\n                                        (default: true)\\n--rope-scaling {none,linear,yarn}       RoPE frequency scaling method, defaults to linear unless specified by\\n                                        the model\\n                                        (env: LLAMA_ARG_ROPE_SCALING_TYPE)\\n--rope-scale N                          RoPE context scaling factor, expands context by a factor of N\\n                                        (env: LLAMA_ARG_ROPE_SCALE)\\n--rope-freq-base N                      RoPE base frequency, used by NTK-aware scaling (default: loaded from\\n                                        model)\\n                                        (env: LLAMA_ARG_ROPE_FREQ_BASE)\\n--rope-freq-scale N                     RoPE frequency scaling factor, expands context by a factor of 1/N\\n                                        (env: LLAMA_ARG_ROPE_FREQ_SCALE)\\n--yarn-orig-ctx N                       YaRN: original context size of model (default: 0 = model training\\n                                        context size)\\n                                        (env: LLAMA_ARG_YARN_ORIG_CTX)\\n--yarn-ext-factor N                     YaRN: extrapolation mix factor (default: -1.0, 0.0 = full\\n                                        interpolation)\\n                                        (env: LLAMA_ARG_YARN_EXT_FACTOR)\\n--yarn-attn-factor N                    YaRN: scale sqrt(t) or attention magnitude (default: -1.0)\\n                                        (env: LLAMA_ARG_YARN_ATTN_FACTOR)\\n--yarn-beta-slow N                      YaRN: high correction dim or alpha (default: -1.0)\\n                                        (env: LLAMA_ARG_YARN_BETA_SLOW)\\n--yarn-beta-fast N                      YaRN: low correction dim or beta (default: -1.0)\\n                                        (env: LLAMA_ARG_YARN_BETA_FAST)\\n-kvo,  --kv-offload, -nkvo, --no-kv-offload\\n                                        whether to enable KV cache offloading (default: enabled)\\n                                        (env: LLAMA_ARG_KV_OFFLOAD)\\n--repack, -nr, --no-repack              whether to enable weight repacking (default: enabled)\\n                                        (env: LLAMA_ARG_REPACK)\\n--no-host                               bypass host buffer allowing extra buffers to be used\\n                                        (env: LLAMA_ARG_NO_HOST)\\n-ctk,  --cache-type-k TYPE              KV cache data type for K\\n                                        allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1\\n                                        (default: f16)\\n                                        (env: LLAMA_ARG_CACHE_TYPE_K)\\n-ctv,  --cache-type-v TYPE              KV cache data type for V\\n                                        allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1\\n                                        (default: f16)\\n                                        (env: LLAMA_ARG_CACHE_TYPE_V)\\n-dt,   --defrag-thold N                 KV cache defragmentation threshold (DEPRECATED)\\n                                        (env: LLAMA_ARG_DEFRAG_THOLD)\\n-np,   --parallel N                     number of parallel sequences to decode (default: 1)\\n                                        (env: LLAMA_ARG_N_PARALLEL)\\n--mlock                                 force system to keep model in RAM rather than swapping or compressing\\n                                        (env: LLAMA_ARG_MLOCK)\\n--mmap, --no-mmap                       whether to memory-map model (if disabled, slower load but may reduce\\n                                        pageouts if not using mlock) (default: enabled)\\n                                        (env: LLAMA_ARG_MMAP)\\n--numa TYPE                             attempt optimizations that help on some NUMA systems\\n                                        - distribute: spread execution evenly over all nodes\\n                                        - isolate: only spawn threads on CPUs on the node that execution\\n                                        started on\\n                                        - numactl: use the CPU map provided by numactl\\n                                        if run without this previously, it is recommended to drop the system\\n                                        page cache before using this\\n                                        see https://github.com/ggml-org/llama.cpp/issues/1437\\n                                        (env: LLAMA_ARG_NUMA)\\n-dev,  --device <dev1,dev2,..>          comma-separated list of devices to use for offloading (none = don\'t\\n                                        offload)\\n                                        use --list-devices to see a list of available devices\\n                                        (env: LLAMA_ARG_DEVICE)\\n--list-devices                          print list of available devices and exit\\n--override-tensor, -ot <tensor name pattern>=<buffer type>,...\\n                                        override tensor buffer type\\n--cpu-moe, -cmoe                        keep all Mixture of Experts (MoE) weights in the CPU\\n                                        (env: LLAMA_ARG_CPU_MOE)\\n--n-cpu-moe, -ncmoe N                   keep the Mixture of Experts (MoE) weights of the first N layers in the\\n                                        CPU\\n                                        (env: LLAMA_ARG_N_CPU_MOE)\\n-ngl,  --gpu-layers, --n-gpu-layers N   max. number of layers to store in VRAM (default: -1)\\n                                        (env: LLAMA_ARG_N_GPU_LAYERS)\\n-sm,   --split-mode {none,layer,row}    how to split the model across multiple GPUs, one of:\\n                                        - none: use one GPU only\\n                                        - layer (default): split layers and KV across GPUs\\n                                        - row: split rows across GPUs\\n                                        (env: LLAMA_ARG_SPLIT_MODE)\\n-ts,   --tensor-split N0,N1,N2,...      fraction of the model to offload to each GPU, comma-separated list of\\n                                        proportions, e.g. 3,1\\n                                        (env: LLAMA_ARG_TENSOR_SPLIT)\\n-mg,   --main-gpu INDEX                 the GPU to use for the model (with split-mode = none), or for\\n                                        intermediate results and KV (with split-mode = row) (default: 0)\\n                                        (env: LLAMA_ARG_MAIN_GPU)\\n-fit,  --fit [on|off]                   whether to adjust unset arguments to fit in device memory (\'on\' or\\n                                        \'off\', default: \'on\')\\n                                        (env: LLAMA_ARG_FIT)\\n-fitt, --fit-target MiB                 target margin per device for --fit option, default: 1024\\n                                        (env: LLAMA_ARG_FIT_TARGET)\\n-fitc, --fit-ctx N                      minimum ctx size that can be set by --fit option, default: 4096\\n                                        (env: LLAMA_ARG_FIT_CTX)\\n--check-tensors                         check model tensor data for invalid values (default: false)\\n--override-kv KEY=TYPE:VALUE            advanced option to override model metadata by key. may be specified\\n                                        multiple times.\\n                                        types: int, float, bool, str. example: --override-kv\\n                                        tokenizer.ggml.add_bos_token=bool:false\\n--op-offload, --no-op-offload           whether to offload host tensor operations to device (default: true)\\n--lora FNAME                            path to LoRA adapter (can be repeated to use multiple adapters)\\n--lora-scaled FNAME SCALE               path to LoRA adapter with user defined scaling (can be repeated to use\\n                                        multiple adapters)\\n--control-vector FNAME                  add a control vector\\n                                        note: this argument can be repeated to add multiple control vectors\\n--control-vector-scaled FNAME SCALE     add a control vector with user defined scaling SCALE\\n                                        note: this argument can be repeated to add multiple scaled control\\n                                        vectors\\n--control-vector-layer-range START END\\n                                        layer range to apply the control vector(s) to, start and end inclusive\\n-m,    --model FNAME                    model path to load\\n                                        (env: LLAMA_ARG_MODEL)\\n-mu,   --model-url MODEL_URL            model download url (default: unused)\\n                                        (env: LLAMA_ARG_MODEL_URL)\\n-dr,   --docker-repo [<repo>/]<model>[:quant]\\n                                        Docker Hub model repository. repo is optional, default to ai/. quant\\n                                        is optional, default to :latest.\\n                                        example: gemma3\\n                                        (default: unused)\\n                                        (env: LLAMA_ARG_DOCKER_REPO)\\n-hf,   -hfr, --hf-repo <user>/<model>[:quant]\\n                                        Hugging Face model repository; quant is optional, case-insensitive,\\n                                        default to Q4_K_M, or falls back to the first file in the repo if\\n                                        Q4_K_M doesn\'t exist.\\n                                        mmproj is also downloaded automatically if available. to disable, add\\n                                        --no-mmproj\\n                                        example: unsloth/phi-4-GGUF:q4_k_m\\n                                        (default: unused)\\n                                        (env: LLAMA_ARG_HF_REPO)\\n-hfd,  -hfrd, --hf-repo-draft <user>/<model>[:quant]\\n                                        Same as --hf-repo, but for the draft model (default: unused)\\n                                        (env: LLAMA_ARG_HFD_REPO)\\n-hff,  --hf-file FILE                   Hugging Face model file. If specified, it will override the quant in\\n                                        --hf-repo (default: unused)\\n                                        (env: LLAMA_ARG_HF_FILE)\\n-hfv,  -hfrv, --hf-repo-v <user>/<model>[:quant]\\n                                        Hugging Face model repository for the vocoder model (default: unused)\\n                                        (env: LLAMA_ARG_HF_REPO_V)\\n-hffv, --hf-file-v FILE                 Hugging Face model file for the vocoder model (default: unused)\\n                                        (env: LLAMA_ARG_HF_FILE_V)\\n-hft,  --hf-token TOKEN                 Hugging Face access token (default: value from HF_TOKEN environment\\n                                        variable)\\n                                        (env: HF_TOKEN)\\n--log-disable                           Log disable\\n--log-file FNAME                        Log to file\\n                                        (env: LLAMA_LOG_FILE)\\n--log-colors [on|off|auto]              Set colored logging (\'on\', \'off\', or \'auto\', default: \'auto\')\\n                                        \'auto\' enables colors when output is to a terminal\\n                                        (env: LLAMA_LOG_COLORS)\\n-v,    --verbose, --log-verbose         Set verbosity level to infinity (i.e. log all messages, useful for\\n                                        debugging)\\n--offline                               Offline mode: forces use of cache, prevents network access\\n                                        (env: LLAMA_OFFLINE)\\n-lv,   --verbosity, --log-verbosity N   Set the verbosity threshold. Messages with a higher verbosity will be\\n                                        ignored. Values:\\n                                         - 0: generic output\\n                                         - 1: error\\n                                         - 2: warning\\n                                         - 3: info\\n                                         - 4: debug\\n                                        (default: 1)\\n\\n                                        (env: LLAMA_LOG_VERBOSITY)\\n--log-prefix                            Enable prefix in log messages\\n                                        (env: LLAMA_LOG_PREFIX)\\n--log-timestamps                        Enable timestamps in log messages\\n                                        (env: LLAMA_LOG_TIMESTAMPS)\\n-ctkd, --cache-type-k-draft TYPE        KV cache data type for K for the draft model\\n                                        allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1\\n                                        (default: f16)\\n                                        (env: LLAMA_ARG_CACHE_TYPE_K_DRAFT)\\n-ctvd, --cache-type-v-draft TYPE        KV cache data type for V for the draft model\\n                                        allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1\\n                                        (default: f16)\\n                                        (env: LLAMA_ARG_CACHE_TYPE_V_DRAFT)\\n\\n\\n----- sampling params -----\\n\\n--samplers SAMPLERS                     samplers that will be used for generation in the order, separated by\\n                                        \';\'\\n                                        (default:\\n                                        penalties;dry;top_n_sigma;top_k;typ_p;top_p;min_p;xtc;temperature)\\n-s,    --seed SEED                      RNG seed (default: -1, use random seed for -1)\\n--sampling-seq, --sampler-seq SEQUENCE\\n                                        simplified sequence for samplers that will be used (default:\\n                                        edskypmxt)\\n--ignore-eos                            ignore end of stream token and continue generating (implies\\n                                        --logit-bias EOS-inf)\\n--temp N                                temperature (default: 0.8)\\n--top-k N                               top-k sampling (default: 40, 0 = disabled)\\n                                        (env: LLAMA_ARG_TOP_K)\\n--top-p N                               top-p sampling (default: 0.9, 1.0 = disabled)\\n--min-p N                               min-p sampling (default: 0.1, 0.0 = disabled)\\n--top-nsigma N                          top-n-sigma sampling (default: -1.0, -1.0 = disabled)\\n--xtc-probability N                     xtc probability (default: 0.0, 0.0 = disabled)\\n--xtc-threshold N                       xtc threshold (default: 0.1, 1.0 = disabled)\\n--typical N                             locally typical sampling, parameter p (default: 1.0, 1.0 = disabled)\\n--repeat-last-n N                       last n tokens to consider for penalize (default: 64, 0 = disabled, -1\\n                                        = ctx_size)\\n--repeat-penalty N                      penalize repeat sequence of tokens (default: 1.0, 1.0 = disabled)\\n--presence-penalty N                    repeat alpha presence penalty (default: 0.0, 0.0 = disabled)\\n--frequency-penalty N                   repeat alpha frequency penalty (default: 0.0, 0.0 = disabled)\\n--dry-multiplier N                      set DRY sampling multiplier (default: 0.0, 0.0 = disabled)\\n--dry-base N                            set DRY sampling base value (default: 1.75)\\n--dry-allowed-length N                  set allowed length for DRY sampling (default: 2)\\n--dry-penalty-last-n N                  set DRY penalty for the last n tokens (default: -1, 0 = disable, -1 =\\n                                        context size)\\n--dry-sequence-breaker STRING           add sequence breaker for DRY sampling, clearing out default breakers\\n                                        (\'\\\\n\', \':\', \'\\"\', \'*\') in the process; use \\"none\\" to not use any\\n                                        sequence breakers\\n--dynatemp-range N                      dynamic temperature range (default: 0.0, 0.0 = disabled)\\n--dynatemp-exp N                        dynamic temperature exponent (default: 1.0)\\n--mirostat N                            use Mirostat sampling.\\n                                        Top K, Nucleus and Locally Typical samplers are ignored if used.\\n                                        (default: 0, 0 = disabled, 1 = Mirostat, 2 = Mirostat 2.0)\\n--mirostat-lr N                         Mirostat learning rate, parameter eta (default: 0.1)\\n--mirostat-ent N                        Mirostat target entropy, parameter tau (default: 5.0)\\n-l,    --logit-bias TOKEN_ID(+/-)BIAS   modifies the likelihood of token appearing in the completion,\\n                                        i.e. `--logit-bias 15043+1` to increase likelihood of token \' Hello\',\\n                                        or `--logit-bias 15043-1` to decrease likelihood of token \' Hello\'\\n--grammar GRAMMAR                       BNF-like grammar to constrain generations (see samples in grammars/\\n                                        dir) (default: \'\')\\n--grammar-file FNAME                    file to read grammar from\\n-j,    --json-schema SCHEMA             JSON schema to constrain generations (https://json-schema.org/), e.g.\\n                                        `{}` for any JSON object\\n                                        For schemas w/ external $refs, use --grammar +\\n                                        example/json_schema_to_grammar.py instead\\n-jf,   --json-schema-file FILE          File containing a JSON schema to constrain generations\\n                                        (https://json-schema.org/), e.g. `{}` for any JSON object\\n                                        For schemas w/ external $refs, use --grammar +\\n                                        example/json_schema_to_grammar.py instead\\n\\n\\n----- example-specific params -----\\n\\n--display-prompt, --no-display-prompt   whether to print prompt at generation (default: true)\\n-co,   --color [on|off|auto]            Colorize output to distinguish prompt and user input from generations\\n                                        (\'on\', \'off\', or \'auto\', default: \'auto\')\\n                                        \'auto\' enables colors when output is to a terminal\\n--ctx-checkpoints, --swa-checkpoints N\\n                                        max number of context checkpoints to create per slot (default: 8)\\n                                        [(more info)](https://github.com/ggml-org/llama.cpp/pull/15293)\\n                                        (env: LLAMA_ARG_CTX_CHECKPOINTS)\\n--cache-ram, -cram N                    set the maximum cache size in MiB (default: 8192, -1 - no limit, 0 -\\n                                        disable)\\n                                        [(more info)](https://github.com/ggml-org/llama.cpp/pull/16391)\\n                                        (env: LLAMA_ARG_CACHE_RAM)\\n--context-shift, --no-context-shift     whether to use context shift on infinite text generation (default:\\n                                        disabled)\\n                                        (env: LLAMA_ARG_CONTEXT_SHIFT)\\n-sys,  --system-prompt PROMPT           system prompt to use with model (if applicable, depending on chat\\n                                        template)\\n--show-timings, --no-show-timings       whether to show timing information after each response (default: true)\\n                                        (env: LLAMA_ARG_SHOW_TIMINGS)\\n-sysf, --system-prompt-file FNAME       a file containing the system prompt (default: none)\\n-r,    --reverse-prompt PROMPT          halt generation at PROMPT, return control in interactive mode\\n-sp,   --special                        special tokens output enabled (default: false)\\n-cnv,  --conversation, -no-cnv, --no-conversation\\n                                        whether to run in conversation mode:\\n                                        - does not print special tokens and suffix/prefix\\n                                        - interactive mode is also enabled\\n                                        (default: auto enabled if chat template is available)\\n-st,   --single-turn                    run conversation for a single turn only, then exit when done\\n                                        will not be interactive if first turn is predefined with --prompt\\n                                        (default: false)\\n-mli,  --multiline-input                allows you to write or paste multiple lines without ending each in \'\\\\\'\\n--warmup, --no-warmup                   whether to perform warmup with an empty run (default: enabled)\\n-mm,   --mmproj FILE                    path to a multimodal projector file. see tools/mtmd/README.md\\n                                        note: if -hf is used, this argument can be omitted\\n                                        (env: LLAMA_ARG_MMPROJ)\\n-mmu,  --mmproj-url URL                 URL to a multimodal projector file. see tools/mtmd/README.md\\n                                        (env: LLAMA_ARG_MMPROJ_URL)\\n--mmproj-auto, --no-mmproj, --no-mmproj-auto\\n                                        whether to use multimodal projector file (if available), useful when\\n                                        using -hf (default: enabled)\\n                                        (env: LLAMA_ARG_MMPROJ_AUTO)\\n--mmproj-offload, --no-mmproj-offload   whether to enable GPU offloading for multimodal projector (default:\\n                                        enabled)\\n                                        (env: LLAMA_ARG_MMPROJ_OFFLOAD)\\n--image, --audio FILE                   path to an image or audio file. use with multimodal models, can be\\n                                        repeated if you have multiple files\\n--image-min-tokens N                    minimum number of tokens each image can take, only used by vision\\n                                        models with dynamic resolution (default: read from model)\\n                                        (env: LLAMA_ARG_IMAGE_MIN_TOKENS)\\n--image-max-tokens N                    maximum number of tokens each image can take, only used by vision\\n                                        models with dynamic resolution (default: read from model)\\n                                        (env: LLAMA_ARG_IMAGE_MAX_TOKENS)\\n--override-tensor-draft, -otd <tensor name pattern>=<buffer type>,...\\n                                        override tensor buffer type for draft model\\n--cpu-moe-draft, -cmoed                 keep all Mixture of Experts (MoE) weights in the CPU for the draft\\n                                        model\\n                                        (env: LLAMA_ARG_CPU_MOE_DRAFT)\\n--n-cpu-moe-draft, -ncmoed N            keep the Mixture of Experts (MoE) weights of the first N layers in the\\n                                        CPU for the draft model\\n                                        (env: LLAMA_ARG_N_CPU_MOE_DRAFT)\\n--chat-template-kwargs STRING           sets additional params for the json template parser\\n                                        (env: LLAMA_CHAT_TEMPLATE_KWARGS)\\n--jinja, --no-jinja                     whether to use jinja template engine for chat (default: enabled)\\n                                        (env: LLAMA_ARG_JINJA)\\n--reasoning-format FORMAT               controls whether thought tags are allowed and/or extracted from the\\n                                        response, and in which format they\'re returned; one of:\\n                                        - none: leaves thoughts unparsed in `message.content`\\n                                        - deepseek: puts thoughts in `message.reasoning_content`\\n                                        - deepseek-legacy: keeps `<think>` tags in `message.content` while\\n                                        also populating `message.reasoning_content`\\n                                        (default: auto)\\n                                        (env: LLAMA_ARG_THINK)\\n--reasoning-budget N                    controls the amount of thinking allowed; currently only one of: -1 for\\n                                        unrestricted thinking budget, or 0 to disable thinking (default: -1)\\n                                        (env: LLAMA_ARG_THINK_BUDGET)\\n--chat-template JINJA_TEMPLATE          set custom jinja chat template (default: template taken from model\'s\\n                                        metadata)\\n                                        if suffix/prefix are specified, template will be disabled\\n                                        only commonly used templates are accepted (unless --jinja is set\\n                                        before this flag):\\n                                        list of built-in templates:\\n                                        bailing, bailing-think, bailing2, chatglm3, chatglm4, chatml,\\n                                        command-r, deepseek, deepseek2, deepseek3, exaone3, exaone4, falcon3,\\n                                        gemma, gigachat, glmedge, gpt-oss, granite, grok-2, hunyuan-dense,\\n                                        hunyuan-moe, kimi-k2, llama2, llama2-sys, llama2-sys-bos,\\n                                        llama2-sys-strip, llama3, llama4, megrez, minicpm, mistral-v1,\\n                                        mistral-v3, mistral-v3-tekken, mistral-v7, mistral-v7-tekken, monarch,\\n                                        openchat, orion, pangu-embedded, phi3, phi4, rwkv-world, seed_oss,\\n                                        smolvlm, vicuna, vicuna-orca, yandex, zephyr\\n                                        (env: LLAMA_ARG_CHAT_TEMPLATE)\\n--chat-template-file JINJA_TEMPLATE_FILE\\n                                        set custom jinja chat template file (default: template taken from\\n                                        model\'s metadata)\\n                                        if suffix/prefix are specified, template will be disabled\\n                                        only commonly used templates are accepted (unless --jinja is set\\n                                        before this flag):\\n                                        list of built-in templates:\\n                                        bailing, bailing-think, bailing2, chatglm3, chatglm4, chatml,\\n                                        command-r, deepseek, deepseek2, deepseek3, exaone3, exaone4, falcon3,\\n                                        gemma, gigachat, glmedge, gpt-oss, granite, grok-2, hunyuan-dense,\\n                                        hunyuan-moe, kimi-k2, llama2, llama2-sys, llama2-sys-bos,\\n                                        llama2-sys-strip, llama3, llama4, megrez, minicpm, mistral-v1,\\n                                        mistral-v3, mistral-v3-tekken, mistral-v7, mistral-v7-tekken, monarch,\\n                                        openchat, orion, pangu-embedded, phi3, phi4, rwkv-world, seed_oss,\\n                                        smolvlm, vicuna, vicuna-orca, yandex, zephyr\\n                                        (env: LLAMA_ARG_CHAT_TEMPLATE_FILE)\\n--simple-io                             use basic IO for better compatibility in subprocesses and limited\\n                                        consoles\\n--draft-max, --draft, --draft-n N       number of tokens to draft for speculative decoding (default: 16)\\n                                        (env: LLAMA_ARG_DRAFT_MAX)\\n--draft-min, --draft-n-min N            minimum number of draft tokens to use for speculative decoding\\n                                        (default: 0)\\n                                        (env: LLAMA_ARG_DRAFT_MIN)\\n--draft-p-min P                         minimum speculative decoding probability (greedy) (default: 0.8)\\n                                        (env: LLAMA_ARG_DRAFT_P_MIN)\\n-cd,   --ctx-size-draft N               size of the prompt context for the draft model (default: 0, 0 = loaded\\n                                        from model)\\n                                        (env: LLAMA_ARG_CTX_SIZE_DRAFT)\\n-devd, --device-draft <dev1,dev2,..>    comma-separated list of devices to use for offloading the draft model\\n                                        (none = don\'t offload)\\n                                        use --list-devices to see a list of available devices\\n-ngld, --gpu-layers-draft, --n-gpu-layers-draft N\\n                                        number of layers to store in VRAM for the draft model\\n                                        (env: LLAMA_ARG_N_GPU_LAYERS_DRAFT)\\n-md,   --model-draft FNAME              draft model for speculative decoding (default: unused)\\n                                        (env: LLAMA_ARG_MODEL_DRAFT)\\n--spec-replace TARGET DRAFT             translate the string in TARGET into DRAFT if the draft model and main\\n                                        model are not compatible\\n--gpt-oss-20b-default                   use gpt-oss-20b (note: can download weights from the internet)\\n--gpt-oss-120b-default                  use gpt-oss-120b (note: can download weights from the internet)\\n--vision-gemma-4b-default               use Gemma 3 4B QAT (note: can download weights from the internet)\\n--vision-gemma-12b-default              use Gemma 3 12B QAT (note: can download weights from the internet)\\n```\\n"}]}],"stream":false,"model":"ggml-org/gpt-oss-20b-GGUF","reasoning_format":"auto","temperature":0.3,"max_tokens":-1,"dynatemp_range":0,"dynatemp_exponent":1,"top_k":30,"top_p":0.95,"min_p":0.05,"xtc_probability":0,"xtc_threshold":0.1,"typ_p":1,"repeat_last_n":64,"repeat_penalty":1.1,"presence_penalty":0,"frequency_penalty":0,"dry_multiplier":0,"dry_base":1.75,"dry_allowed_length":2,"dry_penalty_last_n":8192,"samplers":["penalties","dry","top_n_sigma","top_k","typ_p","top_p","min_p","xtc","temperature"],"timings_per_token":true}' 

Here #17636 (comment) also there is the same message.

First Bad Commit

No response

Relevant log output

srv    operator(): http client error: Failed to read connection

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions