Repository navigation
Unable to get 32K–64K output tokens with OpenRouter / DeepSeek despite increasing Max Tokens #16533
Replies: 1 comment
|
An 8,192-token response alone does not establish a LibreChat-wide output cap. I would first check the actual outbound parameter and the provider's termination reason. Endpoint/key configuration For a custom endpoint, A configuration check LibreChat's default-parameters documentation notes that addParams:
max_tokens: 65536This is only the addition to your endpoint, not a complete config. Ensure Railway loads that edited file after redeployment. Check that Inspect the provider response
OpenRouter documents that reasoning and visible output share the completion budget on most providers. Thus 20K usage does not mean 20K final-answer text. The current direct DeepSeek API docs also list an 8K default in non-thinking mode when Finally, run the same prompt/model/limit directly against the provider. If that also ends early, a LibreChat setting cannot force a longer answer. If only LibreChat differs, compare the outbound payloads. Could you share your LibreChat version/image tag, exact model ID, and redacted endpoint block (no API keys)? I have not reproduced your Railway deployment, so I would not claim a specific clamp without those details. |
Uh oh!
There was an error while loading. Please reload this page.
Hi,
I deployed LibreChat on Railway and I am trying to generate very long responses (around 32K–64K output tokens) for detailed research and long-form Q&A.
My setup
I deployed LibreChat using Railway.
I added my OpenRouter and DeepSeek API keys separately through Railway Variables. Both providers appear to work correctly in LibreChat and I can chat with their models without authentication or connection problems.
However, I noticed that under LibreChat's API management section, I don't see an option to add/connect OpenRouter or DeepSeek directly.
Is this expected when the API keys are provided through environment variables, or am I missing part of the endpoint configuration?
Main problem: output token limit
My main issue is response length.
When using models through OpenRouter, as well as DeepSeek directly, responses usually seem to stop somewhere around 6K–10K tokens.
In many tests, the output appeared to hit approximately 8,192 tokens.
I tried changing Max Tokens in LibreChat's Advanced Parameters to:
but this did not seem to increase the actual output length. Setting Max Tokens to 32K or 64K in the UI did not result in responses anywhere near those lengths.
I also experimented multiple times with "librechat.yaml", including configurations using "addParams" / "max_tokens" and related token/context settings.
Unfortunately, I could not get this to work either. The responses still stopped far below the configured 32K or 64K value.
With reasoning/thinking enabled, I managed to see total token usage around 20K in some cases, but I still could not get anything close to a 32K or 64K final response.
What I am trying to achieve
I would like to use models that support large output limits to generate extremely detailed answers of approximately:
32,000–64,000 output tokens in a single response.
I understand that the model/provider itself must support that output length. My question is specifically about how LibreChat handles these limits.
Questions
Is there an internal LibreChat output limit that could cause responses to stop around 8K tokens even when Max Tokens is set to 32K or 64K?
Does the Max Tokens value in Advanced Parameters get passed directly to OpenRouter/DeepSeek, or can LibreChat override/clamp this value?
Is there a difference between the Max Tokens UI setting and "max_tokens" / "max_completion_tokens" actually sent to the provider?
When using "OPENROUTER_KEY" and "DEEPSEEK_API_KEY" through Railway environment variables, is it normal that OpenRouter and DeepSeek don't appear as options under the API management/add API section?
Could my environment-variable setup be causing LibreChat to use some default token configuration (for example, an 8,192-token output limit)?
Is "addParams" supposed to be able to override this limit? I tried configuring it several times, but the actual response length did not change.
Assuming the selected model and provider support 32K–64K output, what is the correct way to make LibreChat actually request and receive a response of that length?
At this point I am not sure whether I am encountering a LibreChat limitation, a configuration problem, an OpenRouter/DeepSeek limitation, or simply misunderstanding how LibreChat handles "max_tokens".
If 32K–64K single-response outputs are currently not possible through LibreChat, I would also appreciate confirmation of that.
Thanks for any guidance.
All reactions