Tuning chunk size and top-k against a context window you are not being served #20245
behrnt-slatgng
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Disclosure first: I build Grunz, a hosted chat + coding agent on open-weight models. Nothing to sell — RAGFlow gives users a lot of control over chunking and retrieval, which makes it one of the places this problem is most worth knowing about.
The budget you tune against may not be the budget you get
Chunk size, top-k and reranking are tuned against an assumed context budget, and that budget almost always comes off the model card.
The card number is a property of the weights. The number you receive is whatever the operator configured — a KV-cache budget traded against concurrency. Across hosted endpoints I have tested, most serve around 32K regardless of what the card claims. A few reach 256K. I have not found one actually serving the 1M figures that appear in release posts.
It is rarely documented and it fails silently. No error. The request is served against whichever ceiling is lowest and the front of the prompt is gone.
Why this is worse in RAG than in chat
In chat, losing early context degrades gracefully. In RAG it degrades misleadingly:
It is also provider-dependent, so the same knowledge base and the same question can answer differently across two providers serving the same model. That reads as non-determinism rather than as configuration.
The specific risk for a heavily-tunable system
RAGFlow encourages people to tune. Someone tunes chunk size and top-k against an assumed window, gets a configuration that works, then switches model provider — and the tuning silently stops being valid. Nothing in the UI would tell them.
Questions
Does RAGFlow know the effective served context for the configured model, or does it work from static metadata? If something already probes this I would like to be pointed at it.
Could the assembled prompt be checked against the real ceiling before sending — dropping whole chunks deliberately and saying so, rather than letting the provider silently cut the front? Deliberately dropping chunk 8 is much better than losing half of chunk 1, and the user could be shown it happened.
Has anyone compared answer quality for the same knowledge base across providers? If the ceiling effect is as large as I think, provider choice should show up as a retrieval-quality difference that is currently attributed to model quality.
Discovery method in the meantime: send a deliberately oversized prompt and read the error text — providers usually leak the true maximum there even when the docs do not.
All reactions