Your retrieval budget may be based on a context window you are not actually being served #12969
behrnt-slatgng
started this conversation in
General
Replies: 1 comment 1 reply
|
Haystack currently gives you the pieces for a fail-closed guard, but not provider-side context attestation:
A practical pipeline boundary is:
Keep provider rejection as a second guard and record the observed limit, but do not retry by silently dropping context. This also makes cross-provider retrieval tests separable: hold the index/query fixed and report retrieved evidence, rendered tokens, accepted/rejected, and answer citations. Relevant docs: |
1 reply
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Disclosure up front: I build Grunz, a hosted chat + coding agent on open-weight models. Not pitching anything — this is a retrieval-sizing problem I keep hitting and I suspect it silently affects more Haystack pipelines than people realise.
Your retrieval budget is probably based on a number you are not actually being served
Splitter length, top_k, and reranking are all tuned against an assumed context budget. That budget almost always comes off the model card.
The model card number is a property of the weights. The number you actually receive is a property of whoever serves them — a KV-cache budget traded against concurrency. Across the hosted endpoints I have tested, most serve around 32K regardless of what the card claims. A few reach 256K. I have not found one actually serving the 1M figures in release posts.
If you tuned top_k against an assumed 128K and you are being served 32K, you are over-retrieving. And the failure is not an error — it is truncation.
Why this is nastier in RAG than in chat
In chat, losing early context degrades gracefully. In RAG it degrades misleadingly:
It is also provider-dependent, so the same index and the same query can behave differently across two providers serving the same model. That reads as non-determinism rather than as a configuration difference.
Questions
Does Haystack expose the effective served context for the generator in use, as opposed to a static value from model metadata? If a component or the token counter is already doing something smarter here I would like to be pointed at it.
Is there a pattern for failing loudly rather than truncating silently? What I want is for an assembled prompt exceeding the real ceiling to raise, not get quietly cut, even at the cost of a failed query. A silently truncated RAG answer is worse than no answer because it looks fine.
Has anyone benchmarked retrieval quality across providers holding the index constant? If the served-ceiling effect is as large as I think, provider choice should show up as a retrieval-quality difference, which is a slightly alarming result worth having.
Discovery method in the meantime, in case it helps anyone: send a deliberately oversized prompt and read the error text — providers usually leak the true maximum there even when the docs do not.
All reactions