Multi-Pass Re-Ranking - Add native support for a generalized, configurable multi-pass reranking pipeline rather than limiting the system to a single global provider. #5068
jasonpsimon
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Feature Description
Add native support for a generalized, configurable multi-pass reranking pipeline (e.g., Local/Pointwise -> Remote/Listwise/Generative LLM) rather than limiting the system to a single global provider.
Core Motivation
Currently, Hindsight requires selecting a single provider via
HINDSIGHT_API_RERANKER_PROVIDER. However, modern retrieval architectures rely heavily on progressive filtering to optimize the tradeoff between accuracy, context window limitations, API costs, and latency.For instance, a developer might want to perform an aggressive, cheap first pass on a massive candidate pool (30–100 chunks) and then use a highly precise, context-heavy model on a tightly refined subset (5–15 chunks). Different combinations serve different architectural needs:
local(Fast, free MiniLM pointwise) ──>typesafe(Jev listwise compression).zeroentropy(Pinpoint cross-encoder precision) ──>cohereorsiliconflow(Qwen3 / Gemma2 generative reasoning heavy-lifters).Having to choose only one provider forces a compromise between raw recall depth and downstream context window efficiency.
Proposed Architecture / UX
Introduce an array or ordered pipeline configuration string along with progressive limits for each tier:
The framework would sequentially route the top
Ncandidates from the prior pass into the input structure required by the next provider in the pipeline.Alternative Approaches Tried
Implementing this currently requires developers to break out of the native framework orchestration. We must set Hindsight to return an intentionally inflated chunk count, intercept the output array via the SDK client, and manually handle the subsequent reranking steps via external API requests. Native pipeline support would drastically clean up production RAG and agent codebases.
All reactions