Replies: 5 comments 1 reply
|
throw my existing thought first
|
|
Will be officially integrated, but maybe, with an open-source interface standard. |
|
Another promising use case is letting Jev choose the retrieval strategy for each query : semantic, keywords, or hybrid, and whether reranking is needed. We can map its decision to a few pre-calibrated weight profiles. This could sit between query planning and retrieval as an optional policy layer. |
|
Three places where a bounded decision model has the clearest leverage, and all of them sit at the recall boundary rather than inside ranking: 1. At user-input time: decide whether to recall at all. As far as I can tell from the code paths cited above, retrieval runs per turn rather than being gated on whether the turn needs memory at all. A large share of turns in a long session — "run the tests", "continue", a follow-up on a tool result — reference nothing in storage. A single 2. After recall: decide per memory whether it is actually needed. Ranking produces an ordering; it cannot say "none of these belong in the prompt". A per-candidate judgment ("is this memory required to answer this request?") fits the independent, parallel shape Jev is suited to, avoiding the failure mode of a single large 3. Query rewriting: decide whether the input needs rewriting at all. Rewriting is enabled by default in the plugin today, so every turn pays the cost whether or not the input is under-specified. Many inputs are already self-contained; those that are not typically depend on a referent from earlier in the thread. A Context width is the accuracy knob, and it is priced. All three decisions improve significantly with more agent context: the last few turns, the active task, and what the agent already has loaded. Because these are bounded decisions rather than generation, that context can be scaled per call — a one-line summary when downstream work is cheap, or the recent transcript when downstream work involves an expensive generation or a memory write. It is worth treating context width as an explicit cost/accuracy parameter of the integration rather than a fixed prompt. None of this is measured yet. It is simply where I would look first, because all three decisions are observable: you can replay a session, compare what the gate refused against what retrieval would have returned, and count the turns where it was right. |
|
On point 1 (Jev vs an LLM prompted into the same response format): we measured that trade-off against human labels. I maintain Jevals, an independent benchmark.
Boards: https://jevals.com/choice/ |
Uh oh!
There was an error while loading. Please reload this page.
Background
This is an architectural discussion based on OpenViking source and TypeSafe documentation. The scenarios and benefits below are proposals; no OpenViking/Jev integration results are available yet.
Why consider Jev?
Is it worth integrating Jev into OpenViking? Which scenarios would benefit, and would those benefits justify the integration and maintenance cost?
OpenViking manages resources, memories, and skills for agents. Across their lifecycle, it needs semantic decisions: what context helps with this request, which skill applies, whether a candidate memory is worth retaining, and how new information relates to existing memory.
Jev is a possible component for these decisions. It returns a selection (
Choice), a probability that a condition holds (Noul), or a graded judgment (Score). The architectural opportunity is to let a specialized decision model handle suitable bounded judgments while OpenViking retains retrieval, workflow, and storage control and generative models continue producing queries, summaries, and memory text.The potential payoff is better context and memory quality, less unnecessary downstream work, or cheaper/faster semantic decisions. None of these is automatic: OpenViking already has relevant mechanisms, and another model call may add cost without improving outcomes. This discussion asks where the combination makes sense, rather than assuming Jev should become a core dependency or treating reranking as the only possible use case.
Possible applications and benefits
Scenarios worth discussing
A Jev rerank provider is one possible implementation of the first scenario, not the whole proposal. Likewise, a single large
Choiceshould not substitute for independent relevance judgments when multiple memories or skills may be useful.Where the value would come from
There are two different integration cases:
Structured answers can make outputs easier to consume and individual decisions easier to inspect. That is useful engineering behavior, but it does not by itself establish accuracy or distinguish Jev from every existing structured-output alternative.
Fit with the existing architecture
Source references are pinned to
a55afb001bca201f09830b90dc2f64eed20cb799:If we proceed, integration should be optional, reuse existing boundaries where possible, and preserve baseline behavior on provider failure. Provider availability, data-handling requirements, deployment compatibility, and ongoing model/prompt calibration are part of the cost. A general decision abstraction should follow demonstrated needs rather than precede them.
Initial assessment
There is a reasonable case for exploring Jev as an optional component, but insufficient evidence to recommend a default or broad core integration.
Context selection and skill matching are approachable starting points because they operate on observable candidates and their outputs can be compared with existing behavior. Memory admission and relationship classification could offer longer-term memory-quality benefits, but errors persist and their interaction with generation makes the payoff harder to establish. Routing is attractive only where there is meaningful unnecessary retrieval to avoid.
These are prioritization suggestions for discussion, not a decision to limit the work to reranking. Maintainers' experience with actual bottlenecks should determine the first scenario.
Alternatives Considered
Existing rerankers, tuned retrieval/assembly policies, deterministic rules, and general-purpose models remain valid alternatives. If they deliver the same benefit with less complexity, there is no reason to add Jev. Jev is also not a replacement for embeddings, open-ended reasoning, query rewriting, summary generation, or memory-text synthesis.
Use Case
For deployment assistance, the desired result may be a relevant guide, a prior failure containing an environment constraint, and an applicable deployment skill. The value of Jev would be helping OpenViking supply this useful combination—or maintain cleaner memories for later requests—with a better quality/cost tradeoff. Simply returning valid selections is not the outcome we care about.
Additional Context
Questions for core developers:
After choosing a scenario, a small representative replay/shadow comparison against the existing approach can validate the hypothesis. Measure task-specific quality and consequential errors together with total latency/cost; for context-related changes, include final answer quality and actual injected tokens. Keep tuning and held-out samples separate. Detailed evaluation design can follow agreement on the use case.
Related issues: #2880 (rerank budgets) and #4948 (closed historical latency report), relevant to the retrieval scenario but not evidence that Jev resolves it.
References: TypeSafe primitives, reranking example, skill suggestion example, confidence semantics. These document feasibility patterns, not measured OpenViking benefits.
Originally proposed in #5210; continuing here because the integration scope and value are still open for discussion.
All reactions