Skip to content

Retrieval Benchmark

JanYork edited this page Sep 13, 2026 · 3 revisions

LongMemEval-S retrieval results (v0.18.5)

Language: English · 简体中文

This is an untuned retrieval baseline, not a ceiling on LWC’s effectiveness over continued use. In real workflows, the model can proactively assess evidence and react to user corrections or relevance feedback, continually updating document weights and query feedback. With reliable feedback and sustained tuning, retrieval is expected to improve beyond this static baseline; that additional gain was not quantified here. Reactive adjustment means feedback-triggered Agent updates, not automatic background learning on every query.

On September 13, 2026, two complete local runs processed 500/500 questions, with 470 retrieval-scored questions, using the pinned LongMemEval-S dataset and four concurrent workers on an Apple M5 Pro Mac. The second run reused the first run's stored sources; all 500 ranked result lists were identical.

Metric First run Second run
Questions processed 500 / 500 500 / 500
Retrieval-scored questions 470 470
Excluded abstention questions 30 30
Recall@1 83.83% (394/470) 83.83% (394/470)
Recall@3 91.49% (430/470) 91.49% (430/470)
Recall@5 95.11% (447/470) 95.11% (447/470)
Recall@10 97.66% (459/470) 97.66% (459/470)
Recall@30 99.15% (466/470) 99.15% (466/470)
Recall@50 99.15% (466/470) 99.15% (466/470)
MRR 0.883668 0.883668
Mean latency 509.883 ms 679.759 ms
P50 507.697 ms 665.454 ms
P90 663.061 ms 931.238 ms
P95 731.766 ms 990.375 ms
P99 888.555 ms 1094.886 ms

These results exclude proactive tuning. The adapter retrieves raw session sources without model-led knowledge curation, relevance feedback, or weight updates. In actual Agent workflows, a model can proactively assess evidence, curate knowledge, and explicitly adjust document weights or query-specific feedback to improve subsequent retrieval. This is an evidence-driven Agent workflow, not automatic weight changes on every search; gains depend on feedback quality and were not measured here. Evaluate tuning on held-out questions rather than feeding test answers back into the same evaluation.

These are local retrieval results, not official leaderboard scores or answer accuracy. Latencies describe four-worker load, not a controlled version-speedup comparison. First-run data · Second-run data · Reproduction and scoring.

LWC Wiki

English · 简体中文


Start here · 开始使用

Core capabilities · 核心能力

Practical guides · 实战指南

Capability configuration · 能力配置

Technical design · 技术设计

Operations · 运行与维护

Reference · 参考资料

Contributing · 参与贡献


Repository · Releases

Clone this wiki locally