-
Notifications
You must be signed in to change notification settings - Fork 6
Retrieval Benchmark
Language: English · 简体中文
This is an untuned retrieval baseline, not a ceiling on LWC’s effectiveness over continued use. In real workflows, the model can proactively assess evidence and react to user corrections or relevance feedback, continually updating document weights and query feedback. With reliable feedback and sustained tuning, retrieval is expected to improve beyond this static baseline; that additional gain was not quantified here. Reactive adjustment means feedback-triggered Agent updates, not automatic background learning on every query.
On September 13, 2026, two complete local runs processed 500/500 questions, with 470 retrieval-scored questions, using the pinned LongMemEval-S dataset and four concurrent workers on an Apple M5 Pro Mac. The second run reused the first run's stored sources; all 500 ranked result lists were identical.
| Metric | First run | Second run |
|---|---|---|
| Questions processed | 500 / 500 | 500 / 500 |
| Retrieval-scored questions | 470 | 470 |
| Excluded abstention questions | 30 | 30 |
| Recall@1 | 83.83% (394/470) | 83.83% (394/470) |
| Recall@3 | 91.49% (430/470) | 91.49% (430/470) |
| Recall@5 | 95.11% (447/470) | 95.11% (447/470) |
| Recall@10 | 97.66% (459/470) | 97.66% (459/470) |
| Recall@30 | 99.15% (466/470) | 99.15% (466/470) |
| Recall@50 | 99.15% (466/470) | 99.15% (466/470) |
| MRR | 0.883668 | 0.883668 |
| Mean latency | 509.883 ms | 679.759 ms |
| P50 | 507.697 ms | 665.454 ms |
| P90 | 663.061 ms | 931.238 ms |
| P95 | 731.766 ms | 990.375 ms |
| P99 | 888.555 ms | 1094.886 ms |
These results exclude proactive tuning. The adapter retrieves raw session sources without model-led knowledge curation, relevance feedback, or weight updates. In actual Agent workflows, a model can proactively assess evidence, curate knowledge, and explicitly adjust document weights or query-specific feedback to improve subsequent retrieval. This is an evidence-driven Agent workflow, not automatic weight changes on every search; gains depend on feedback quality and were not measured here. Evaluate tuning on held-out questions rather than feeding test answers back into the same evaluation.
These are local retrieval results, not official leaderboard scores or answer accuracy. Latencies describe four-worker load, not a controlled version-speedup comparison. First-run data · Second-run data · Reproduction and scoring.
LWC Wiki
- Architecture overview · 总体架构
- Storage and data model · 存储与数据模型
- Retrieval and indexing · 检索与索引设计
- Graph projection and performance · 图投影与性能设计
- MCP, Hooks, and AgentTarget design · MCP、Hook 与 AgentTarget 设计
- Safety and trust boundaries · 安全模型与信任边界
- Maintenance and diagnostics · 维护与诊断
- Troubleshooting and FAQ · 故障排查与常见问题
- Migration and compatibility · 迁移与版本兼容
- Support and issue reporting · 获取帮助与问题反馈
- Contributing and development · 贡献与开发指南
- Testing and release process · 测试与发布流程
- Wiki style guide · Wiki 编写规范