You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
There is now a way to score a LongMemEval-S retrieval run without running the tool that published the number. -score takes a JSON file mapping each question_id to the session ids your system ranked, best first, and prints hit@1/@5/@10/@20, MRR and the dataset's own per-evidence recall. No index, no ingest, nothing of ours in the loop — it reads the dataset and your file, and finishes in seconds.
Disclosure: I maintain deja-vu, which has a row in that table, and that is why the driver is the part worth having rather than the row. Our figures are 85.3% hit@1 on the 470-question cleaned set (-skip-abs) and 84.8% on all 500, both committed as artifacts with the dataset's sha256 beside them. Retrieval only, no answering model — a different metric from an end-to-end QA score and not comparable to one.
If cognee's number on the same questions is better, a pull request against that table is worth more to somebody choosing between memory tools than another paragraph from either of us. And if there is a setting I have wrong that changes the result, say which one and I will re-run it.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
There is now a way to score a LongMemEval-S retrieval run without running the tool that published the number.
-scoretakes a JSON file mapping eachquestion_idto the session ids your system ranked, best first, and prints hit@1/@5/@10/@20, MRR and the dataset's own per-evidence recall. No index, no ingest, nothing of ours in the loop — it reads the dataset and your file, and finishes in seconds.Format, the two file shapes it accepts, and the rules: https://github.com/vshulcz/deja-vu/blob/main/docs/benchmarks/SUBMISSION.md
Disclosure: I maintain deja-vu, which has a row in that table, and that is why the driver is the part worth having rather than the row. Our figures are 85.3% hit@1 on the 470-question cleaned set (
-skip-abs) and 84.8% on all 500, both committed as artifacts with the dataset's sha256 beside them. Retrieval only, no answering model — a different metric from an end-to-end QA score and not comparable to one.If cognee's number on the same questions is better, a pull request against that table is worth more to somebody choosing between memory tools than another paragraph from either of us. And if there is a setting I have wrong that changes the result, say which one and I will re-run it.
All reactions