Head-to-head benchmark thread — Cognee vs nautilus-compass (LongMemEval-S, evidence open) #5070
chunxiaoxx
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
We published a full-500 LongMemEval-S head-to-head vs mem0 (retrieval P@1 0.890 vs 0.774, identical questions/criteria, ~$3.50 to reproduce: https://github.com/chunxiaoxx/nautilus-compass — reproduction script + per-question evidence in the repo). Our bet is no-extraction-at-write (verbatim + local embedding, graph-free); Cognee's is knowledge-graph ECL.
Rather than claim victory from one benchmark, we'd like to invite a cross-run: we run Cognee through our harness, you run us through yours, publish both. Interested?
(Context: independent benchmark threads are welcome on our Reproducibility Wall too — contradicting numbers get published with the same prominence as favorable ones.)
All reactions