Skip to content

v0.3.3

Choose a tag to compare

@tehw0lf tehw0lf released this 10 Apr 19:52
· 19 commits to main since this release
382fb8c

What's Changed

Added

  • Benchmark: Evaluated on LongMemEval (ICLR 2025) — 90.2% R@5 retrieval, 38.3% end-to-end QA accuracy with GPT-5 as judge over 470 non-abstention questions
  • Benchmark scripts: scripts/longmemeval_eval.py and scripts/requirements-benchmark.txt for reproducing results
  • Benchmark results: Raw results in results/ directory

Changed

  • Dependencies: Updated dev dependencies (ajv, brace-expansion, flatted, js-yaml, minimatch, picomatch, vite) via npm audit fix

See CHANGELOG.md for full details.