Skip to content

v1.2.0

Choose a tag to compare

@github-actions github-actions released this 09 Oct 06:26
· 10 commits to main since this release
c50d555

Highlights

Issue localisation. Paste an issue or bug report and get_context_packet returns where_to_look: the files to open next, best first. Measured on SWE-bench Lite test, run once after tuning only on the SWE-bench dev split (276 strict-holdout issues, file-level Acc@1/3/5/10):

@1 @3 @5 @10
plain BM25 0.301 0.507 0.587 0.721
1.2.0 0.486 0.710 0.750 0.804

CPU only, no LLM, no embeddings — p50 2.8 s per issue including a cold index.

Cold index ~8× faster on large repositories (django: first packet 3m20s → 26 s, same graph).

Details: docs/CHANGELOG.md, docs/measured.md.

What's Changed

  • Plain-language: trust the whole-question ranking; fix skeleton panic; baseline comparison script by @ParsaVictor in #128
  • Measured: outside baselines + fresh click holdout; packet: folds once, server guesses not 'missing' by @ParsaVictor in #129
  • Optional code-aware embeddings (jina-code v2): ripgrep plain-language 0.500 → 0.667 recall, 0.156 → 0.347 precision by @ParsaVictor in #130
  • Issue mode: long reports use the whole-question ranking (SWE-bench Lite dev hit@3 0.250 → 0.583) by @ParsaVictor in #131
  • Index speed: django cold index 3m20s → 26s, same graph by @ParsaVictor in #132
  • Disk: bound target/ (debug 43 GB, index cache 6.7 GB) by @ParsaVictor in #133
  • Issue mode: definition-level ranking + report hygiene (SWE dev Acc@5 0.456 → 0.596) by @ParsaVictor in #134
  • 1.2.0: issue localisation (SWE-bench Lite holdout Acc@5 0.587 → 0.750 vs BM25), where_to_look, 8× faster cold index by @ParsaVictor in #135

Full Changelog: v1.1.0...v1.2.0

What's Changed

  • Plain-language: trust the whole-question ranking; fix skeleton panic; baseline comparison script by @ParsaVictor in #128
  • Measured: outside baselines + fresh click holdout; packet: folds once, server guesses not 'missing' by @ParsaVictor in #129
  • Optional code-aware embeddings (jina-code v2): ripgrep plain-language 0.500 → 0.667 recall, 0.156 → 0.347 precision by @ParsaVictor in #130
  • Issue mode: long reports use the whole-question ranking (SWE-bench Lite dev hit@3 0.250 → 0.583) by @ParsaVictor in #131
  • Index speed: django cold index 3m20s → 26s, same graph by @ParsaVictor in #132
  • Disk: bound target/ (debug 43 GB, index cache 6.7 GB) by @ParsaVictor in #133
  • Issue mode: definition-level ranking + report hygiene (SWE dev Acc@5 0.456 → 0.596) by @ParsaVictor in #134
  • 1.2.0: issue localisation (SWE-bench Lite holdout Acc@5 0.587 → 0.750 vs BM25), where_to_look, 8× faster cold index by @ParsaVictor in #135

Full Changelog: v1.1.0...v1.2.0

What's Changed

  • Plain-language: trust the whole-question ranking; fix skeleton panic; baseline comparison script by @ParsaVictor in #128
  • Measured: outside baselines + fresh click holdout; packet: folds once, server guesses not 'missing' by @ParsaVictor in #129
  • Optional code-aware embeddings (jina-code v2): ripgrep plain-language 0.500 → 0.667 recall, 0.156 → 0.347 precision by @ParsaVictor in #130
  • Issue mode: long reports use the whole-question ranking (SWE-bench Lite dev hit@3 0.250 → 0.583) by @ParsaVictor in #131
  • Index speed: django cold index 3m20s → 26s, same graph by @ParsaVictor in #132
  • Disk: bound target/ (debug 43 GB, index cache 6.7 GB) by @ParsaVictor in #133
  • Issue mode: definition-level ranking + report hygiene (SWE dev Acc@5 0.456 → 0.596) by @ParsaVictor in #134
  • 1.2.0: issue localisation (SWE-bench Lite holdout Acc@5 0.587 → 0.750 vs BM25), where_to_look, 8× faster cold index by @ParsaVictor in #135

Full Changelog: v1.1.0...v1.2.0