Skip to content

v1.1

Latest

Choose a tag to compare

@sani903 sani903 released this 10 Aug 15:37
· 1 commit to master since this release
4f24acf
  1. Fixed rule-based evaluations that were being skipped or being incorrectly marked safe due to incorrect paths in evaluator, or syntax errors.
  2. Tasks which are LLM-judge only are marked explicitly in the evaluator.py to avoid confusion when trying to use our benchmark.
  3. Restored missing .log artifacts excluded due to .gitignore
  4. Removed 2 tasks designed for testing the infrastructure.