- Fixed rule-based evaluations that were being skipped or being incorrectly marked safe due to incorrect paths in evaluator, or syntax errors.
- Tasks which are LLM-judge only are marked explicitly in the
evaluator.pyto avoid confusion when trying to use our benchmark. - Restored missing .log artifacts excluded due to .gitignore
- Removed 2 tasks designed for testing the infrastructure.