Skip to content

[evals] Daily Evals Feature Report - 2026-07-25 #47955

Description

@github-actions

Executive Summary

Across 82 evals-enabled workflows and 206 runs inspected in the last 7 full days, no run produced an evals.jsonl artifact. That leaves the evals feature with no usable scoring data and indicates the evals job is broken rather than just low-quality.

Caution

Status: BROKEN - the evals job is not producing results.

Key Metrics

Metric Value
Workflows with evals 82
Runs analyzed 206
Runs with evals results 0
Evals job success rate 0%
Overall YES rate 0%

Investigation Checklist

  1. Run gh aw audit 30150039650 on the most recent evals-declaring workflow run to inspect the evals job steps.
  2. Check whether the evals artifact is present; the representative run only exposed activation and agent artifacts.
  3. Look for errors in the evals engine execution step, especially Parse BinEval results.
  4. Verify that the evals frontmatter config is valid and that each question has both id and question.
  5. Compare max-ai-credits against actual usage to confirm the run is not budget-starved before evals executes.
  6. Check whether the evals/<workflow-id> branch exists and contains recent commits.

Affected Runs

References

Generated by 🧪 Daily Evals Feature Report · gpt54 · 25.9 AIC · ⌖ 2.68 AIC · ⊞ 10.3K ·

  • expires on Aug 1, 2026, 12:03 AM UTC-08:00

Metadata

Metadata

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions