| PR Sous Chef |
4 |
100.0% |
25.0% |
"Did the agent add a comment to at least one pull request?" (25% YES) |
| AI Moderator |
3 |
100.0% |
0.0% |
"Does the agent output include a rationale explaining why the label(s) were applied or why noop was called?" (0% YES) |
| Artifacts Usage Report |
1 |
100.0% |
0.0% |
"Was a comprehensive summary report of artifacts usage produced?" (0% YES) |
| Auto-Triage Issues |
1 |
100.0% |
0.0% |
"Was a summary discussion created listing the issues processed and the labels applied?" (0% YES) |
| CLI Version Checker |
1 |
100.0% |
100.0% |
"Did the agent check for new versions of agentic CLI tools (Claude Code, GitHub Copilot CLI, Codex, MCP servers, etc.)?" (100% YES) |
| Code Scanning Fixer |
1 |
100.0% |
0.0% |
"Was a pull request created with a remediation for a code scanning alert, or was noop used when no fixable alerts existed?" (0% YES) |
| Code Simplifier |
1 |
100.0% |
100.0% |
"Was a pull request created with simplifications, or was noop used when no improvements were needed?" (100% YES) |
| Contribution Check |
1 |
100.0% |
0.0% |
"Was a report issue created summarizing PR compliance with the contributing guidelines?" (0% YES) |
| Copilot CLI Deep Research Agent |
1 |
100.0% |
0.0% |
"Does the agent output identify specific missed optimization opportunities for Copilot CLI usage in this repository?" (0% YES) |
| Daily Assign Issue To User |
1 |
100.0% |
0.0% |
"Did the agent post a comment explaining the assignment decision?" (0% YES) |
| Daily AstroStyleLite Markdown Spellcheck |
1 |
100.0% |
100.0% |
"Did the agent run American English spellcheck on AstroStyleLite docs content?" (100% YES) |
| Daily Cli Tools Tester |
1 |
100.0% |
100.0% |
"Did the agent run exploratory tests on the audit, logs, and compile CLI tools?" (100% YES) |
| Daily Community Attribution Updater |
1 |
100.0% |
100.0% |
"Did the agent scan community-labeled issues and attribute contributions using the five-tier strategy?" (100% YES) |
| Daily Compiler Quality Check |
1 |
100.0% |
0.0% |
"Was a discussion or report created with quality findings, or was noop used when all analyzed files met the quality standards?" (0% YES) |
| Daily Container Image Security Scan |
1 |
100.0% |
100.0% |
"Did the agent report actionable image findings, or use noop when no findings required action?" (100% YES) |
| Daily Go Test Parallelizer |
1 |
100.0% |
100.0% |
"Did the agent analyze Go tests to identify a safe candidate for t.Parallel?" (100% YES) |
| Daily Safe Outputs Conformance Checker |
1 |
100.0% |
0.0% |
"Were agentic tasks created for Copilot to address critical/high/medium/low issues, or was noop used when implementation was compliant?" (0% YES) |
| Daily Safe Outputs Git Simulator |
1 |
100.0% |
100.0% |
"Did the agent report the simulator results and any systematic safe-output issues it found?" (100% YES) |
| Daily Semgrep Scan |
1 |
100.0% |
0.0% |
"Did the agent complete a Semgrep security scan and report on the findings?" (0% YES) |
| Daily Skill Optimizer Improvements |
1 |
100.0% |
0.0% |
"Does the agent output include at least three specific skill improvement recommendations?" (0% YES) |
| Daily VulnHunter Scan |
1 |
100.0% |
0.0% |
"Did the agent download the prepared VulnHunter bundle artifact, load its vulnhunt skill instructions, and complete a repository scan?" (0% YES) |
| Daily Windows Terminal Integration Builder |
1 |
100.0% |
100.0% |
"Did the agent assess the Windows CLI integration build and test workflow?" (100% YES) |
| Daily Workflow Updater |
1 |
100.0% |
0.0% |
"Did the agent create a pull request for required updates, or report that no changes were needed?" (0% YES) |
| Dictation Prompt Generator |
1 |
100.0% |
0.0% |
"Did the agent create a pull request containing the generated dictation prompt update?" (0% YES) |
| Discussion Task Miner - Code Quality Improvement Agent |
1 |
100.0% |
0.0% |
"Does the agent output show that actionable tasks were identified from the analyzed discussions?" (0% YES) |
| Documentation Noob Tester |
1 |
100.0% |
0.0% |
"Does the agent output reflect the perspective of a new user rather than an expert reviewer?" (0% YES) |
| ESLint Refiner |
1 |
100.0% |
0.0% |
"Did the agent analyze ESLint diagnostics trends to identify rule refinement opportunities?" (0% YES) |
| GPL Dependency Cleaner (gpclean) |
1 |
100.0% |
0.0% |
"Does the agent output show that the objective for experiment tool_verbosity was successfully completed?" (0% YES) |
| Go Logger Enhancement |
1 |
100.0% |
0.0% |
"Does the agent output show that it ran validation commands to verify the logging changes compile correctly?" (0% YES) |
| Issue Arborist |
1 |
100.0% |
0.0% |
"Did the agent analyze recent issues and identify related issue relationships?" (0% YES) |
| Multi-Device Docs Tester |
1 |
100.0% |
100.0% |
"Did the agent report the multi-device test results and any responsive design or functionality findings?" (100% YES) |
| PR Triage Agent |
1 |
100.0% |
0.0% |
"Does the agent output confirm that category, risk, and action data were determined for each processed PR?" (0% YES) |
| Sub-Issue Closer |
1 |
100.0% |
100.0% |
"Did the agent check parent issues for the completion status of all their sub-issues?" (100% YES) |
| Tidy |
1 |
100.0% |
100.0% |
"Was a pull request created with formatting and tidying fixes, or was noop used when no changes were needed?" (100% YES) |
| [aw] Failure Investigator (6h) |
1 |
100.0% |
100.0% |
"Were fix sub-issues created for unresolved failures, or were resolved tracking issues closed?" (100% YES) |
Executive Summary
40 result-producing runs across 35 workflows were analyzed from the last 7 days. 26 runs passed every eval question, but the overall YES rate was 51.6%, so the feature is DEGRADED. There were 34 non-binary answers (
UNKNOWN) out of 97 total eval answers, which is a separate contract-quality signal.Note
Status: DEGRADED — evals jobs are producing results, but the overall YES rate is only 51.6%.
Key Metrics
Per-Workflow Pass Rates
Per-Question Breakdown per Workflow
AI Moderator
Artifacts Usage Report
Auto-Triage Issues
CLI Version Checker
Code Scanning Fixer
Code Simplifier
Contribution Check
Copilot CLI Deep Research Agent
Daily Assign Issue To User
Daily AstroStyleLite Markdown Spellcheck
Daily Cli Tools Tester
Daily Community Attribution Updater
Daily Compiler Quality Check
Daily Container Image Security Scan
Daily Go Test Parallelizer
Daily Safe Outputs Conformance Checker
Daily Safe Outputs Git Simulator
Daily Semgrep Scan
Daily Skill Optimizer Improvements
Daily VulnHunter Scan
Daily Windows Terminal Integration Builder
Daily Workflow Updater
Dictation Prompt Generator
Discussion Task Miner - Code Quality Improvement Agent
Documentation Noob Tester
ESLint Refiner
GPL Dependency Cleaner (gpclean)
Go Logger Enhancement
Issue Arborist
Multi-Device Docs Tester
PR Sous Chef
PR Triage Agent
Sub-Issue Closer
Tidy
[aw] Failure Investigator (6h)
Quality Signals
30733952293is at 100% NO across 3 run(s): PR Sous Chef30735700695is at 67% NO across 3 run(s): PR Sous Chef30737490610is at 67% NO across 3 run(s): PR Sous Chefpr-evaluatedin PR Sous Chef is the weakest repeated question, at 25% YES over 4 runs.Recommendations
UNKNOWNanswers. They are not binary eval outputs and should be reduced or explicitly handled in the evals contract.PR Sous Chefrubric aroundpr-evaluated; it is the only multi-run workflow with sustained partial failure.References