Added extensibility-request triage and implementation categories, including datasets, prompts, skills, runners, and LMChecklist evaluation.
Added AI red-team scanning for BCal with built-in or custom attack objectives and scorecard output.
Made code-review a dedicated category backed by the BC PR Review engine and BCQuality. Corrected collection to evaluate comments against the exact commit reviewed, added a neutral ignored-comments bucket to scoring, and expanded the gold set with vetted entries.
Moved tag pinning for multiple runs earlier in the workflow.
Pinned and persisted LLM judge models to prevent incompatible scores from being combined. Updated Copilot CLI, added MAI Code 1.1 Flash, and switched Copilot metrics collection to JSON output.
Versions
• GitHub Copilot CLI 1.0.80
• Claude Code 2.1.220
• Microsoft.Dynamics.BusinessCentral.Development.Tools 18.0.37.11445-beta