Added Data Query, an execution-based category answering how good an agent harness/model is at retrieving data from Business Central environment. It also enabled the capability to connect with a BC MCP server.
Changed NL2AL/BCal category: agents now get one attempt, empty diffs count as failures so we don't keep getting instabilities, and two safety-focused tasks moved from the accuracy dataset to the red-team dataset.
Reworked contamination measurement to a SWE-Bench Illusion-style top-1 exact file-path probe for bug-fix tasks and fixed parsing of valid fenced answers
Updated Code Review evaluation dataset, including added gold labels and expected comments.
Improved PR review agent's performance metrics, now including metrics like AI Credits.
Versions
• GitHub Copilot CLI 1.0.82
• Claude Code 2.1.221
• Microsoft.Dynamics.BusinessCentral.Development.Tools 18.0.37.11445-beta
• BC-ALAgents e9d9249eaa25be88388b60106b4cbf8cd7aa8a7d