Repository navigation
Feedback: Efficiency Workflows MSFT adoption (ADO Specific) #58324
Paullyoung
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi team - got some good feedback from a MSFT engineer who adopted the 'Efficiency Improver' Agentic Workflow in ADO:
I ran the efficiency improver tool on our ADO repo. Here is my analysis of the tool so far:
Operational analysis:
The setup with ADO is little complex, specifically, with repositories in large ADO organizations. Few challenges that had to be work around:
The ADO-AW tool requires the service connection to read/write access to the project and pull requests for having enough context for recommendations. However, in many case, like ours, we will not have the access to add service connections with roles to the ADO organization or will be time-consuming. For now, we have configured the efficiency improver job to use the in-place memory system to keep track of the recommendations it provided and assuming the action taken based on existing code state in the
mainbranch - all checked out through only the available build identity context. PR history, comments, and other actions cannot be tracked without the service connection.The job should run as an ADO pipeline. However, in almost all ADO organizations (Microsoft internal) having direct Azure DevOps pipeline is non-compliant and raises S360s. The auto-generated pipeline YAML should cater to the compliant OneBranch pipeline format and template so that repository owners can configure them as compliant pipelines in the DevOps for continuous evaluation of the code.
I am yet to try this out, however, with ADO, the docs need to identify how we can set the GITHUB_COPILOT_TOKEN through Foundry over having Copilot PATs being configured right now.
Use case analysis: Output task: Task 6638847 [efficiency-improver] Monthly Activity 2026-08 (Opus 4.7), Task 6670716 [efficiency-improver] Monthly Activity 2026-09 (Opus 5)
Both the tasks provide recommendations on the Spark code improvements, but majority of these recommendations are related to minor enhancements or with smaller Dataframes. Example: It suggests avoiding shuffle operations or collect() calls repeatedly, however, most of this are applied on very small Dataframes (implying that the agent should have focused on data-sizing interpretations too).
So, in our previous attempts at improving the Spark code efficiency (before coming across this tool), some of these pointers were provided out of the box in the prompt that we wrote, example: "Please modify notebook X such that the given operation does not get Spark FetchFailedException due to its large shuffle size". In this case, models like Opus 4.8+ were used and provided significant improvements to notebook execution time, around ~60%.
One of our team members ran the efficiency improver markdown by passing it as a prompt locally to VS Code GitHub copilot and was able to raise following changes: Pull request 1592614: perf: reduce redundant Spark work in monthly pipeline notebooks - Repos. With this, while many improvements were still focusing on "generic" Spark improvements (independent of the scale of impact), few of the changes were actually meaningful (though still need to evaluate the impact of it).
I am guessing the markdown is generic "for any codebase" - probably we need to take following actions to use it beneficially:
Generate "AGENTS.md" which is repo-specific covering details around the repo code and some of the key use cases on how the code runs.
Probably ask Copilot to modify the efficiency improver markdown in a way that aligns it closer to the type of code base being looked at or the scale of different datasets being handled (as suggested 1.a.).
I am going to try both these approaches in coming week - but please let me know if you see any other changes that might also help us further.
Regards,
Chintan Rajvir
All reactions