Batch Mode: Failure Scenarios, Recovery, and Operational Gotchas #20
Closed
AlexDeMichieli
started this conversation in
Ideas
Replies: 1 comment
|
Closing in favor of #17 |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Context
During E2E deployment testing of batch mode (
submit-repos.yml) on an EMU enterprise (volcano-coffee), I encountered multiple failure scenarios that required manual intervention. This discussion documents each failure mode, its root cause, recovery steps, and whether the proposed architecture in #14 would address it.Test environment: EMU enterprise, SAML SSO enforced, single org (
alexdemichieli-migrations), Copilot Enterprise license.Failure Scenarios
1.Sessions from Failed Assignments
Symptom:
github.com/copilot/agentsshows N "Active" sessions. New Copilot assignments fail with the same generic error "The agent encountered an error and was unable to start working on this issue: Please try again later, or contact support if the problem persists.".Recovery:
github.com/copilot/agents. The active count should dropWould #14 fix this? Yes, the agentic workflows wouldn't create per-repo issues, eliminating the session scenario.
2. Tag Clear Fails. Repos Re-Processed on Next Batch Run
Symptom: Batch workflow logs show
Failed to update custom property for <repo>: Resource not accessible by integration. TheGH_MIGRATION_TYPEtag stays set, so the repo is picked up again on the next batch run, creating duplicate issues.Root cause: This happened because I have incorrectly set the app's permissions. The GitHub App had
organization_custom_properties: write(for managing the org-level schema) but notcustom_propertiesat the repository level. The endpointPATCH /repos/{owner}/{repo}/properties/valuesrequires Repository: Custom properties (Read & Write).Recovery:
Prevention: Add this permission to the Phase 1 deployment checklist. The upstream
deployment.mddoesn't mention it.Would #14 fix this? Yes. The proposal uses a tracking file in
.github-privateinstead of custom properties.3. Duplicate Issues from Partial Batch Failures
Symptom: Multiple identical
[Actions Migration]issues on the same repo, each with the full agent prompt as the body.Root cause: The batch workflow has no idempotency check. If tag clear fails (Scenario 3) and the workflow is re-run, it finds the same tagged repos and creates new issues without checking for existing open migration issues.
Recovery:
Prevention: Add a pre-check to
submit-repositories.js: before creating an issue, search for open issues titled[Actions Migration]on the target repo. Skip if one exists.Would #14 fix this? Yes — eliminates the IssueOps pattern entirely as the Coding agent will start the process via API, without creating a new issue.
Related
#coding-agent-team— Platform bug tracking for SSO token mintingAll reactions