Jobs wait ~60 s before the broker offers them to idle self-hosted runners, then go to several runners at once #209463
Replies: 1 comment
|
Thank you for your interest in contributing to our community! We currently only accept discussions created through the GitHub UI using our provided discussion templates. Please re-submit your discussion by navigating to the appropriate category and using the template provided. This discussion has been closed because it was not submitted through the expected format. If you believe this was a mistake, please reach out to the maintainers. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Why are you starting this discussion?
Bug
What GitHub Actions topic or product is this about?
Actions Runner
Discussion Details
Jobs sometimes wait about 60 seconds before GitHub offers them to any self-hosted runner, even though idle runners with matching labels are online and long-polling the broker the whole time. The job then starts by itself, so it is not the indefinitely-stuck case in #186811, but the wait is the same length every time and nothing on the runner side accounts for it.
What the runner
_diag/Runner_*.logfiles show in both traced cases:Acknowledging runner requestfor the job for about 60 s after the job'screated_at.RunnerRequestJobNotFoundException: Job not foundon/acknowledge, then get HTTP 409 onacquirejob. One runner wins and runs the job.The ~60 s gap matches the documented re-queue for an assigned job that is not picked up within 60 seconds, but none of our runners logged an earlier offer of the job, so if a first assignment happened, no runner received it.
Environment
self-hosted, Linux, X64, <lane label>).GET /orgs/{org}/actions/runnerslists only our runners, all online, so there are no stale registrations that could have taken the first assignment.runs-onwith that label set. The delayed jobs had no pendingneeds:dependency.Incidents (UTC)
Case 1, job ID 111313202003, runner request
ea0da1be-acb4-5d02-8246-12da9e1579b2created_at2026-10-03T23:06:00Z, noneeds:.Acknowledging runner request 'ea0da1be-…', thenJob not foundon/acknowledge, then 409 onacquirejob.started_at23:07:03Z.Case 2, job ID 111327411034, runner request
547006ec-a918-5c71-862d-71c784fb6d06needs:dependency completed at 2026-10-04T00:36:58Z. Jobcreated_at00:36:59Z.Job not foundand 409 onacquirejob, D wins. Jobstarted_at00:38:11Z.In other runs on the same runners the broker offers a job within 1–2 s of creation.
Ruled out
needs:dependencies: the run's jobs API timings show both jobs were ready atcreated_at.TaskCanceledException/Back offblock logged when a job finishes: it is the listener cancelling itsstatus=Busypoll to re-poll asOnline(as in Runner Fails to Connect to GitHub Actions Broker - TaskCanceledException / SocketException actions/runner#3904), and it does not line up with either gap.What we need
Confirmation of whether this ~60 s hold is expected broker behavior (for example, a first assignment to a runner session whose message is never delivered, followed by the 60 s re-queue), and whether anything on the runner or configuration side avoids it. The request IDs and timestamps above should locate both jobs in the broker logs.
All reactions