workflow_job "queued" event fires for every matrix job at run start, even when max-parallel is holding them back.
#209687
Unanswered
andrewdibiasio6
asked this question in
Actions
Replies: 1 comment
|
💬 Your Product Feedback Has Been Submitted 🎉 Thank you for taking the time to share your insights with us! Your feedback is invaluable as we build a better GitHub experience for all our users. Here's what you can expect moving forward ⏩
Where to look to see what's shipping 👀
What you can do in the meantime 💻
As a member of the GitHub community, your participation is essential. While we can't promise that every suggestion will be implemented, we want to emphasize that your feedback is instrumental in guiding our decisions and priorities. Thank you once again for your contribution to making GitHub even better! We're grateful for your ongoing support and collaboration in shaping the future of our platform. ⭐ |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
🏷️ Discussion Type
Bug
💬 Feature/Topic Area
Workflow Configuration
Discussion Details
Summary
With
strategy.max-parallel: 1, GitHub creates all matrix jobs when the run starts and sends aworkflow_jobevent withaction: queuedfor each of them immediately, even though only one can run at a time. Autoscalers that launch a runner perqueuedevent (we use the terraform-aws-github-runner module) therefore launch a runner for every 'cell' up front. We haven't tested it, but this may affect other autoscalers that react toqueuedevents too, for example Actions Runner Controller in its webhook-driven mode. The runners for later cells sit idle while earlier cells run, and may be reaped by idle-timeout before their cell becomes eligible.Nothing in the event payload tells the autoscaler that the job is blocked by
max-parallelrather than waiting for runner capacity.What we observed (8 cells,
max-parallel: 1, each cell sleeps 6 minutes, ephemeral runners)created_at. 8queuedwebhook events were received within about 250 ms of each other.queued → startedper cell grows by about 6 minutes per cell (0 to 43 min), but the real runner wait after the previous cell finished was 0 to 0.5 min. The "queue time" is almost entirely earlier cells running.Contrast: matrix job that calls a reusable workflow
When the matrix job
uses:a reusable workflow instead of havingruns-on, the cells are not all created up front. Samemax-parallel: 1, same 8 cells, same runner pool:queuedevent, 1 runner launch. (The direct matrix had 8 jobs, 8 events and 8 launches at that point.)So the same
max-parallelproduces very differentqueuedtiming depending on whether the matrix job hasruns-onor calls a reusable workflow. We saw the same pattern in a separate scheduled workflow of ours that uses the reusable shape.Reproduction
Attached workflows
matrix-direct.ymlandmatrix-reusable.yml(plus_cell.yml). Run each withworkflow_dispatchon a self-hosted runner pool. To see the events, register a repository or App webhook forworkflow_jobpointed at any request inspector, or compare jobcreated_atwithGET /repos/{owner}/{repo}/actions/runs/{run_id}/jobs.matrix-direct.yml
matrix-reusable.yml
_cell.yml
Questions
queuedfor a matrix job thatmax-parallelis still holding back intended?waitingfor environment protection rules) for "created but not yet eligible to run", so autoscalers can avoid launching runners early?(Posted as a question first; happy to move to the roadmap as a feature request.)
All reactions