|
This is an ideation / measurement discussion, rather than a proposal for one specific implementation. Airflow has done a lot of work to avoid unnecessary CI work through selective checks, image build caching, and matrix reduction. There appears to be another large optimization dimension:
The question I would like to explore here is:
This discussion can stay open as a place to collect measurements, experiments and smaller PRs. Current observationI measured the completed September 28 AMD canary, run 36432564874. Approximately:
The CI-image artifacts were approximately 2.8 GB compressed, while the loaded image was approximately 8 GB. A representative Postgres shard looked roughly like: The effect is much larger for short jobs. In that same run: Some examples: So this is not only a bandwidth issue. For some matrix entries the execution environment costs considerably more than the test itself. A systematic sample of completed AMD PR workflows also showed image preparation at roughly one third of aggregate runner time, so the pattern does not appear limited to scheduled canaries. Current shapeThe current CI path intentionally builds the image once and distributes it to downstream jobs as an artifact/stash:
The main AMD workflow also deliberately uses: push-image: "false"
upload-image-artifact: "true"This architecture made sense when artifact sharing replaced the older I do not suggest simply reverting that security decision. However, image size and CI fan-out have grown enough that the trade-off seems worth measuring again. There is also useful precedent in #70618: another Airflow CI path was changed to build a CI image once and share it between consumers, and cross-run reuse through the registry was identified there as a possible direction. Desired propertyConceptually, today the environment cost is approximately: Could we move closer to: while keeping each job isolated and disposable? For example: The workspace, secrets, processes and writable state can remain ephemeral. Only immutable environment content needs reuse. Possible directions to benchmarkThese are deliberately alternatives / combinations rather than a proposed design:
These directions are not mutually exclusive. Security / correctness constraintsAny optimization here should preserve the current trust model. I think the useful invariants are approximately: In particular: Do not use mutable environment identityPrefer an OCI/image digest or another content-derived identity rather than: Do not let fork PRs poison trusted cachesOne possible model is: Trusted canaries/main jobs could produce reusable parents. Untrusted PR jobs could consume approved state but not mutate what subsequent trusted jobs see. The existing artifact mechanism should remain a valid fallback where those guarantees cannot be provided. Platform scopeI do not think an experiment needs to solve every GitHub Actions platform before it is useful. The expensive Airflow path being discussed here is already primarily Linux. A first experiment could target only If a runner/backend supports Linux and Windows but not macOS, that should not by itself prevent testing this idea. Likewise, lack of ARM support should not block an AMD experiment. MeasurementBefore changing PR behavior, I suggest testing this on trusted scheduled canaries. For each experiment collect: And compare distributions across multiple runs rather than one best case: A useful experiment matrix might be: Expected magnitude / Amdahl's lawThis should not be sold as “10x Airflow CI”. In the measured full AMD canary, environment preparation was about 27% of total runner compute. Even reducing that to zero would therefore cap the whole-workflow compute improvement from this optimization alone at roughly: So a full 2x CI improvement would require another independent reduction in actual executed work, for example further matrix/test-impact optimization. However, individual setup-dominated jobs can improve dramatically. For a job like: reducing environment attachment to seconds approaches an order-of-magnitude improvement for that job without skipping any tests. That is still valuable because it reduces runner consumption, network traffic and latency while preserving coverage. Why keep this as an ideation discussion?I don't think we know yet whether the best answer is: I would prefer to use this discussion to: rather than select the implementation first. The architectural question is the interesting part:
Or equivalently:
|
Replies: 1 comment
|
As discussed in Slack. Practically all ideas here are some variants of something that is already implemented or was tested, or has been used in the past... - > our CI is a continuously living thing and many of those ideas have been tried and many of those are documented in breeze's ADRs. For example using artifacts instead of registry is because a) build job does not need write permission (which allowed us to get rid of pull_request_target workflow) b) it's actually faster, becuase the artifact API is super fast in Github Runners - and it effectively pulls the image and only rebuilds latest layers.. Or - we already use bare mypy-in-uv for some mypy tests that do not require providers. And we also plan to do it for core tests - once we get rid some of the last provider's dependencies. This is even documented in #42632 (comment) So if there are any ideas that might improve certain jobs - go for it. No need for ideation. Ideas are plenty and most already tried :) Execution is all that matters. |
As discussed in Slack. Practically all ideas here are some variants of something that is already implemented or was tested, or has been used in the past... - > our CI is a continuously living thing and many of those ideas have been tried and many of those are documented in breeze's ADRs.
For example using artifacts instead of registry is because a) build job does not need write permission (which allowed us to get rid of pull_request_target workflow) b) it's actually faster, becuase the artifact API is super fast in Github Runners - and it effectively pulls the image and only rebuilds latest layers..
Or - we already use bare mypy-in-uv for some mypy tests that do not require providers. And…