ci: purge the Cloudflare cache after each site deploy - #1189
Merged
Conversation
webjs.dev's static assets are served with max-age=14400 at stable urls, so after a deploy the edge keeps serving the previous copy for up to four hours. That shipped two visible regressions in one day (a pre-redesign tailwind.css after #1179, then the un-fixed logo marks after #1185), each cleared by a manual dashboard purge somebody had to remember. Staleness is also per-asset rather than all-or-nothing, so the site can sit half-updated with no signal. The workflow deliberately does not purge on push. Railway auto-deploys from the same push and the build takes minutes, so an immediate purge would evict the cache while the origin still served the old bytes and the next visitor would repopulate the edge with exactly those bytes, leaving the site stale for another four hours and burning the run for nothing. Instead it polls /__webjs/version (#239) on the Railway origin, bypassing the cache it is about to purge, until the reported uptime is lower than the elapsed time since the run began, which proves the restart belongs to THIS push rather than some earlier deploy. On timeout it fails loudly instead of purging. Purge is zone-wide: all four hostnames are proxied in the one webjs.dev zone, so a single call covers them and cannot silently miss an asset the way a hand-maintained path list does. curl uses --fail-with-body and the response is asserted on .success, since a bare curl exits 0 on a 4xx and would report a rejected purge as a success. workflow_dispatch is wired for on-demand purges, which is what the two incidents above actually needed. Closes #1188
Self-review of the new workflow turned up four real defects, all in the
wait step's failure handling.
1. Under `set -e`, `uptime=$(... | jq ...)` aborts the whole step when jq
fails, so a single non-JSON body from the origin (exactly what a
mid-deploy error page returns) turned a transient blip into a hard
failure instead of another retry. Every extraction now ends in
`|| true`, and the numeric comparison defaults to false.
2. The loop was 60 iterations of (10s curl + 15s sleep), a 25 minute
worst case against `timeout-minutes: 25`, so the job-level kill could
fire before the step printed why it gave up. The loop is now
deadline-driven on an explicit 15 minute WAIT_BUDGET, which holds
whatever the per-request latency turns out to be, leaving 10 minutes
of headroom for the diagnostic.
3. The verify step reported "Edge matches the origin" when EITHER side
sent no ETag, and Cloudflare strips ETags under some zone settings, so
the reassuring message was reachable with no comparison behind it. A
missing header is now reported as inconclusive.
4. The job declared no `permissions`, inheriting the default token scopes
despite never checking out the repo or calling the GitHub API. Now
`permissions: {}`.
Verified by running the real wait step against stub origins: the happy
path (uptime below elapsed) exits 0 after detecting the restart, and a
dead origin exits 1 with the refuse-to-purge guidance.
vivek7405
marked this pull request as ready for review
July 30, 2026 09:44
This was referenced Jul 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
.github/workflows/purge-cdn.yml, which evicts the Cloudflare edge cache after each site deploy, plus aworkflow_dispatchtrigger for on-demand purges.webjs.devserves its static assets withcache-control: public, max-age=14400at stable urls, so the edge keeps the previous copy for up to four hours after a deploy. That caused two visible regressions in one day (staletailwind.cssafter #1179, stale logo marks after #1185), each cleared by a manual dashboard purge.The workflow waits before purging, and that is the point. Railway auto-deploys from the same push and the build takes minutes. Purging immediately would evict the cache while the origin still serves the OLD bytes, the next visitor repopulates the edge with those bytes, and the site is stale for another four hours with nothing to show for the run. So the job polls
GET /__webjs/version(#239) on the Railway origin (bypassing the cache it is about to purge) until the reporteduptimeis lower than the elapsed time since the run began, which proves the restart belongs to this push. On timeout it fails loudly rather than purging.Purge is zone-wide: all four hostnames (
webjs.dev,example-blog,docs,ui) are proxied inside the onewebjs.devzone, so one call covers them all and cannot silently miss an asset the way a path list does.Action required before this is useful
Add a repo secret
CLOUDFLARE_API_TOKEN, scoped to Zone / Cache Purge / Purge on thewebjs.devzone only (not an account-wide token). Without it the purge step fails with an explicit message rather than passing silently.Test plan
runblock passesbash -nuptime=459s, elapsed 5s and 100s correctly evaluate "not yet deployed", a large elapsed evaluates "deployed"webjs.dev)mainNote on the permanent fix
tailwind.cssand the brand marks are referenced by hand-written urls inapp/layout.tsandlib/design/brand.ts, so the framework's content-hash?v=fingerprinting (#243) never touches them. Fingerprinting those urls would remove the need for a purge entirely. That is a separate, larger change; this workflow is worth having regardless as the safety net for anything on a stable url. Documented inframework-dev.md.Closes #1188