Hardening squawk after Mini Shai-Hulud #264
neilcochran
announced in
Design notes
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
The incident itself - what got compromised, how the worm reached the registry, what got cleaned up - is documented in #251. This post is the follow-up: the hardening that close the exfil path, and the decisions behind each change.
The exfil path that ran
To make the rest of this concrete, here is the specific chain that ran on 2026-05-11:
@tanstack/router-cliand@tanstack/router-pluginto versions that had been compromised earlier the same day.main.publish.ymlonmainwithNPM_TOKENin scope across the whole job.npm ciresolved the compromised tarballs and executed theirpreparelifecycle scripts.preparescript readNPM_TOKENfrom the environment.Each hardening below narrows or eliminates a specific step in that chain (with one exception, called out where it applies).
Lifecycle scripts: off by default
The first hardening was the smallest:
--ignore-scriptson everynpm cicall across all four workflows (ci.yml,codeql.yml,docs.yml,publish.yml). This blocks step 4. A compromised dependency cannot run apreparescript during install if scripts do not run at install.The cost is that any dependency which genuinely needs an install-time build step stops working. An audit of the dependency graph at the time turned up exactly two scripts:
fsevents(macOS-only, already skipped on Linux runners) andunrs-resolver(a non-load-bearing post-install verifier). Both safe to skip. If a future dependency needs a real install script - a native module likebetter-sqlite3, for example - the install will fail loud and obvious, and the right answer is a per-step exception rather than a global re-enable.Applied to all four workflows, not just
publish.yml. The exfil ran frompublish.yml, but the same chain could have run fromci.ymlif the next attacker decided to target a different secret. Defense in depth at the workflow level is cheap.The token is no longer in scope when
npm cirunsThe second hardening was structural:
publish.ymlwas split into two jobs.The build job runs
npm ci --ignore-scriptsandnpm run build, then uploadspackages/libs/*/distas a workflow artifact. It hascontents: readand no secrets.The publish job downloads the artifact and hands off to
changesets/action. It does check out source and runnpm ci --ignore-scriptsfor the changesets tooling itself, but it does not rebuild any@squawk/*package - the tarballs that publish are the dist artifacts the build job produced. The workflow-level permissions block is empty (permissions: {}), so each job opts in to only what it specifically needs.This blocks step 3. Even if a malicious lifecycle script were to run in the build job (it cannot, because of
--ignore-scripts, but if it did), the publish credential would not be in the environment to steal.The cost is one upload-artifact + download-artifact round trip per release. Sub-minute overhead. Worth it.
No long-lived publishing credential exists
The third hardening eliminated the long-lived token entirely.
publish.ymlno longer carriesNPM_TOKENat all. Instead, the publish job uses npm's Trusted Publisher feature: theid-token: writepermission is granted on the job, npm CLI auto-detects the OIDC environment, and exchanges the GitHub OIDC token for a short-lived publishing credential that lives roughly 15 minutes.Per-package configuration on npm.com locks the trust to one specific workflow path -
neilcochran/squawkpluspublish.yml. A different repository, or a different workflow file in the same repository, cannot mint a publishing credential for these packages.As a belt-and-suspenders step, the "Publishing access" radio on each
@squawk/*package is also set to "Require two-factor authentication and disallow tokens." OIDC and Trusted Publisher continue to work normally under this setting; what it eliminates is the residual risk that a future maintainer mistake reintroduces a long-lived token. Even ifNPM_TOKENcame back somehow, it could not publish.This blocks step 3 differently from the job split - not just "the token is not in scope here" but "the token does not exist anywhere to be exfiltrated." If someone compromised the publish job today and pulled the OIDC token from memory, the worst they could do is publish for the next 15 minutes from a workflow that is already authorized to publish. The token cannot be saved for later, used from another machine, or used by another workflow.
The setup cost was 22 one-time clicks on npm.com (one Trusted Publisher entry per
@squawk/*package) plus a smallpublish.ymldiff to remove theNPM_TOKENline. The interim granular token issued during recovery has been revoked. TheNPM_TOKENrepo secret has been deleted.Every publish run pauses for human approval
The fourth hardening adds friction on purpose: every publish workflow run pauses between merge and publish for a one-tap manual approval.
The gate is a GitHub Environment named
production-publish, attached to the publish job. After CI succeeds onmainand the build job has uploaded its dist artifact, the publish job waits in a pending state. GitHub notifies the required reviewer; one tap on the approval link lets the run continue and the publish executes against the prebuilt artifact under OIDC. The build job is intentionally not gated - it runs without secrets, produces the dist artifact, and exits. Gating it would just waste a runner re-running the build later if approval lingered.The environment's
deployment_branch_policy.protected_branchesis set totrue, so aworkflow_dispatchfrom an unprotected branch cannot bypass the gate.Unlike the hardenings above, this one does not narrow a step in the original exfil chain - it adds a new step that did not exist before. The category it catches is a malicious PR that got merged anyway. Branch protection cannot save a maintainer from themselves; the original incident was a Dependabot PR I reviewed and approved. The approval gate gives a deliberate "wait, what am I actually shipping right now?" moment after the merge has happened, when the publish workflow shows exactly what is about to land on npm. One tap per release in exchange for one more place to notice something is wrong.
Actions cannot be tag-rewritten or come from arbitrary sources
Two settings on the Actions side narrow what code can run in a workflow at all:
The repository's Actions allowlist is set to "Allow neilcochran, and select non-neilcochran, actions and reusable workflows," with an explicit curated allowlist for non-GitHub actions:
actions/*(covered by the "Allow actions created by GitHub" checkbox),changesets/action@*, andlycheeverse/lychee-action@*. Anything outside the allowlist refuses to run."Require actions to be pinned to a full-length commit SHA" is enforced. A workflow using
uses: someorg/someaction@v3is rejected; onlyuses: someorg/someaction@<40-char-sha>is accepted. This blocks tag-rewrite attacks, where a previously-trusted version of an action gets re-pointed at a malicious commit.Every
uses:reference in the workflow files is a 40-character SHA followed by a# v<x.y.z>comment for readability. Dependabot keeps both in sync; manual edits to one without the other drift the comment from the SHA, which is caught in review.These do not address the Shai-Hulud exfil path specifically (the attack came in through an npm dependency, not an Action), but they close an adjacent path. A malicious Action can read repo secrets the same way a malicious npm script can, and the next attacker may go in through the Actions ecosystem instead.
The path from PR to
mainis gatedTwo GitHub Rulesets target
main:ciandCodeQL; deletion and force-push blocked; "require branches up to date" enabled so a stale PR cannot land.These do not stop a compromised dependency from being merged - the original incident involved a PR I reviewed and approved myself, and any branch protection regime that lets the maintainer merge anything is necessarily defeated by a maintainer mistake. They do stop an unreviewed force-push, a direct commit, or a merge with failing CI, and they make the merge audit trail explicit.
Account perimeter
npm 2FA enabled via passkey. All GitHub Personal Access Tokens revoked. No SSH or GPG keys on the GitHub account.
~/.npmrcabsent from the local machine. With OIDC in place, no long-lived publishing credential lives anywhere - not on the local machine, not in repo secrets, not in any CI environment between runs.Versions built under the hardened pipeline
The first publish through the new pipeline end-to-end was
@squawk/icao-registry-data@0.8.9on 2026-05-14: OIDC token exchange against npm, provenance signed and attested, publish run paused for approval before executing. Every subsequent@squawk/*publish follows the same path.The pre-incident clean versions listed in #251 predate the new pipeline but are bit-for-bit identical to what they were on 2026-05-11 before the worm ran - they were restored by GitHub Trust and Safety from the registry's pre-incident state, not republished.
All reactions