v0.7.2
HTTP retry + resilient endpoint fallthrough
Transient network blips on one of ~2500 packages no longer fail the build when a later endpoint successfully resolves the same name.
Changes
- New
_http.with_retryhelper — retriesURLError,TimeoutError,OSError, HTTP 5xx, HTTP 429. HonoursRetry-After(capped at 60s). HTTP 404, other 4xx, andJSONDecodeErrorare never retried. cas cooldown+cas audit: per-name error tracking across endpoint fallthrough. The build only fails on a transient error when the name errored on at least one endpoint and was never resolved on any other.- New
--retries Nflag (default2) on both commands. Shared env:CAS_RETRIES.
Why
Closes two distinct v0.7.1 CI failures observed in production:
- A single
URLErroronregistry.npmjs.orgfailing the cooldown gate. - A probe-registry error short-circuiting the audit CodeArtifact fallback (
Private-package probe failed for <pkg> (probe-registry)).
Both are inherent to large lockfiles (~2500 packages → 2500 HTTP calls) where a transient blip on any single request is effectively guaranteed.
Resilient fallthrough semantics
- Transient error on endpoint A → retry exponentially (
base_delay * 4^attempt). - Retries exhausted on A but B resolves the same
(name, version)→ error discarded, build passes. - Errored on at least one endpoint and never resolved on any other → surface as the build-failing finding (so genuine multi-endpoint outages still fail loudly).
225 tests pass (216 prior + 12 retry-helper + 9 resilient-fallthrough). Ruff + mypy strict clean.