Skip to content

v0.7.2

Choose a tag to compare

@DragonStuff DragonStuff released this 13 May 01:20
· 7 commits to main since this release
v0.7.2

HTTP retry + resilient endpoint fallthrough

Transient network blips on one of ~2500 packages no longer fail the build when a later endpoint successfully resolves the same name.

Changes

  • New _http.with_retry helper — retries URLError, TimeoutError, OSError, HTTP 5xx, HTTP 429. Honours Retry-After (capped at 60s). HTTP 404, other 4xx, and JSONDecodeError are never retried.
  • cas cooldown + cas audit: per-name error tracking across endpoint fallthrough. The build only fails on a transient error when the name errored on at least one endpoint and was never resolved on any other.
  • New --retries N flag (default 2) on both commands. Shared env: CAS_RETRIES.

Why

Closes two distinct v0.7.1 CI failures observed in production:

  1. A single URLError on registry.npmjs.org failing the cooldown gate.
  2. A probe-registry error short-circuiting the audit CodeArtifact fallback (Private-package probe failed for <pkg> (probe-registry)).

Both are inherent to large lockfiles (~2500 packages → 2500 HTTP calls) where a transient blip on any single request is effectively guaranteed.

Resilient fallthrough semantics

  • Transient error on endpoint A → retry exponentially (base_delay * 4^attempt).
  • Retries exhausted on A but B resolves the same (name, version) → error discarded, build passes.
  • Errored on at least one endpoint and never resolved on any other → surface as the build-failing finding (so genuine multi-endpoint outages still fail loudly).

225 tests pass (216 prior + 12 retry-helper + 9 resilient-fallthrough). Ruff + mypy strict clean.