Skip to content

v0.71.0

Choose a tag to compare

@steipete steipete released this 05 Oct 05:22
· 47 commits to main since this release
Immutable release. Only release title and notes can be modified.
v0.71.0
b4673cc

0.71.0 - 2026-10-04

Highlights

  • Retry cloud allocations with the same lease ID. RunPod, Scaleway, and Linode now support fixed idempotent lease IDs, and the provider catalog advertises support so orchestrators can check before creating a lease. PR 2680, PR 2684, PR 2678, PR 2674.
  • Get timely coordinator responses under load. Lease creates that cannot start committing admission within 30 seconds return a retryable 503; provider deadlines, compact polling, and bounded history reads reduce stalls and memory pressure. PR 2687, PR 2689, PR 2672, PR 2671, PR 2693, PR 2695, Issue 1561, Issue 2694. Thanks @shakkernerd.
  • Spend less time scanning local claims. Generated slugs avoid a full claim-directory scan, exact lease IDs use one-file lookups, ambiguous stop targets fail clearly, and resumable pruning removes eligible expired claims. PR 2685.
  • Choose the CPU and memory your workload needs. DigitalOcean, Scaleway, and Linode now map portable machine classes to distinct real sizes, with explicit native type overrides preserved. Review the higher-cost standard defaults below before upgrading. PR 2676.

Upgrade notes

  • The standard class now selects larger, higher-cost machines: DigitalOcean s-1vcpu-1gb → s-4vcpu-8gb, Scaleway DEV1-S → DEV1-L, and Linode g6-standard-1 → g6-standard-4. Keep the old size with --type <old type> or class: tiny; explicitly configured native types still take precedence. PR 2676.
  • Successful stop/release now triggers bounded daily pruning of expired E2B and sufficiently old Blacksmith Testbox local claims. Set claims.autoPrune: false or CRABBOX_CLAIMS_AUTO_PRUNE=false to retain them; recovery claims and provider resources remain protected. PR 2685.
  • Generated slugs now end in an eight-hex lease-ID suffix. Update scripts that assume the previous generated format; explicit --slug values retain their existing behavior. Ambiguous stop/release slugs now require an exact lease ID. PR 2685.

Added

  • Support fixed idempotent RunPod lease IDs with account-bound replay, lost-create recovery, and durable stop receipts for fixed-lease orchestrators. PR 2680.
  • Support fixed idempotent Scaleway lease IDs with replay, journaled SSH-key and root-volume recovery, and terminal release receipts for fixed-lease orchestrators. PR 2684.
  • Support fixed idempotent Linode lease IDs with account-bound replay, lost-create recovery, and terminal stop receipts. PR 2678.
  • Advertise backend-derived fixed-lease-id support in the provider catalog so fixed-lease orchestrators can preflight caller-supplied lease IDs. PR 2674.
  • Add claims prune and resumable daily maintenance after successful stop/release for providers with proven remote expiry bounds (E2B, and Blacksmith Testbox past GitHub's 35-day run cap); preserve recovery claims, concurrent updates, and provider resources, with claims.autoPrune / CRABBOX_CLAIMS_AUTO_PRUNE opt-out. PR 2685.
  • Let the guarded AWS Linux image publisher bake a smaller root snapshot from stock Ubuntu with linux_root_gb, keeping candidate, baseline, and promoted proofs on normal sizing. PR 2668.

Changed

  • Make --class select real machine sizes on DigitalOcean, Scaleway, and Linode, with CPU/RAM profiles and explicit type overrides preserved. The standard class mapping increases size and cost (DigitalOcean s-1vcpu-1gb → s-4vcpu-8gb, Scaleway DEV1-S → DEV1-L, Linode g6-standard-1 → g6-standard-4); keep the old size with --type <old type> or class: tiny. PR 2676.
  • Avoid the local claim-directory scan before coordinator create by giving generated slugs an eight-hex ID suffix, use O(1) exact canonical lease-ID claim lookups, and make AWS/coordinator acquisition, core claim routing, and Testbox ownership scans cancellable. PR 2685.
  • Page coordinator CLI lists with compact summaries, bound slug lookup and durable admission memory, keep identity/list reads out of the lifecycle queue, and arm cleanup recovery before provider I/O; preserve legacy list and full inspect responses. PR 2672.

Fixes

  • Bound runner synchronization history reads and pending writes while preserving complete stale responses, legacy identities, and retry behavior. PR 2695, Issue 2694. Thanks @shakkernerd.
  • Prevent queued coordinator uploads from retaining every run log in memory, page run-event reads at storage, and return structured failures for interrupted event appends instead of uncaught Worker errors. PR 2693.
  • Buffer POSIX sync manifest writes, time out rsync only after I/O inactivity, and retain workspace-witness stop requests when interrupted transfers are still settling. PR 2692.
  • Return a retryable 503 when coordinator lease creates cannot begin committing admission within 30 seconds instead of hanging. PR 2687; related Issue 1561.
  • Bound coordinator AWS, Azure, and Tailscale requests with retryable 503 deadlines, and verify Mac host ownership outside the coordinator lock, so stalled image or Mac host operations cannot indefinitely block queued lease creates. PR 2689.
  • Tolerate up to ±5 seconds of ASCII Box (Boat) create/read timestamp skew while preserving the original fixed-lease witness, and keep failed own creates inspectable and stoppable. PR 2681.
  • Report fleet, org, and owner lease headroom with the blocking cap in capacity, and identify fleet limits in allocation errors. PR 2679.
  • End interrupted coordinator provisioning waits on the first recovery tick, retain safe cleanup of uncertain cloud resources, and show the current attempt phase in lease diagnostics. PR 2677.
  • Refresh Tenki SSH gateway certificates before new connections so long-running commands retain workspace-owner renewal, collection, and cleanup access; report credential expiry when refresh fails. PR 2675.
  • Allow heartbeat and status --wait to renew owned direct fixed leases on DigitalOcean and Azure by validating the acquired account scope and exact resource identity. PR 2673.
  • Reject ambiguous local slugs before stop/release instead of selecting the first claim, and extend coordinator cancel-create recovery and final-attempt windows from 10 to 30 seconds each. PR 2685.
  • Bound coordinator lease-list memory as retained history grows, preserving visibility, filters, and result ordering. PR 2671, Issue 1561. Thanks @shakkernerd.
  • Retain Blacksmith Testbox claims and keys until the exact associated GitHub work settles, and report exact native queue and terminal status. PR 2670, Issue 2669. Thanks @shakkernerd.
  • Authenticate hosted Homebrew verifier metadata reads with the workflow's read-only token to avoid anonymous GitHub API rate limits while keeping installation credential-free. PR 2665.
  • Retry dropped connections in release asset downloads during public release and Homebrew verification instead of failing on one transient network error. PR 2667.
  • Expose the exact Blacksmith workflow association and verified remote settlement through read-only status, including after local claim finalization. PR 2683, Issue 2682. Thanks @shakkernerd.

Credits

  • Thanks @shakkernerd for PR 2650 (cold Actions hydration from GitHub SSH checkouts) and PR 2651 (AWS SSH access refresh over IPv4), which shipped in 0.70.0 without the thanks in its notes.