Repository navigation
A safety release. The headline is uncomfortable: the dangerous configuration was also the default one.
Changed
--useage now defaults to 24 hours and enforces it as a hard floor. It previously defaulted to 0, and if args.useage else "" emits no age predicate at all for 0 — so anyone who never set the flag ran with no age window.
That window is the only thing between a run and live data. ClickHouse uploads a part's blobs to S3 and registers them in system.remote_data_paths a moment later; in that gap a live blob is absent from the reference table and looks orphaned, and there is no per-object re-check before the S3 delete.
Below the floor a run is now refused outright, including --dry-run — a preview computed over a wider set than the delete would honour is worse than no preview, because the reviewed number is the one the customer approves. Enforced in both s3gc.py and render.py, so a direct CLI run is caught as well as a rendered Job. PHASE=dev-automation is the one exemption, since it seeds and deletes its own fixtures within minutes.
Operators who relied on the old default of 0 must now set USEAGE_HOURS explicitly, to 24 or higher. Raising it is the "be more careful" lever — a cluster with slow merges may want 72 hours or a week.
Added
A durable run log in ClickHouse, <COLLECTTABLEPREFIX><disk>_log, never truncated. Pod logs are not a record: the kubelet rotates container output and ttlSecondsAfterFinished deletes the Job with everything it printed, so a cleanup that reclaimed terabytes left no evidence once that window closed. Rows carry run_id, phase, event, running objects/bytes, and the scope the run was pointed at. A logging failure never fails the run, and a missing CREATE TABLE grant degrades to stdout only. Opt out with RUNLOG=false.
Collect progress at INFO, throttled to every 100k objects. Previously a multi-hour collect emitted about four lines at --verbose, while --debug emitted one line per object.
Fixed
- Cluster topology is re-checked before every sample, not once per run. The preflight was point-in-time while the anti-join loop can run for hours; a replica dropping out mid-run took its references with it, so blobs it alone held started looking orphaned.
--useafteris quoted as a SQL string literal. It was interpolated bare, landing as an identifier — the only unquoted value in the anti-joinWHEREclause.PYTHONUNBUFFERED=1in the image. stdout is a pipe under Kubernetes, so a Job killed atactiveDeadlineSecondslost its buffered tail.
Tests
The deletion scope is now asserted directly. Mutation testing established the gap rather than assuming it: against the previous suite, nine of eleven deliberate breakages of the delete scope passed fully green — including turning LEFT ANTI JOIN into a plain LEFT JOIN, and making --dry-run delete for real. All eleven are now caught.
Image
ghcr.io/altinity/s3gc:0.7.0
ghcr.io/altinity/s3gc@sha256:c6e90ebb2ae2e9a5ad6ab9356b614deb2fe0ad6d5b1c5f556f025e14d3c1872d
Multi-arch linux/amd64 + linux/arm64. Pin the digest in deployments — render.py rejects any image that is not digest-pinned.