Repository navigation
Releases: Altinity/s3gc
Release list
v0.7.0
A safety release. The headline is uncomfortable: the dangerous configuration was also the default one.
Changed
--useage now defaults to 24 hours and enforces it as a hard floor. It previously defaulted to 0, and if args.useage else "" emits no age predicate at all for 0 — so anyone who never set the flag ran with no age window.
That window is the only thing between a run and live data. ClickHouse uploads a part's blobs to S3 and registers them in system.remote_data_paths a moment later; in that gap a live blob is absent from the reference table and looks orphaned, and there is no per-object re-check before the S3 delete.
Below the floor a run is now refused outright, including --dry-run — a preview computed over a wider set than the delete would honour is worse than no preview, because the reviewed number is the one the customer approves. Enforced in both s3gc.py and render.py, so a direct CLI run is caught as well as a rendered Job. PHASE=dev-automation is the one exemption, since it seeds and deletes its own fixtures within minutes.
Operators who relied on the old default of 0 must now set USEAGE_HOURS explicitly, to 24 or higher. Raising it is the "be more careful" lever — a cluster with slow merges may want 72 hours or a week.
Added
A durable run log in ClickHouse, <COLLECTTABLEPREFIX><disk>_log, never truncated. Pod logs are not a record: the kubelet rotates container output and ttlSecondsAfterFinished deletes the Job with everything it printed, so a cleanup that reclaimed terabytes left no evidence once that window closed. Rows carry run_id, phase, event, running objects/bytes, and the scope the run was pointed at. A logging failure never fails the run, and a missing CREATE TABLE grant degrades to stdout only. Opt out with RUNLOG=false.
Collect progress at INFO, throttled to every 100k objects. Previously a multi-hour collect emitted about four lines at --verbose, while --debug emitted one line per object.
Fixed
- Cluster topology is re-checked before every sample, not once per run. The preflight was point-in-time while the anti-join loop can run for hours; a replica dropping out mid-run took its references with it, so blobs it alone held started looking orphaned.
--useafteris quoted as a SQL string literal. It was interpolated bare, landing as an identifier — the only unquoted value in the anti-joinWHEREclause.PYTHONUNBUFFERED=1in the image. stdout is a pipe under Kubernetes, so a Job killed atactiveDeadlineSecondslost its buffered tail.
Tests
The deletion scope is now asserted directly. Mutation testing established the gap rather than assuming it: against the previous suite, nine of eleven deliberate breakages of the delete scope passed fully green — including turning LEFT ANTI JOIN into a plain LEFT JOIN, and making --dry-run delete for real. All eleven are now caught.
Image
ghcr.io/altinity/s3gc:0.7.0
ghcr.io/altinity/s3gc@sha256:c6e90ebb2ae2e9a5ad6ab9356b614deb2fe0ad6d5b1c5f556f025e14d3c1872d
Multi-arch linux/amd64 + linux/arm64. Pin the digest in deployments — render.py rejects any image that is not digest-pinned.
v0.6.0
The first release published by CI, and the first image whose provenance can be recovered from the image itself.
Fixed
The delete phase died on its first batch with SESSION_IS_LOCKED (ClickHouse error 373), after the objects were already removed from S3.
connect_to_ch() built a single clickhouse_connect client, which the driver gives an auto-generated session_id, and ClickHouse permits one query at a time per session. do_use() holds that session for the entire anti-join while consuming query_row_block_stream, and insert() issues its own DESCRIBE TABLE before writing — a second concurrent query on the held session. The job exited non-zero with up to --deletebatchsize objects deleted and no tombstone recorded, so a resumed run could not tell they were done.
Tombstone writes now go to a second client built in the same call.
The defect was not new; it existed in every build back to the original single-client design and had simply never fired. Every earlier delete that reclaimed data ran with --order-by-objpath, which sorts the whole result server-side before streaming, and against ClickHouse 25.x. The first run without global ordering — the documented default for Kubernetes Jobs — hit it 78 minutes in, on the first block the anti-join produced.
Added
USETOTAL is now settable from the Kubernetes Job template. --usetotal already existed on the command line and, via env_prefix="S3GC", in the environment, but the renderer never emitted it — so the only delete available to an operator deploying with the renderer was unbounded.
Bound the first delete against a newly published image or an unfamiliar cluster to a few thousand objects: it exercises anti-join, S3 deletion and tombstone write-back end to end in minutes. The key is optional and renders no variable when empty, since S3GC_USETOTAL is parsed as an integer. Existing environment files render unchanged.
Release process
Two defects had to be fixed to publish anything at all, recorded because neither is visible from the code:
- The workflow triggers on
tags: ['v*.*.*'], while the repository's tags werev0.5,v_0.1andv_0.2. None can match, so the publish job had never run for a release. ghcr.io/altinity/s3gcalready existed from manual pushes. A GHCR package created by a user push is not linked to its repository, soGITHUB_TOKENwas refused withdenied: permission_denied: write_packageuntil the package's Manage Actions access granted the repository the Write role. A renamed or new package will need that grant again.
Image
ghcr.io/altinity/s3gc:0.6.0
ghcr.io/altinity/s3gc@sha256:9579513319ce35aee21b646a28f83273802b87b03f1f182af9c0566854f80273
Multi-arch linux/amd64 + linux/arm64, carrying org.opencontainers.image.revision. Pin the digest in deployments — render.py rejects any image that is not digest-pinned.
V0.5
s3gc v0.5
This release adds a safer Kubernetes Job workflow, flexible S3 authentication, and important correctness fixes for orphan
collection and deletion.
Highlights
-
Added Kubernetes Job phases for collect, dry-run, approved delete, and verify.
-
Added development-only dev-automation phase for end-to-end fixture testing.
-
Added S3 authentication modes:
- static credentials, including session tokens
- aws boto3 credential chain / AWS SSO profiles
- iam workload identity for IRSA, EC2 instance profiles, and ECS task roles
-
Published public multi-architecture images to ghcr.io/altinity/s3gc.
-
Added GCS compatibility: GCS endpoints automatically use per-object deletion because batch deletion is unsupported.
-
CI now runs offline tests, renders the Kubernetes Job, and validates it with strict Kubeconform schemas.
Fixes
- Fixed boolean environment parsing.
- Fixed --age for objects older than 24 hours.
- --usecollected now fails clearly when the auxiliary table is missing or empty.
- Added per-delete checkpointing and cumulative deletion totals.
- Added clearer S3 listing-permission errors and broader secret redaction.
Breaking change
S3GC_S3USEIAM / --s3useiam was removed. Use:
S3GC_S3AUTH=iam
Use S3AUTH=static|aws|iam in Kubernetes rendering configuration.
Operational notes
- Use a per-replica ClickHouse Service consistently across all phases.
- Always follow:
collect → dry-run → explicit approval → delete → verify
- Use digest-pinned images and a dedicated ClickHouse user with the documented minimum grants.