fix(inkless:retention): Limit the size of AWS delete error messages - #733
Merged
Conversation
jeqo
approved these changes
Aug 6, 2026
jeqo
added a commit
that referenced
this pull request
Aug 11, 2026
Two error paths interpolated an unbounded value into the message, the same class of bug as the S3 bulk-delete fix (#733): - FetchCompleter: the catch-all wrapped the failure with the full fetchInfos.keySet(), which spans every client-requested partition (up to thousands). Report only the partition count. - FileCommitter: logged the whole ClosedFile record, whose toString dumps every aggregated produce request and commit batch. Log a bounded summary (start, request count, batch count, bytes). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
jeqo
added a commit
that referenced
this pull request
Aug 11, 2026
A delete pass that fails for every key it submitted logged one line per key (with a stack trace per key on Azure) and repeated it on every cleanup cycle, since undeleted keys stay marked for deletion. GCS still interpolated the whole key set into the exception message, the bug #733 fixed for S3. Aggregate per-key failures into DeleteErrorSummary: a count by error code, at most 3 sampled hard failures, and the first hard cause. One line per call, bounded by the distinct error codes rather than the key count, at INFO when every failure was a throttle and WARN otherwise. GCS cannot report per-key results, so its message just carries the key count. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
jeqo
added a commit
that referenced
this pull request
Aug 11, 2026
Two error paths interpolated an unbounded value into the message, the same class of bug as the S3 bulk-delete fix (#733): - FetchCompleter: the catch-all wrapped the failure with the full fetchInfos.keySet(), which spans every client-requested partition (up to thousands). Report only the partition count. - FileCommitter: logged the whole ClosedFile record, whose toString dumps every aggregated produce request and commit batch. Log a bounded summary (start, request count, batch count, bytes). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
jeqo
added a commit
that referenced
this pull request
Aug 12, 2026
Two error paths interpolated an unbounded value into the message, the same class of bug as the S3 bulk-delete fix (#733): - FetchCompleter: the catch-all wrapped the failure with the full fetchInfos.keySet(), which spans every client-requested partition (up to thousands). Report only the partition count. - FileCommitter: logged the whole ClosedFile record, whose toString dumps every aggregated produce request and commit batch. Log a bounded summary (start, request count, batch count, bytes). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
jeqo
added a commit
that referenced
this pull request
Aug 12, 2026
A delete pass that fails for every key it submitted logged one line per key (with a stack trace per key on Azure) and repeated it on every cleanup cycle, since undeleted keys stay marked for deletion. GCS still interpolated the whole key set into the exception message, the bug #733 fixed for S3. Aggregate per-key failures into DeleteErrorSummary: a count by error code, at most 3 sampled hard failures, and the first hard cause. One line per call, bounded by the distinct error codes rather than the key count, at INFO when every failure was a throttle and WARN otherwise. GCS cannot report per-key results, so its message just carries the key count. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
jeqo
added a commit
that referenced
this pull request
Aug 13, 2026
…741) Two error paths interpolated an unbounded value into the message, the same class of bug as the S3 bulk-delete fix (#733): - FetchCompleter: the catch-all wrapped the failure with the full fetchInfos.keySet(), which spans every client-requested partition (up to thousands). Report only the partition count. - FileCommitter: logged the whole ClosedFile record, whose toString dumps every aggregated produce request and commit batch. Log a bounded summary (start, request count, batch count, bytes). Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The number of keys passed to the ObjectDeleter:delete method may be extremely large. If any key fails to delete, an error message including every key is assembled. This error message may be several megabytes in size, putting excessive memory pressure on the broker node.
Instead, the error message should be scoped to just the batch of keys that has just been deleted. This is limited by MAX_DELETE_KEYS_LIMIT (1000) which should keep the size of the message bounded.
This same failure mode appears in the other cloud implementations, those bugs/fixes are out of scope but will be addressed in an immediate follow-up. We will also investigate limiting the number of keys distributed to nodes by the control plane.