Skip to content

fix(inkless:retention): Limit the size of AWS delete error messages - #733

Merged
jeqo merged 1 commit into
mainfrom
delete-error-message
Aug 6, 2026
Merged

fix(inkless:retention): Limit the size of AWS delete error messages#733
jeqo merged 1 commit into
mainfrom
delete-error-message

Conversation

@gharris1727

Copy link
Copy Markdown
Contributor

The number of keys passed to the ObjectDeleter:delete method may be extremely large. If any key fails to delete, an error message including every key is assembled. This error message may be several megabytes in size, putting excessive memory pressure on the broker node.

Instead, the error message should be scoped to just the batch of keys that has just been deleted. This is limited by MAX_DELETE_KEYS_LIMIT (1000) which should keep the size of the message bounded.

This same failure mode appears in the other cloud implementations, those bugs/fixes are out of scope but will be addressed in an immediate follow-up. We will also investigate limiting the number of keys distributed to nodes by the control plane.

@gharris1727
gharris1727 requested review from ivanyu and jeqo August 6, 2026 19:43
@jeqo
jeqo merged commit f751436 into main Aug 6, 2026
@jeqo
jeqo deleted the delete-error-message branch August 6, 2026 19:51
jeqo added a commit that referenced this pull request Aug 11, 2026
Two error paths interpolated an unbounded value into the message, the same
class of bug as the S3 bulk-delete fix (#733):

- FetchCompleter: the catch-all wrapped the failure with the full
  fetchInfos.keySet(), which spans every client-requested partition
  (up to thousands). Report only the partition count.
- FileCommitter: logged the whole ClosedFile record, whose toString dumps
  every aggregated produce request and commit batch. Log a bounded summary
  (start, request count, batch count, bytes).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
jeqo added a commit that referenced this pull request Aug 11, 2026
A delete pass that fails for every key it submitted logged one line per key
(with a stack trace per key on Azure) and repeated it on every cleanup cycle,
since undeleted keys stay marked for deletion. GCS still interpolated the whole
key set into the exception message, the bug #733 fixed for S3.

Aggregate per-key failures into DeleteErrorSummary: a count by error code, at
most 3 sampled hard failures, and the first hard cause. One line per call,
bounded by the distinct error codes rather than the key count, at INFO when
every failure was a throttle and WARN otherwise. GCS cannot report per-key
results, so its message just carries the key count.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
jeqo added a commit that referenced this pull request Aug 11, 2026
Two error paths interpolated an unbounded value into the message, the same
class of bug as the S3 bulk-delete fix (#733):

- FetchCompleter: the catch-all wrapped the failure with the full
  fetchInfos.keySet(), which spans every client-requested partition
  (up to thousands). Report only the partition count.
- FileCommitter: logged the whole ClosedFile record, whose toString dumps
  every aggregated produce request and commit batch. Log a bounded summary
  (start, request count, batch count, bytes).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
jeqo added a commit that referenced this pull request Aug 12, 2026
Two error paths interpolated an unbounded value into the message, the same
class of bug as the S3 bulk-delete fix (#733):

- FetchCompleter: the catch-all wrapped the failure with the full
  fetchInfos.keySet(), which spans every client-requested partition
  (up to thousands). Report only the partition count.
- FileCommitter: logged the whole ClosedFile record, whose toString dumps
  every aggregated produce request and commit batch. Log a bounded summary
  (start, request count, batch count, bytes).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
jeqo added a commit that referenced this pull request Aug 12, 2026
A delete pass that fails for every key it submitted logged one line per key
(with a stack trace per key on Azure) and repeated it on every cleanup cycle,
since undeleted keys stay marked for deletion. GCS still interpolated the whole
key set into the exception message, the bug #733 fixed for S3.

Aggregate per-key failures into DeleteErrorSummary: a count by error code, at
most 3 sampled hard failures, and the first hard cause. One line per call,
bounded by the distinct error codes rather than the key count, at INFO when
every failure was a throttle and WARN otherwise. GCS cannot report per-key
results, so its message just carries the key count.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
jeqo added a commit that referenced this pull request Aug 13, 2026
…741)

Two error paths interpolated an unbounded value into the message, the same
class of bug as the S3 bulk-delete fix (#733):

- FetchCompleter: the catch-all wrapped the failure with the full
  fetchInfos.keySet(), which spans every client-requested partition
  (up to thousands). Report only the partition count.
- FileCommitter: logged the whole ClosedFile record, whose toString dumps
  every aggregated produce request and commit batch. Log a bounded summary
  (start, request count, batch count, bytes).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants