Skip to content

Flaky test: S3StorageClientSpec intermittently fails 4 tests in CI (shared MinIO contention) #7049

Description

@Yicong-Huang

Summary

S3StorageClientSpec intermittently fails 4 tests in the build / amber CI job (step Run amber and common module tests with coverage), scope WorkflowCore / Test / test. It is flaky: the same commit passes on re-run and on other PRs.

Observed:

[info] Tests: succeeded 688, failed 4, canceled 0, ignored 0, pending 1
[error] *** 4 TESTS FAILED ***
[error] Failed tests:
[error] 	org.apache.texera.service.util.S3StorageClientSpec
[error] (WorkflowCore / Test / test) sbt.TestsFailedException: Tests unsuccessful

Example failing run: https://github.com/apache/texera/actions/runs/30486266380/job/90692899618 (PR #7022, a CI-only change that touches no Scala — so the failure cannot be caused by the diff).

The per-test assertion messages are not in the console log (ScalaTest writes them only to target/test-reports/*.xml, which is uploaded to Codecov but not kept as a downloadable CI artifact), so the exact 4 failing test names are not recoverable from the run log.

When it started

The S3 test coverage was substantially expanded in #6987 (merged 2026-07-29, b291f75), which added heavier/new cases to S3StorageClientSpec — including a concurrent 1001-object upload (common/workflow-core/src/test/scala/org/apache/texera/service/util/S3StorageClientSpec.scala:433) and two ~5 MiB multipart cases (:622, :690). The flake surfaced after this.

Root cause (hypothesis)

Contention on the single shared MinIO container and the single shared S3StorageClient.s3Client (one connection pool) under concurrent load, intermittently exceeding the AWS SDK's default socket/connection timeouts on a resource-constrained runner:

  1. Four suites mix S3StorageTestBase and share one MinIO container + one s3Client: S3StorageClientSpec, LargeBinaryManagerSpec, LargeBinaryInputStreamSpec, LargeBinaryOutputStreamSpec (common/workflow-core/src/test/scala/org/apache/texera/service/util/S3StorageTestBase.scala:29).
  2. ScalaTest runs suites in parallel via its own Distributor thread pool (CI log shows pool-24-thread-N-ScalaTest-running-...Spec). The Global / concurrentRestrictions += Tags.limit(Tags.Test, 1) in common/workflow-core/build.sbt:35 only limits sbt task concurrency; it does not stop ScalaTest from running suites concurrently within a single unforked test task.
  3. On top of that, S3StorageClientSpec.scala:439 fires 1001 uploads through a 16-thread pool at the same shared container.

The combined burst overwhelms MinIO / starves the shared client's connection pool → some operations time out → the owning Future/assert fails (and can cascade to neighbouring sequential tests in the suite while the container recovers). When the runner has enough headroom the whole suite passes, hence the flakiness.

Suggested fix directions

  • Stop the MinIO-backed suites from running concurrently (they share one container + one client), e.g. serialize just these suites, and/or
  • Reduce the concurrent-upload burst in the >1000-objects test (:439) to a modest bounded pool.
  • Optionally, upload target/test-reports/*.xml as a CI artifact so future flakes surface the exact failing test + assertion.

Environment

  • Job: build / amber (ubuntu-latest, 17); JDK 17 (Temurin 17.0.19); sbt; MinIO testcontainer minio/minio:RELEASE.2025-02-28T09-55-16Z.

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions