Skip to content

Throughput and Capacity

Gabor Galazzo edited this page Jul 25, 2026 · 1 revision

Throughput and Capacity

Behaviour introduced in 2.4.0. Public guide: docs.eventoframework.com → Throughput and Capacity.

This page exists because exceeding server capacity does not degrade gracefully — it collapses. Understanding why takes five minutes and saves an incident.


1. The incident that produced this page

A bundle running a bulk catalogue import, 20 concurrent writers against an 8-core server:

Throughput fell from 16 writes/s to 0.07 writes/s
Errors ~250 local_handler_threw … TimeoutException per hour
CPU under 1%
Databases idle
int_lock empty
Blocked queries none
A single command sent by hand returned in 0.1 s

Nothing was wedged. Re-running the same job at 8 concurrent writers sustained 25 writes/s with zero timeouts.

Less concurrency, more throughput. That inversion is the signature of this failure mode.


2. Root cause: a pool that never grew

busBusinessExecutor was a stock ThreadPoolExecutor(core = cores×2, max = cores×8, ArrayBlockingQueue(1024), CallerRunsPolicy).

A ThreadPoolExecutor starts a thread beyond core only when the queue is full. With a 1024-deep queue the pool never left 16 threads — max = 64 was unreachable under precisely the load it had been configured for, and CallerRunsPolicy never engaged either.

Then the second half. Every request carries a client deadline (30 s by default) and nothing cancels the work when it expires. So the queue filled with requests whose callers had already given up, and the server stayed fully busy producing responses nobody would read.

That is congestion collapse. Retries sustain it, which is why it does not self-heal when the triggering load stops.


3. What changed

  • BusBusinessExecutor grows before it queues. Its GrowthFirstQueue (Tomcat's approach) refuses the offer while the pool can still grow, so the pool reaches max before anything queues. submittedCount is tracked in execute / afterExecute to avoid taking getActiveCount()'s pool lock on every submission.
  • Queue default 1024 → 256. Deeper than deadline ÷ service time is dead work by construction.
  • Typed exceptions — RequestTimeoutException and TooManyPendingRequestsException in com.evento.transport, replacing the generic IllegalStateException.
  • BundleClientConfig.Builder.maxInFlightRequests(n) (default 2048, 0 disables) — refuse rather than queue work that will expire unsent.
  • Meters evento.server.bus.executor.{pool.size,max,active,queue.depth,saturated}.
  • Expiries moved INFO → WARN, and the bundle-side line gained byType={...}.

4. The three limits in series

A request is capped by the smallest of:

Limit Default Notes
evento.server.bus.business-executor-max-size cores × 8 Now actually reachable
evento.server.bus.business-executor-queue-capacity 256 The buffer behind a fully-grown pool
evento.es.fetch.concurrency 4 A fair semaphore — no number of bus threads raises it

evento.es.fetch.concurrency is usually the binding constraint on a consumer-heavy cluster. It is deliberately still 4: it guards heap against concurrent EventFetchRequest result sets, and the OOM it was added for is real. Raise it against measured headroom, not on principle.


5. Why enlarging the queue makes it worse

The queue is the buffer behind a fully-grown pool, not a shock absorber. Anything deeper than deadline ÷ service time is dead work by construction — those requests will expire before they are served, and serving them burns the exact capacity the live requests need.

Enlarging it to absorb load deepens the collapse. This is the most common instinctive wrong move.


6. Detection playbook

Signal Where Meaning
evento.server.bus.executor.saturated /actuator/prometheus The alerting signal. Only moves when the pool is at max and the queue is full. Alert on any sustained increase
evento.server.bus.executor.queue.depth same The earlier warning
event=bus_business_executor_saturated server log, WARN Rate-limited by business-executor-saturation-warn-interval
event=bundle_correlation_expired … byType={CatalogProductAddCommand=12} bundle log, WARN Names the payload type that is drowning. Read as capacity exhaustion, not unlucky individual requests

The tell: idle CPU and idle databases alongside collapsing throughput. If everything looks under-utilised and nothing is completing, you are here — not in a resource shortage.


7. Recovery

  1. Reduce client concurrency. This is what actually breaks the cycle. Fewer concurrent clients frequently complete strictly more work.
  2. Stop the retry storm — retries are what keep the collapse self-sustaining.
  3. Then size the server against the arithmetic above.

8. Handling a timeout correctly in application code

RequestTimeoutException           → the caller stopped waiting. The work MAY HAVE BEEN APPLIED.
                                    Report 504 (indeterminate), not 500. Retry only if idempotent.

TooManyPendingRequestsException   → nothing was transmitted. Definite refusal. Always retryable.

In the incident, roughly 20 commands reported as failures had in fact been fully applied. Treating a timeout as a definite failure and retrying it would have double-applied every one of them.


9. Bounding a bundle from the client side

BundleClientConfig.Builder.maxInFlightRequests(n)   // default 2048; 0 disables

Refusing at the client is strictly better than queueing work that will expire unsent: the refusal is immediate, definite, and safe to retry.

The consume path has its own separate bound — see Parallel Consumers.


See also

Clone this wiki locally