-
Notifications
You must be signed in to change notification settings - Fork 0
Throughput and Capacity
Behaviour introduced in 2.4.0. Public guide: docs.eventoframework.com → Throughput and Capacity.
This page exists because exceeding server capacity does not degrade gracefully — it collapses. Understanding why takes five minutes and saves an incident.
A bundle running a bulk catalogue import, 20 concurrent writers against an 8-core server:
| Throughput | fell from 16 writes/s to 0.07 writes/s |
| Errors | ~250 local_handler_threw … TimeoutException per hour |
| CPU | under 1% |
| Databases | idle |
int_lock |
empty |
| Blocked queries | none |
| A single command sent by hand | returned in 0.1 s |
Nothing was wedged. Re-running the same job at 8 concurrent writers sustained 25 writes/s with zero timeouts.
Less concurrency, more throughput. That inversion is the signature of this failure mode.
busBusinessExecutor was a stock
ThreadPoolExecutor(core = cores×2, max = cores×8, ArrayBlockingQueue(1024), CallerRunsPolicy).
A ThreadPoolExecutor starts a thread beyond core only when the queue is full. With a
1024-deep queue the pool never left 16 threads — max = 64 was unreachable under precisely the load
it had been configured for, and CallerRunsPolicy never engaged either.
Then the second half. Every request carries a client deadline (30 s by default) and nothing cancels the work when it expires. So the queue filled with requests whose callers had already given up, and the server stayed fully busy producing responses nobody would read.
That is congestion collapse. Retries sustain it, which is why it does not self-heal when the triggering load stops.
-
BusBusinessExecutorgrows before it queues. ItsGrowthFirstQueue(Tomcat's approach) refuses the offer while the pool can still grow, so the pool reachesmaxbefore anything queues.submittedCountis tracked inexecute/afterExecuteto avoid takinggetActiveCount()'s pool lock on every submission. -
Queue default
1024 → 256. Deeper thandeadline ÷ service timeis dead work by construction. -
Typed exceptions —
RequestTimeoutExceptionandTooManyPendingRequestsExceptionincom.evento.transport, replacing the genericIllegalStateException. -
BundleClientConfig.Builder.maxInFlightRequests(n)(default 2048,0disables) — refuse rather than queue work that will expire unsent. -
Meters
evento.server.bus.executor.{pool.size,max,active,queue.depth,saturated}. -
Expiries moved
INFO → WARN, and the bundle-side line gainedbyType={...}.
A request is capped by the smallest of:
| Limit | Default | Notes |
|---|---|---|
evento.server.bus.business-executor-max-size |
cores × 8 |
Now actually reachable |
evento.server.bus.business-executor-queue-capacity |
256 |
The buffer behind a fully-grown pool |
evento.es.fetch.concurrency |
4 |
A fair semaphore — no number of bus threads raises it |
evento.es.fetch.concurrency is usually the binding constraint on a consumer-heavy cluster. It is
deliberately still 4: it guards heap against concurrent EventFetchRequest result sets, and the OOM
it was added for is real. Raise it against measured headroom, not on principle.
The queue is the buffer behind a fully-grown pool, not a shock absorber. Anything deeper than
deadline ÷ service time is dead work by construction — those requests will expire before they are
served, and serving them burns the exact capacity the live requests need.
Enlarging it to absorb load deepens the collapse. This is the most common instinctive wrong move.
| Signal | Where | Meaning |
|---|---|---|
evento.server.bus.executor.saturated |
/actuator/prometheus |
The alerting signal. Only moves when the pool is at max and the queue is full. Alert on any sustained increase |
evento.server.bus.executor.queue.depth |
same | The earlier warning |
event=bus_business_executor_saturated |
server log, WARN
|
Rate-limited by business-executor-saturation-warn-interval
|
event=bundle_correlation_expired … byType={CatalogProductAddCommand=12} |
bundle log, WARN
|
Names the payload type that is drowning. Read as capacity exhaustion, not unlucky individual requests |
The tell: idle CPU and idle databases alongside collapsing throughput. If everything looks under-utilised and nothing is completing, you are here — not in a resource shortage.
- Reduce client concurrency. This is what actually breaks the cycle. Fewer concurrent clients frequently complete strictly more work.
- Stop the retry storm — retries are what keep the collapse self-sustaining.
- Then size the server against the arithmetic above.
RequestTimeoutException → the caller stopped waiting. The work MAY HAVE BEEN APPLIED.
Report 504 (indeterminate), not 500. Retry only if idempotent.
TooManyPendingRequestsException → nothing was transmitted. Definite refusal. Always retryable.
In the incident, roughly 20 commands reported as failures had in fact been fully applied. Treating a timeout as a definite failure and retrying it would have double-applied every one of them.
BundleClientConfig.Builder.maxInFlightRequests(n) // default 2048; 0 disablesRefusing at the client is strictly better than queueing work that will expire unsent: the refusal is immediate, definite, and safe to retry.
The consume path has its own separate bound — see Parallel Consumers.
- Server Configuration § 2 — the properties named here
- Observability — full meter list
- Bundle Client § 3 — typed request outcomes
- Troubleshooting
Evento Framework — Copyright 2020–2026 © Gabor Galazzo. Dual-licensed under AGPL-3.0 and a commercial licence.
This wiki documents the implementation; the repository is authoritative where the two disagree. Found something out of date? Open an issue.
Getting oriented
Internals
Operations
- Server Configuration
- Throughput and Capacity
- Observability
- Security Model
- Server REST API
- Troubleshooting
Project