Skip to content

Troubleshooting

Gabor Galazzo edited this page Jul 25, 2026 · 1 revision

Troubleshooting

A symptom-first index of the failure modes that are documented in the codebase. Each entry links to the page with the full mechanism.


Throughput collapses while CPU and databases sit idle

Symptom. Throughput falls by orders of magnitude — the reference incident went from 16 writes/s to 0.07 writes/s. CPU under 1%, databases idle, no blocked queries, no lock contention. A single request sent by hand returns in 0.1 s. Timeouts everywhere. Reducing client concurrency increases throughput.

Cause. Congestion collapse. Requests carry a client deadline and nothing cancels the work when it expires, so the server spends its capacity producing responses nobody is waiting for. Retries keep it there, so it does not self-heal when the triggering load stops.

Check. evento.server.bus.executor.saturated (rising), …queue.depth, and WARN event=bus_business_executor_saturated.

Fix. Reduce client concurrency first, stop the retry storm, then size the server. Do not enlarge the queue — that deepens the collapse.

→ Throughput and Capacity


A bundle is connected and consuming, but the dashboard shows none of its handlers

Symptom. The bundle registers, enables and processes events normally, yet its handlers, catalog entries and flows are missing from the GUI. Nearly invisible — nothing fails.

Cause. The rich evento:bundle-discovery notification was rejected at decode, killing the entire BundleDiscoveryInfo. Historically caused by a peer emitting a property the other side did not know — a derived getter on a wire DTO is enough.

Check. event=listener_error … decode failed for BundleDiscoveryInfo in the log.

Fix. Ensure all four codecs keep FAIL_ON_UNKNOWN_PROPERTIES disabled, and give new record components a null-normalising @JsonCreator.

→ Wire Protocol § 8 · Contributing and Conventions § 5


Connection-acquisition timeouts, and events dead-letter immediately

Symptom. JDBC connection-acquisition timeouts under load. Events go straight to the dead-letter queue rather than retrying. Often appears right after enabling a parallel consumer.

Cause. Pool under-sizing, not a lock bug. Each held LockHandle pins one pooled connection for its whole lifetime, and every concurrently-executing transactional handler holds another. An executor of capacity 64 against a 10-connection pool cannot work. The acquisition timeout is classified as transient, and retry is coerced to 0 under an executor — hence the immediate dead-letter.

Fix.

pool ≥ concurrent consumers + Σ(capacity of transactional executors) + headroom

Cap a transactional handler's executor at its share of the pool. Set an explicit retry on parallel handlers that call remote dependencies.

→ Consumer State Store § 5 · Parallel Consumers § 6


A freshly started async consumer is slow to catch up, then fine

Cause. A cold Hikari pool, not the executor. Connections open lazily and cost more to open than a short handler costs to run.

Fix. Raise minimumIdle if catch-up burst latency matters. (Virtual threads are not the problem — 32 concurrent pg_sleep transactions reach full capacity on an 8-core box, so there is no carrier pinning.)

→ Parallel Consumers § 6


Events were lost after a hard kill

Cause. Expected, if the handler uses a consumer executor: delivery is at-most-once for the in-flight window, because the checkpoint advances when a task starts. A graceful stop drains; kill -9 does not. Idempotency protects against duplicates, not against loss.

Fix. Use CheckpointMode.WATERMARK to restore at-least-once, at the cost of replaying the in-flight window after a crash.

→ Parallel Consumers § 3–4


Events are being processed out of order

Cause. Expected under a consumer executor — there is no ordering between parallel events.

Fix. ConsumerExecutors.partitioned(name, lanes) pins events sharing an aggregate id to one lane and applies them in sequence order. No handler change needed. Cost: a hot aggregate serialises.

→ Parallel Consumers § 4


A bundle reconnects but stops receiving messages

Cause. The superseded-session race (Fix A) — the old session's disconnect callback wiping the reconnected session's handlers.

Check. event=disconnect_superseded_skip shows the guard working correctly; its absence around a reconnect is the interesting case.

→ Server Bus § 4


Bundle start-up fails with an executor error

Cause. An @EventHandler(executor = "name") naming an executor never registered via addConsumerExecutor. This is deliberate: ConsumerExecutorValidator fails start-up rather than silently degrading to sequential execution.

Fix. Register the executor, or remove the executor attribute.

→ Parallel Consumers § 1


Bundle start-up fails with IllegalArgumentException

EventoBundle.Builder.start() requires three things:

Missing Message
basePackage Invalid basePackage
bundleId (null or blank) Invalid bundleId
eventoServerMessageBusConfiguration Invalid messageBusConfiguration

→ Getting Started § 2


Registration is rejected

Cause. TokenValidator refused the Hello. The client's start() future fails with Reject(CODE_AUTH_FAILED) — it does not retry into a rejection and does not call System.exit.

Fix. Match the bundle's auth token to the server's evento.server.bus.auth-token.

→ Security Model § 1


Server start-up fails on water-mark configuration

Cause. write-buffer-low-water-mark must be strictly less than write-buffer-high-water-mark; BusProperties throws otherwise.

→ Server Configuration § 2


Consumers are slow and adding bus threads does not help

Cause. evento.es.fetch.concurrency (default 4) is a fair semaphore — no number of bus threads raises it. On a consumer-heavy cluster it is very likely the binding constraint.

Fix. Raise it against measured headroom. It guards heap against concurrent EventFetchRequest result sets and the OOM it was added for is real.

→ Throughput and Capacity § 4


A retried command was applied twice

Cause. A RequestTimeoutException treated as a definite failure. It means the caller stopped waiting — the handler may have applied the work in full. In the reference incident about 20 commands reported as failures had in fact been fully applied.

Fix. Report a timeout as indeterminate (HTTP 504), not failed (500), and retry only what is safe to retry. TooManyPendingRequestsException is the one that is always retryable — nothing was transmitted.

→ Bundle Client § 3 · Throughput and Capacity § 8


evento.server.forwarded{path=reencoded} is rising

Cause. Broker relays are re-encoding CBOR instead of taking the zero-copy raw path. Every Netty-to-Netty forward should be raw.

Fix. Treat it as a regression — the contract is pinned by tests.

→ Wire Protocol § 7 · Observability § 1


Still stuck?

Clone this wiki locally