-
Notifications
You must be signed in to change notification settings - Fork 0
Troubleshooting
A symptom-first index of the failure modes that are documented in the codebase. Each entry links to the page with the full mechanism.
Symptom. Throughput falls by orders of magnitude — the reference incident went from 16 writes/s to 0.07 writes/s. CPU under 1%, databases idle, no blocked queries, no lock contention. A single request sent by hand returns in 0.1 s. Timeouts everywhere. Reducing client concurrency increases throughput.
Cause. Congestion collapse. Requests carry a client deadline and nothing cancels the work when it expires, so the server spends its capacity producing responses nobody is waiting for. Retries keep it there, so it does not self-heal when the triggering load stops.
Check. evento.server.bus.executor.saturated (rising), …queue.depth, and
WARN event=bus_business_executor_saturated.
Fix. Reduce client concurrency first, stop the retry storm, then size the server. Do not enlarge the queue — that deepens the collapse.
Symptom. The bundle registers, enables and processes events normally, yet its handlers, catalog entries and flows are missing from the GUI. Nearly invisible — nothing fails.
Cause. The rich evento:bundle-discovery notification was rejected at decode, killing the
entire BundleDiscoveryInfo. Historically caused by a peer emitting a property the other side did not
know — a derived getter on a wire DTO is enough.
Check. event=listener_error … decode failed for BundleDiscoveryInfo in the log.
Fix. Ensure all four codecs keep FAIL_ON_UNKNOWN_PROPERTIES disabled, and give new record
components a null-normalising @JsonCreator.
→ Wire Protocol § 8 · Contributing and Conventions § 5
Symptom. JDBC connection-acquisition timeouts under load. Events go straight to the dead-letter queue rather than retrying. Often appears right after enabling a parallel consumer.
Cause. Pool under-sizing, not a lock bug. Each held LockHandle pins one pooled connection for its
whole lifetime, and every concurrently-executing transactional handler holds another. An executor
of capacity 64 against a 10-connection pool cannot work. The acquisition timeout is classified as
transient, and retry is coerced to 0 under an executor — hence the immediate dead-letter.
Fix.
pool ≥ concurrent consumers + Σ(capacity of transactional executors) + headroom
Cap a transactional handler's executor at its share of the pool. Set an explicit retry on parallel
handlers that call remote dependencies.
→ Consumer State Store § 5 · Parallel Consumers § 6
Cause. A cold Hikari pool, not the executor. Connections open lazily and cost more to open than a short handler costs to run.
Fix. Raise minimumIdle if catch-up burst latency matters. (Virtual threads are not the problem —
32 concurrent pg_sleep transactions reach full capacity on an 8-core box, so there is no carrier
pinning.)
Cause. Expected, if the handler uses a consumer executor: delivery is at-most-once for the
in-flight window, because the checkpoint advances when a task starts. A graceful stop drains;
kill -9 does not. Idempotency protects against duplicates, not against loss.
Fix. Use CheckpointMode.WATERMARK to restore at-least-once, at the cost of replaying the
in-flight window after a crash.
Cause. Expected under a consumer executor — there is no ordering between parallel events.
Fix. ConsumerExecutors.partitioned(name, lanes) pins events sharing an aggregate id to one lane
and applies them in sequence order. No handler change needed. Cost: a hot aggregate serialises.
Cause. The superseded-session race (Fix A) — the old session's disconnect callback wiping the reconnected session's handlers.
Check. event=disconnect_superseded_skip shows the guard working correctly; its absence
around a reconnect is the interesting case.
Cause. An @EventHandler(executor = "name") naming an executor never registered via
addConsumerExecutor. This is deliberate: ConsumerExecutorValidator fails start-up rather than
silently degrading to sequential execution.
Fix. Register the executor, or remove the executor attribute.
EventoBundle.Builder.start() requires three things:
| Missing | Message |
|---|---|
basePackage |
Invalid basePackage |
bundleId (null or blank) |
Invalid bundleId |
eventoServerMessageBusConfiguration |
Invalid messageBusConfiguration |
Cause. TokenValidator refused the Hello. The client's start() future fails with
Reject(CODE_AUTH_FAILED) — it does not retry into a rejection and does not call System.exit.
Fix. Match the bundle's auth token to the server's evento.server.bus.auth-token.
Cause. write-buffer-low-water-mark must be strictly less than
write-buffer-high-water-mark; BusProperties throws otherwise.
Cause. evento.es.fetch.concurrency (default 4) is a fair semaphore — no number of bus
threads raises it. On a consumer-heavy cluster it is very likely the binding constraint.
Fix. Raise it against measured headroom. It guards heap against concurrent
EventFetchRequest result sets and the OOM it was added for is real.
Cause. A RequestTimeoutException treated as a definite failure. It means the caller stopped
waiting — the handler may have applied the work in full. In the reference incident about 20 commands
reported as failures had in fact been fully applied.
Fix. Report a timeout as indeterminate (HTTP 504), not failed (500), and retry only what is
safe to retry. TooManyPendingRequestsException is the one that is always retryable — nothing was
transmitted.
→ Bundle Client § 3 · Throughput and Capacity § 8
Cause. Broker relays are re-encoding CBOR instead of taking the zero-copy raw path. Every
Netty-to-Netty forward should be raw.
Fix. Treat it as a regression — the contract is pinned by tests.
→ Wire Protocol § 7 · Observability § 1
- SUPPORT.md — where to ask
- Issues — bugs and feature requests
- Security problems: do not open an issue — follow SECURITY.md
Evento Framework — Copyright 2020–2026 © Gabor Galazzo. Dual-licensed under AGPL-3.0 and a commercial licence.
This wiki documents the implementation; the repository is authoritative where the two disagree. Found something out of date? Open an issue.
Getting oriented
Internals
Operations
- Server Configuration
- Throughput and Capacity
- Observability
- Security Model
- Server REST API
- Troubleshooting
Project