-
Notifications
You must be signed in to change notification settings - Fork 0
Observability
Everything is exposed through Spring Boot Actuator on the HTTP port (3000):
management.endpoints.web.exposure.include=health,info,prometheus,metricsMicrometer with micrometer-registry-prometheus, so /actuator/prometheus and /actuator/metrics
both work.
Wired by BusMetricsBinder (com.evento.server.bus.spring):
| Meter | Meaning |
|---|---|
evento.server.connections |
Connected bundle nodes |
evento.server.connections.available |
Nodes able to receive |
evento.server.correlations.pending |
Outstanding server-initiated correlations |
evento.server.forwarding.table.size |
In-flight relayed requests |
evento.server.forwarded{path=raw|reencoded} |
Zero-copy versus re-encoded forward counts |
forwarded{path=reencoded} rising for Netty-to-Netty traffic means the zero-copy path has been lost —
that is a regression, and tests pin the contract.
Not yet covered: heartbeat lag, per-type processing-duration timers, consumer lag.
| Meter | Meaning |
|---|---|
evento.server.bus.executor.pool.size |
Current threads |
evento.server.bus.executor.max |
Configured maximum |
evento.server.bus.executor.active |
Threads currently executing |
evento.server.bus.executor.queue.depth |
The early warning |
evento.server.bus.executor.saturated |
The alerting signal — only moves when the pool is at max and the queue is full |
Alert on any sustained increase in saturated. Full playbook:
Throughput and Capacity.
Wired by ConsumerMetricsRegistry (com.evento.server.bus.spring) and fed by a periodic
ConsumerStatsMessage that bundles push on the admin notification channel
(setConsumerStatsInterval, default 30 s).
Why push, not poll? The counters live in bundle-side objects. Polling per scrape would make one scrape a round trip per consumer.
| Meter | Tags |
|---|---|
evento.consumer.executor.{capacity,in.flight,admitted,rejected,completed,failed} |
bundle,instance,executor |
evento.consumer.async.{in.flight,submit.timeouts,transient.failures} |
bundle,instance,consumer,component |
evento.consumer.executor.rejectedis the alerting signal — the executor reporting it is the bottleneck.
Meters are removed on BusEvent.NodeLeft, so rolling restarts do not leak a time series per dead
instance. Bundles with no consumer executor push nothing at all.
| Endpoint | Contents |
|---|---|
/actuator/health |
BusHealthIndicator — bus bound port and node counts; DataSource health via Actuator's built-in indicator |
/actuator/health/liveness · /readiness
|
Kubernetes probes (management.endpoint.health.probes.enabled=true) |
management.endpoint.health.show-details=when_authorized, so details need authentication.
PerformanceStoreService persists HandlerInvocationCountPerformance and
HandlerServiceTimePerformance; PerformanceController serves them to the GUI.
Sampling is evento.performance.capture.rate (1 in the shipped properties, 0.1 in the compose
file); retention is evento.telemetry.ttl days.
This is what powers the GUI's performance model and time series — and, per design decision 9, it is all the framework does about load. Scaling belongs to your orchestrator.
SLF4J, key=value on every bus lifecycle transition. Events worth knowing:
| Log event | Level | Meaning |
|---|---|---|
event=bus_business_executor_saturated |
WARN |
Pool at max and queue full. Rate-limited by business-executor-saturation-warn-interval
|
event=bundle_correlation_expired … byType={…} |
WARN |
Requests expired; byType names the drowning payload type |
event=disconnect_superseded_skip |
— | A superseded session's disconnect was correctly ignored (see Server Bus § 4) |
event=listener_error … decode failed for BundleDiscoveryInfo |
— | Discovery metadata was rejected — the bundle works but the dashboard shows nothing for it (see Wire Protocol § 8) |
TracingAgent's default is an honest no-op — telemetry and metrics only. For real distributed
tracing, wire a custom agent or SentryTracingAgent via EventoBundle.Builder.setTracingAgent.
| Alert on | Why |
|---|---|
evento.server.bus.executor.saturated increasing |
Congestion collapse is starting |
evento.consumer.executor.rejected increasing |
A consumer executor is the bottleneck |
evento.server.forwarded{path=reencoded} rising for bundle traffic |
Zero-copy regression |
WARN event=bundle_correlation_expired |
Capacity exhaustion; byType tells you where |
/actuator/health not UP
|
Bus or datasource down |
Evento Framework — Copyright 2020–2026 © Gabor Galazzo. Dual-licensed under AGPL-3.0 and a commercial licence.
This wiki documents the implementation; the repository is authoritative where the two disagree. Found something out of date? Open an issue.
Getting oriented
Internals
Operations
- Server Configuration
- Throughput and Capacity
- Observability
- Security Model
- Server REST API
- Troubleshooting
Project