-
Notifications
You must be signed in to change notification settings - Fork 0
Operations Monitoring
When something looks wrong, the API gateway answers three questions fast: is the system healthy, if not which part is broken, and why. Every check on this page is a read-only HTTP query to mf-api-gateway:9090 - the same commands work from any machine once the Deployment: Basics#setting-up-a-dns-entry-for-your-api-gateway is in place.
Prefer a screen to a shell? The Dashboard presents these same signals in a browser. It is built entirely on the REST APIs below - so it shows the same data, live, without the
curl. (Dashboard user guide to follow.)
Ask every component for its state:
curl -s "http://mf-api-gateway:9090/api/application-state/v1/*/name,role,state,info"
Healthy is every component FULLY_OPERATIONAL:
[
{ "name": "mf-mds-consolidated", "role": "ACTIVE", "state": "FULLY_OPERATIONAL", "info": "All constituents fully operational" },
{ "name": "mf-eta-simulator", "role": "ACTIVE", "state": "FULLY_OPERATIONAL", "info": "All constituents fully operational" }
]Anything other than FULLY_OPERATIONAL on a container is your signal to drill in.
A container's state is a roll-up of its constituents - the servers and controllers inside it. When a container is not fully operational, list them and find the one that is not:
curl -s "http://mf-api-gateway:9090/api/application-state/v1/*/constituents/*/name,state"
[[ { "name": "PacketProtocolServer [protocol=SessionProtocol]", "state": "FULLY_OPERATIONAL" },
{ "name": "AuthenticationController: FILE-AUTHN", "state": "FULLY_OPERATIONAL" },
{ "name": "Memory Monitor", "state": "FULLY_OPERATIONAL" } ]]Each state is one of:
| State | Meaning |
|---|---|
FULLY_OPERATIONAL |
working normally |
INITIALIZING |
still starting - acquiring resources or connections |
RESOURCE_COMPLETE |
resources ready; awaiting activation |
STRESSED |
degraded but still serving |
RECOVERING |
restoring a resource it lost |
INITIALIZATION_ERROR |
failed to start |
NON_FUNCTIONAL |
failed at runtime |
A container reports the worst of its constituents' states, so the roll-up is FULLY_OPERATIONAL only when every constituent is. Constituents can themselves contain constituents, so a broken one may be a level down - follow the same /constituents path into it. The lifecycle these states move through is shown in Architecture: Advanced.
Each container and constituent keeps a timestamped state-history buffer - the record of how it got to its current state. This is usually enough to point at the cause:
curl -s "http://mf-api-gateway:9090/api/application-state/v1/*/history"
2026-07-15 16:09:34 FULLY_OPERATIONAL All constituents fully operational
2026-07-15 16:09:34 RESOURCE_COMPLETE ETATableAdapter-Default.Quotes.LOCAL -> Waiting for projector state resolution
2026-07-15 16:09:34 RESOURCE_COMPLETE ETATableAdapter-Default.Quotes.CurveInputs -> Waiting for projector state resolution
Each line is a timestamp, a state, and a reason, most recent first: the top line is the current state, the lines below show how it got there. A constituent that is stuck or has failed shows its non-FULLY_OPERATIONAL state here with the reason - what it is waiting on, or what failed. For a specific constituent, append its name: .../constituents/<name>/history. Turning a cause into a fix is the job of Operations: Troubleshooting & FAQ.
Three resources - CPU, memory, and throughput - across three levels: containers, servers, and shards.
Containers - CPU and memory. Each container reports its host stats:
curl -s "http://mf-api-gateway:9090/api/application-state/v1/*/name,system"
[
{ "name": "mf-mds-consolidated",
"system": {
"Summary": { "loadAverage": 1.44, "memoryUsedPercent": 69 },
"Details": { "cpuCount": 24, "osName": "Linux" },
"Stats": { "memoryFree": 490384712, "memoryMax": 16793993216 } } }
]Summary.loadAverage and Summary.memoryUsedPercent give load and memory pressure at a glance.
Servers and shards - throughput. The push distribution server reports per-shard delivery stats, and whether the system has had to protect a slow consumer:
curl -s "http://mf-api-gateway:9090/api/push-distribution-server/v1/*/shards/"
[
{ "shardID": "b8a8282a-7f27-4ff5-a456-de572a586bb0",
"offeredUpdates": 2300314,
"autoConflationEnabled": true,
"autoConflationEngaged": false,
"clientProfile": { "smoothedMeanRate": 2004.77, "projectedMaxRateUpdatesPerSec": 125000 } }
]clientProfile.smoothedMeanRate is the current delivery rate (updates per second) and offeredUpdates the running total; autoConflationEngaged tells you the system has engaged conflation to keep a slow consumer from holding the others back. (Each shard carries more - streams, dictionaries, orchestration stats; byte-level I/O lives on the channels.) The full stat surface is in the REST API reference.
Two ways in:
-
On disk. Every component writes to its log directory - the same place the configuration reports land (see Configuration: Basics).
-
Over the gateway, where the deployment exposes the log service - ask for recent errors:
curl -s "http://mf-api-gateway:9090/api/log-service/v1/queryLogs?level=SEVERE&limit=20"On a healthy system that returns an empty list -
[]. Widen the level to see records; each entry is one log line:[ { "l": "WARNING", "ln": "com.metafluent.blueprint.rtc.eta.server.tcpip", "m": "initChannel", "msg": "channel.ioctl() failed: channel not in active state", "sID": "34f1505a-560e-48c3-95bd-8d020889cb6e" } ]Filter with
logger,text,since, andlimit; results are newest-first. (llevel,lnlogger,c/mclass / method,sIDthe reporting container.) For the full set of ways to slice the logs - by component, container, text, and severity - see Operations: Logging.
To investigate a specific component, raise its log level - a change to a running system, so restore it when you are done:
curl -X PATCH "http://mf-api-gateway:9090/api/log/v1/<component>/level?level=FINE"
- Dashboard - the same signals in a browser (user guide to follow).
-
REST API - the
apidoc, field projection, and predicate selection used throughout this page. - Operations: Troubleshooting & FAQ - symptom to cause to fix.
- Glossary - definitions of the terms used here.
Elastic MDS documentation - (c) MetaFluent LLC - Confidential. Tracked in IssueTracking#586.
Getting Started
Deployment Cookbook
Concepts
- Architecture: Basics
- Access Control
- Architecture: Advanced
- Security: Basics
- Security: Advanced
- Glossary
Configuration
Configuration Cookbook
Deployment
Operations
- Monitoring & Diagnostics
- Logging
- Dashboard
- Troubleshooting & FAQ
- AI-Assisted Troubleshooting
- API Token Administration
Diagnostic Cookbook
Developing Applications
Reference