Skip to content

Operations Monitoring

Andrew MacGaffey edited this page Jul 16, 2026 · 17 revisions

Operations: Monitoring & Diagnostics

When something looks wrong, the API gateway answers three questions fast: is the system healthy, if not which part is broken, and why. Every check on this page is a read-only HTTP query to mf-api-gateway:9090 - the same commands work from any machine once the Deployment: Basics#setting-up-a-dns-entry-for-your-api-gateway is in place.

Prefer a screen to a shell? The Dashboard presents these same signals in a browser. It is built entirely on the REST APIs below - so it shows the same data, live, without the curl. (Dashboard user guide to follow.)


Is the system healthy?

Ask every component for its state:

curl -s "http://mf-api-gateway:9090/api/application-state/v1/*/name,role,state,info"

Healthy is every component FULLY_OPERATIONAL:

[
    { "name": "mf-mds-consolidated", "role": "ACTIVE", "state": "FULLY_OPERATIONAL", "info": "All constituents fully operational" },
    { "name": "mf-eta-simulator",    "role": "ACTIVE", "state": "FULLY_OPERATIONAL", "info": "All constituents fully operational" }
]

Anything other than FULLY_OPERATIONAL on a container is your signal to drill in.


Which part is broken?

A container's state is a roll-up of its constituents - the servers and controllers inside it. When a container is not fully operational, list them and find the one that is not:

curl -s "http://mf-api-gateway:9090/api/application-state/v1/*/constituents/*/name,state"
[[ { "name": "PacketProtocolServer [protocol=SessionProtocol]", "state": "FULLY_OPERATIONAL" },
   { "name": "AuthenticationController: FILE-AUTHN",            "state": "FULLY_OPERATIONAL" },
   { "name": "Memory Monitor",                                  "state": "FULLY_OPERATIONAL" } ]]

Each state is one of:

State Meaning
FULLY_OPERATIONAL working normally
INITIALIZING still starting - acquiring resources or connections
RESOURCE_COMPLETE resources ready; awaiting activation
STRESSED degraded but still serving
RECOVERING restoring a resource it lost
INITIALIZATION_ERROR failed to start
NON_FUNCTIONAL failed at runtime

A container reports the worst of its constituents' states, so the roll-up is FULLY_OPERATIONAL only when every constituent is. Constituents can themselves contain constituents, so a broken one may be a level down - follow the same /constituents path into it. The lifecycle these states move through is shown in Architecture: Advanced.


Why? - the state history

Each container and constituent keeps a timestamped state-history buffer - the record of how it got to its current state. This is usually enough to point at the cause:

curl -s "http://mf-api-gateway:9090/api/application-state/v1/*/history"
2026-07-15 16:09:34 FULLY_OPERATIONAL   All constituents fully operational
2026-07-15 16:09:34 RESOURCE_COMPLETE   ETATableAdapter-Default.Quotes.LOCAL -> Waiting for projector state resolution
2026-07-15 16:09:34 RESOURCE_COMPLETE   ETATableAdapter-Default.Quotes.CurveInputs -> Waiting for projector state resolution

Each line is a timestamp, a state, and a reason, most recent first: the top line is the current state, the lines below show how it got there. A constituent that is stuck or has failed shows its non-FULLY_OPERATIONAL state here with the reason - what it is waiting on, or what failed. For a specific constituent, append its name: .../constituents/<name>/history. Turning a cause into a fix is the job of Operations: Troubleshooting & FAQ.


Resource monitoring

Three resources - CPU, memory, and I/O - across three levels: containers, servers, and shards.

Containers - CPU and memory. Each container reports its host stats:

curl -s "http://mf-api-gateway:9090/api/application-state/v1/*/name,system"

Look at Summary.loadAverage and Summary.memoryUsedPercent for load and memory pressure at a glance.

Servers and shards - I/O. The distribution servers report throughput per shard, and whether the system has had to protect a slow consumer:

curl -s "http://mf-api-gateway:9090/api/push-distribution-server/v1/*/shards/"

Per shard, writeChannel/ioSummary/bytesPerSecond is the current output rate, and autoConflationEngaged tells you the system has engaged conflation to keep a slow consumer from holding others back. The full per-channel stat surface is in the REST API reference.


Logs

Two ways in:

  • On disk. Every component writes to its log directory - the same place the configuration reports land (see Configuration: Basics).

  • Over the gateway, where the deployment exposes the log service - ask for recent errors:

    curl -s "http://mf-api-gateway:9090/api/log-service/v1/queryLogs?level=SEVERE&limit=20"
    

    Filter with logger, text, since, and limit; results are newest-first.

To investigate a specific component, raise its log level - a change to a running system, so restore it when you are done:

curl -X PATCH "http://mf-api-gateway:9090/api/log/v1/<component>/level?level=FINE"

Going deeper

Clone this wiki locally