-
Notifications
You must be signed in to change notification settings - Fork 0
Operations Troubleshooting
This is the page to reach for when something is wrong. It is organized the way you actually work a problem: find the symptom that matches, look at the likely cause, and follow the remedy. It builds on Operations: Monitoring & Diagnostics (is it healthy, which part, why) and Operations: Logging (what the component said); every check here is a read-only query to mf-api-gateway:9090.
Audience: Operator (primary), Developer.
Prefer a screen? The Dashboard shows most of these signals live in a browser, and AI-Assisted Troubleshooting is where this is heading - an assistant that runs the checks, captures the evidence, and analyzes it for you.
Jump to the closest match. If nothing quite fits, work the standard diagnostic procedure below.
Startup and availability - something will not come up
User (client) trouble - an application cannot connect, or is not getting what it asked for
Data trouble - the content itself is missing, stale, or wrong
Performance - slow or overloaded
-
Confirm the health first, so you know which component and what state it is in:
curl -s "http://mf-api-gateway:9090/api/application-state/v1/*/name,role,state,info" -
Find the matching symptom in the sections below and follow its likely cause and remedy.
-
If the system was healthy and is now failing, capture the evidence before you restart. A restart clears the very state that explains the failure - capturing first is the difference between a diagnosis and a guess:
- Capture the full application state - state, constituents, and their history - over REST (see Operations: Monitoring & Diagnostics).
- Where the deployment exposes them, pull a diagnostic (DataFabric) dump and export the recent log records from the log service (see Operations: Logging).
- Open an issue with MetaFluent support and attach what you captured; note anything unusual about the timing.
(This capture-and-analyze step is exactly what AI-Assisted Troubleshooting is being built to do for you.)
Confirm: read the state and reason, then drill into the stuck constituent and its history.
curl -s "http://mf-api-gateway:9090/api/application-state/v1/*/name,state,info"
curl -s "http://mf-api-gateway:9090/api/application-state/v1/*/constituents/*/name,state"
curl -s "http://mf-api-gateway:9090/api/application-state/v1/*/constituents/<name>/history"
| Symptom | Likely cause | Remedy |
|---|---|---|
Stuck at INITIALIZING, or terminal INITIALIZATION_ERROR
|
A required setting is missing or wrong (an unset $(VAR), a bad host or path) |
Check the effective configuration (Configuration: Basics); correct the deployment's deployment.properties; restart the container |
INITIALIZING never completes |
A dependency the component needs (a source, store, or peer) is not up or not routable | The info/history reason names it - bring the dependency up or fix the address |
INITIALIZATION_ERROR naming a file or resource |
A file, descriptor, or mount the component expects is not present | Provide the resource the reason names; restart |
| A port/bind error in the logs | Another process holds the configured port | Free the port or reconfigure it; restart |
INITIALIZATION_ERROR is terminal - the component will not retry out of it. Fix the cause and restart.
Redundant deployments run each protected service (session, pub/sub, SQL, orchestration) as an active/standby pair. Before treating a standby as a problem, know what a healthy one looks like - a hot standby is deliberately held short of fully operational.
List the roles directly. The predicate form gives you exactly the standbys, with their state and reason, in one call:
curl -s "http://mf-api-gateway:9090/api/application-state/v1/role=STANDBY/name,role,state,info"
curl -s "http://mf-api-gateway:9090/api/application-state/v1/role=ACTIVE/name,role"
A healthy standby looks like this - STANDBY, RESOURCE_COMPLETE, "awaiting activation":
[ { "name": "mf-session", "role": "STANDBY", "state": "RESOURCE_COMPLETE",
"info": "PacketProtocolServer [protocol=SessionProtocol](self) -> RESOURCE_COMPLETE: Ready; awaiting activation" } ]Redundant services appear as a same-named pair - one ACTIVE, one STANDBY - that share the name and are told apart by system.summary.systemID. (Note the projection quirk: a lone dotted token after a predicate returns []; use the CSV form .../role=STANDBY/name,system.summary.systemID.)
| Symptom | Likely cause | Remedy |
|---|---|---|
A standby sits at RESOURCE_COMPLETE and never reaches FULLY_OPERATIONAL
|
Normal - a hot standby is held at RESOURCE_COMPLETE, awaiting activation |
No action; it is ready and will take over if the active fails |
| A component you expect is absent from the state list or the catalog | Its container did not start, or cannot reach the gateway | Confirm the container is running (Deployment: Basics); check gateway reachability (see the FAQ) |
No instance shows role ACTIVE for a service that should have one |
The active/standby members cannot see each other, so none has claimed the active role | Confirm the members are up and can reach one another over the control network; see active/standby in Architecture: Advanced |
| A standby did not promote after its active failed | The standby was not ready (RESOURCE_COMPLETE), or could not confirm the active was gone |
Read the standby's state and history for the reason |
| Symptom | Likely cause | Remedy |
|---|---|---|
| The connection times out | Wrong host or port | Client connects to the session server (e.g. :8900); verify the connection string (JMS Application Development, JDBC Application Development) |
| The connection is refused or nothing is listening | The server the client points at is not running, or is not reachable on the network | Confirm the target is up and routable |
| The connection is rejected with a context error | The named context is invalid or not loaded | Verify the context name the application selects (JMS Application Development) |
| The connection is rejected on login | Authentication or entitlement failure | Supply valid credentials; confirm the user is entitled (Entitlements Context); login failures are logged server-side |
If your REST calls are what fail to connect (not a client app), that is a gateway/DNS issue - see the FAQ below.
Confirm: find the adapter/constituent that is not fully operational and read its history and logs (see Operations: Logging).
| Symptom | Likely cause | Remedy |
|---|---|---|
A container rolls up to RECOVERING; its data goes stale |
The upstream feed or connection went away (the source, the network path, or credentials) | The system reconnects on its own - confirm the source is reachable and watch for the return to FULLY_OPERATIONAL. RECOVERING is the system working, not a fault; it is a problem only if it does not clear |
| Warnings about fields or data conversion; specific content marked stale | The source metadata (data dictionary) does not match the data received | The affected content is marked stale so consumers are not misled; align the metadata/dictionary, or correct the publisher sending non-conforming data |
The history names the connection and the logs show the drop and retry - for the worked example (an ETA feed loss), see Operations: Logging.
| Symptom | Likely cause | Remedy |
|---|---|---|
| Subscribes successfully but receives no updates | Invalid topic/context syntax - the topic is passive (accepted but matches no content) | The context decides topic syntax; the bare v5 shorthand only works under the v5 context, while MarketData-6.0.0 needs the schema-qualified form (Dynamic Data Conventions) |
| Receives only status messages, or "access denied" | Not entitled - the client, or the server on the client's behalf, is not permitted the content | Confirm the entitlement (Entitlements Context); the status text carries the reason |
Receives STALE
|
The upstream source is disrupted | This is the source-data case above - check the adapter's state |
| Symptom | Likely cause | Remedy |
|---|---|---|
| One client lags while others are fine | The client is a slow consumer; the system engaged auto-conflation to protect the others | Protective, not an error. Confirm with the shards query below; if chronic, the client must read faster or subscribe to less |
A container is sluggish or reports STRESSED
|
CPU or memory pressure - too much subscribed content, an undersized container, or a runaway consumer | Add capacity, or let auto-scaling add a projector where it applies (Architecture: Advanced) |
curl -s "http://mf-api-gateway:9090/api/push-distribution-server/v1/*/shards/autoConflationEngaged=true/shardID,offeredUpdates"
curl -s "http://mf-api-gateway:9090/api/application-state/v1/*/name,system"
What host do I use for REST calls? http://mf-api-gateway:9090, always - never localhost. If a REST call cannot connect, the mf-api-gateway name is not resolving from your machine: set up the DNS entry (Deployment: Basics), and confirm the gateway container is up.
Where are the logs? Each component writes to its own log directory on disk, and the centralized log service answers queries across all of them at once (Operations: Logging).
How do I turn up logging for one component? PATCH its log level, and set it back when you are done (REST API).
Is RECOVERING an error? No - the component lost a resource and is restoring it; it clears on its own. The only terminal states are INITIALIZATION_ERROR and NON_FUNCTIONAL (Architecture: Basics).
What does active/standby failover look like? A standby is hot; on failover its role flips from STANDBY to ACTIVE in seconds and it takes over (Architecture: Advanced).
It was working and now it is not - what do I do first? Capture the evidence before you restart, per the standard diagnostic procedure above.
For the meaning of a specific WARNING or SEVERE message and the recommended action, see the Log Message Reference - a per-component catalog of the messages a deployment can emit (in preparation).
- Operations: Monitoring & Diagnostics - is it healthy, which part, and why.
- Operations: Logging - slicing the logs.
- AI-Assisted Troubleshooting - the assistant that runs these checks for you (preview).
- Glossary - definitions of the terms used here.
Elastic MDS documentation - (c) MetaFluent LLC - Confidential. Tracked in IssueTracking#586.
Getting Started
Deployment Cookbook
Concepts
- Architecture: Basics
- Access Control
- Architecture: Advanced
- Security: Basics
- Security: Advanced
- Glossary
Configuration
Configuration Cookbook
Deployment
Operations
- Monitoring & Diagnostics
- Logging
- Dashboard
- Troubleshooting & FAQ
- AI-Assisted Troubleshooting
- API Token Administration
Diagnostic Cookbook
Developing Applications
Reference