-
-
Notifications
You must be signed in to change notification settings - Fork 589
Managing Background Services
You can use xyOps to start, stop, restart, and monitor long-running services on your servers, including databases, web servers, application servers, and local AI services. The key is choosing how you want to manage the service's lifetime.
For most services, a good arrangement is to let a service manager such as systemd own the running process, and use short xyOps jobs to control it. xyOps can then monitor the service independently, notify you when something goes wrong, and optionally launch a recovery job.
You can also run a service in the foreground and keep a xyOps job open for its entire lifetime. This guide covers both approaches, along with launching a background process directly when a service manager is unavailable.
A xyOps job tracks the code executed by its Event Plugin. If that code starts a service in the background and then exits, the job completes, even if the service keeps running.
For example, a command such as zkServer.sh start may launch ZooKeeper as a daemon and return almost immediately. A successful job means the startup command completed successfully. It does not mean xyOps is still supervising the ZooKeeper process, or that ZooKeeper is ready to accept requests.
There are three separate questions to answer:
| Question | How to answer it |
|---|---|
| Did the startup operation succeed? | Check the Start job's result, including any startup verification. |
| Is the service process still running? | Use process monitoring or a service-manager status check. |
| Is the service actually healthy? | Check a health endpoint or perform a service-specific operation. |
If a daemon detaches from the startup script, xyOps cannot automatically use the daemon's eventual exit code as the original job's result. To keep a job running until the service exits, use the foreground approach described below.
This is the recommended starting point for services that should run independently of individual xyOps jobs. The service manager handles the ongoing process lifecycle, while xyOps provides controls, workflows, monitoring, and alerts.
On Linux, this commonly means systemd. The same pattern applies to other service managers, Windows services, and services managed by a container runtime.
The examples below assume your service is already registered with systemd as my-web-service.service. Replace this with your actual unit name. You can register a custom application as a service once, then control it from xyOps without launching the application binary directly in each job.
The xySat executing these jobs must have access to the service manager on the server hosting the service. A native xySat installation makes host-service control straightforward. If xySat runs inside a container, commands normally operate in that container's environment; managing host services requires additional access and configuration.
Run the control jobs with an account that has permission to manage the service. These commands must work without interactive password prompts. The application's own runtime account can be configured separately in its service definition.
Create separate events using the Shell Plugin, targeting the server that hosts the service. Give each event an enabled Manual Run trigger so you can launch it from the UI or API, and use it in a Run Event action.
Start with these four events:
| Event | Purpose |
|---|---|
| Start Web Service | Ensure the service is started. |
| Stop Web Service | Stop the service gracefully. |
| Restart Web Service | Deliberately restart the service. |
| Check Web Service Status | Report the current service-manager state. |
#!/bin/bash
set -e
# Start the service. If it is already active, leave it running.
systemctl start my-web-service.service
# Verify that the service manager considers the service active.
systemctl is-active --quiet my-web-service.serviceThe Shell Plugin reports the script's exit status to xyOps. With set -e, a failed command causes the job to fail instead of continuing and reporting success.
The script above starts the service only if it is not already active. Running it again leaves an active service running without launching a second copy or restarting it. This makes the Start event suitable for both manual controls and missing-process recovery. Use the separate Restart event when you want to restart the service deliberately.
#!/bin/bash
# Ask the service manager to stop the service gracefully.
# Use its exit status as the job result.
exec systemctl stop my-web-service.serviceStopping an already-stopped service is normally harmless. After this job completes, its success represents the stop operation's result. It does not change the historical result of the Start job.
#!/bin/bash
set -e
# Deliberately restart the service, even if it is currently running.
systemctl restart my-web-service.service
# Verify the resulting service-manager state.
systemctl is-active --quiet my-web-service.service#!/bin/bash
# Include the current state and recent service logs in the job output.
# An inactive or failed service will produce a nonzero job result.
exec systemctl --no-pager --full status my-web-service.serviceThis Status event is useful for an on-demand check. For ongoing visibility and alerting, use the monitoring features below.
See the systemctl documentation for details on these commands.
A service being active does not always mean it is ready. A web server may still be loading a model, or a database may be recovering before it accepts connections.
If other workflow jobs depend on the service, add a bounded readiness check to the Start and Restart scripts, or place it in a separate job before the dependent jobs. For a web service, you can append this after the service-manager check:
# Replace this with your service's actual readiness endpoint.
HEALTH_URL="http://127.0.0.1:8080/health"
# Configure the number of attempts and request timeouts in seconds.
MAX_ATTEMPTS=12
CONNECT_TIMEOUT=2
REQUEST_TIMEOUT=3
# Give the service several chances to become ready.
# Bound each request and the number of attempts.
for (( ATTEMPT=1; ATTEMPT<=MAX_ATTEMPTS; ATTEMPT++ )); do
if curl --fail --silent --output /dev/null \
--connect-timeout "$CONNECT_TIMEOUT" \
--max-time "$REQUEST_TIMEOUT" "$HEALTH_URL"; then
echo "Service is ready."
exit 0
fi
sleep 2
done
# Fail the startup job so dependent workflow jobs do not proceed.
echo "Service did not become ready in time." >&2
exit 1Use a readiness endpoint whose successful response means the service can perform the work you need. For a database, substitute a database-specific readiness command or a small query. The Status event's service-manager check can also be supplemented with a health check.
Once your controls work, you can launch them through the run_event API, connect them to your own dashboard buttons, or combine them into Workflows. For example, a startup workflow can start a database, verify readiness, then start its application server. A shutdown workflow can stop those services in reverse order.
Important
The total and average functions shown below require xyOps v1.1.0 or newer.
xyOps collects a process list from every monitored server every minute. You can use it to graph a service's memory usage, count matching processes, and detect when the process disappears.
First, open the Server Data Explorer and inspect ServerMonitorData.processes on the actual target server. The command value may contain a full path, arguments, or a generic executable such as java or node. Choose a match that identifies your service specifically.
The expressions below assume a process whose command is exactly postgres. Adapt them to the values you find on your server.
Create a Monitor with Type set to bytes, scoped to the server group hosting the service. For an exact command match, use:
total( processes.list[.command == 'postgres'], 'memRss' )For a substring match, use:
total( find(processes.list, 'command', 'postgres'), 'memRss' )These expressions add up the resident-memory values of all matching processes, including services such as PostgreSQL that run multiple processes. The total function returns zero when there are no matches, so no separate count check or Monitor Plugin is needed.
You can also create a second bytes Monitor to track the average resident memory per matching process:
average( find(processes.list, 'command', 'postgres'), 'memRss' )The average function also returns zero when there are no matches. See Math and Array Functions for more examples of total and average.
For an exact match, create an integer Monitor using:
count(processes.list[.command == 'postgres'])For a substring match:
count(find(processes.list, 'command', 'postgres'))This makes the number of matching processes visible over time. A broad substring match can include unrelated applications or multiple instances, so confirm the matches before relying on it for recovery.
Create an Alert, scoped to the server group where the service is expected to run. For an exact match, use this expression:
count(processes.list[.command == 'postgres']) == 0For a substring match:
count(find(processes.list, 'command', 'postgres')) == 0Use a count comparison rather than negating the filtered array: an empty array is truthy, so !processes.list[.command == 'postgres'] does not detect absence.
Monitors and alerts operate on minute samples. Configure the alert's Samples value to tolerate brief startup or restart gaps, and use its Test dialog to check the expression against your server. These checks require an available process list; an offline server needs separate monitoring.
A process can be present while the service is unhealthy. It may be stuck, unable to reach its database, or returning errors to every request. Process monitoring answers whether something exists; a health check answers whether it works.
For a complete example, follow the Web Service Health Alert recipe. It uses a Monitor Plugin to check a local HTTP endpoint every minute, with bounded requests and retries, and an alert that matches the resulting UP or DOWN output. You do not need a numeric Monitor for that recipe because the alert uses the Plugin output directly.
For other services, apply the same pattern with a suitable check, such as a database query or a Redis PING. Choose a check that exercises the behavior your application depends on.
Once your monitoring and control jobs are tested, you can add a Run Event action to an alert.
For a missing-process alert:
- Select the condition for when the alert fires (
alert_new). - Choose your Start event.
- Enable Target Alert Server so the job runs on the server where the process disappeared. This target override applies to ordinary events; a workflow needs its own targeting arrangement.
- Leave Clear Alert on Job Completion unchecked. Let monitoring confirm recovery and clear the alert when its expression becomes false.
- Keep the alert's Job Limit and Job Abort options off for this use case so the alert does not block its own recovery job or abort unrelated work.
The Start event must have an enabled Manual Run trigger. Use its readiness check to make the recovery job fail if startup does not produce a usable service.
A health alert needs a different recovery decision. An idempotent Start operation will leave an existing process alone, even if that process is unhealthy. You may want a dedicated recovery event that checks health again and restarts only if the service is still unhealthy. Consider whether restarting is appropriate for that service and failure condition.
Alert actions run when an alert fires, rather than every minute while it remains active. If the recovery job fails and the alert stays active, xyOps does not automatically keep launching more recovery jobs from that same invocation.
Use a Max Retry Limit on the recovery event if you want a bounded number of additional attempts, with a delay between them. Add an error notification so you know when recovery fails. Configure concurrency limits to prevent overlapping recovery operations on the same service.
For systemd-managed services, you can also let systemd handle local crash recovery with a policy such as Restart=on-failure, while xyOps handles monitoring, notifications, and broader remediation. Coordinate the two policies so they do not compete to restart the same service. See the systemd service documentation.
If your alert starts a missing service automatically, pressing Stop can cause the alert to start it again on a later sample. Decide when the service is expected to be running before enabling recovery.
For an always-on service, you can disable its recovery alert during planned maintenance. Remember that disabling an alert definition affects all servers in its scope.
For a service that users routinely start and stop, use a per-server expected-state flag. For example, your control scripts can maintain a local marker file, and a Monitor Plugin can report whether that marker exists alongside the service's health. The alert should fire only when the service is expected to run and is missing or unhealthy. The recovery job should recheck that expected state before starting anything, including on retries.
The intentional Stop operation should mark the service as no longer expected to run before stopping it. A Start operation should mark it as expected to run. This lets monitoring distinguish an outage from a deliberate shutdown.
If you want the job itself to remain running for the service's entire lifetime, launch the service in its foreground mode. Do not use a startup command that daemonizes and exits.
For example, a Shell Plugin script could contain:
#!/bin/bash
# Replace the shell with the server process and keep it attached.
# Supply the executable path and options appropriate for your service.
exec /opt/llama.cpp/build/bin/llama-server \
--model /opt/models/example.gguf \
--host 127.0.0.1 \
--port 8080In this arrangement, the script remains running while the server runs, its output becomes job output, and the Shell Plugin uses its eventual exit status as the job result. A nonzero exit or termination signal produces a failed result; an explicit xyOps abort is handled as a job abort.
See the llama.cpp server documentation for its setup and command options. For other services, check their documentation for a foreground command. Some service wrappers provide a separate foreground subcommand. Polling an already-detached daemon can keep a script open, but it does not by itself recover that daemon's real exit code.
Long-running jobs are possible, but keeping an indefinite service inside a job has operational consequences:
- Runtime limits: Review every applicable Max Time Limit, including limits inherited from the category and server groups. A timeout can terminate the service.
- Capacity: The running job occupies job concurrency capacity for its lifetime. Account for event limits and Max Jobs Per Server.
- Output: A busy service can generate substantial job output. Review output limits and consider the service's own logging configuration.
- Aborts: Review the Plugin's process termination settings and verify that aborting the job stops the intended service and its children cleanly.
- Workflows: A workflow waiting for this job's success cannot proceed until the service exits. Use separate start-and-readiness jobs when other workflow steps need to run while the service stays up.
If you want the job to relaunch automatically when it completes, see the Continuous Jobs recipe. Plan the completion actions and retry behavior carefully, including how you will stop the service intentionally without immediately relaunching it.
If a service manager is unavailable, a Start job can launch the application in the background and then finish. This can be useful for simple applications, but your scripts must handle the lifecycle details that a service manager would otherwise provide.
On Unix-like systems, a basic launch might look like this:
# Illustrate the background launch only.
# Create the log directory with suitable ownership before using this.
nohup /opt/my-app/bin/server \
</dev/null >>/var/log/my-app/server.log 2>&1 &
# This identifies the process launched by the shell.
# It does not prove that the application became ready.
SERVICE_PID=$!Redirect all three standard streams so the background process does not keep the job's input or output streams open. nohup handles hangup signals; it does not provide supervision, automatic recovery, or protection against explicit termination by xySat's configured abort behavior.
A complete implementation also needs to:
- Prevent duplicate starts, including simultaneous requests.
- Record process identity and verify it before stopping anything. A saved PID alone can become stale and later belong to an unrelated process.
- Handle applications that fork again or change their process identity.
- Verify readiness before reporting startup success.
- Stop the correct process and its children gracefully, with a bounded wait.
- Manage logs, runtime permissions, and any desired startup-at-boot behavior.
Use the application's own start/stop tools if they already handle these details, and add independent monitoring and health checks as described above. If these scripts become complicated, registering the application with a service manager is usually the easier arrangement to maintain.
| What you need | Suggested approach |
|---|---|
| A database or web server that runs independently of jobs | Service manager, short control events, and independent monitoring. |
| Start/stop buttons or API-driven controls | Separate Start and Stop events, with expected-state handling if recovery is enabled. |
| A workflow that starts services in dependency order | Short startup jobs with readiness checks before dependent steps. |
| A running job that represents the service's full lifetime | Foreground execution with deliberate limits and abort behavior. |
| A background application without a service manager | Detached launch scripts that manage identity, readiness, shutdown, and logs. |
| Automatic recovery from an outage | A scoped process or health alert with a bounded recovery event. |