Repository navigation
User Manual
Sentinel is an AI-powered co-pilot for DevOps/SRE engineers. It automatically detects incidents, analyzes them with an AI agent, proposes a corrective action, and asks for your approval before executing anything. You are always in control.
Target audience: DevOps or SRE engineer responsible for responding to production incidents.
- Logging In
- System Configuration & Health Check
- Dashboard — Main View
- Understanding an Incident
- Approving, Rejecting, or Postponing an Action
- Creating an Incident Manually
- Runbooks & Similar Incidents
- Exporting the Post-Mortem
- Quick Reference
When you open Sentinel, you will see the login screen.
Steps:
- Enter your institutional email in the top field.
- Enter your password in the bottom field.
- Click "Iniciar sesión" (Sign in).
If this is your first time using Sentinel, click the "Ver guía de configuración" (Setup guide) link at the bottom of the screen. It will take you to the configuration page where you can verify all services are active before you begin.
The SoftServe logo at the bottom confirms you are on the correct instance of the system.
The configuration page (/setup) has four tabs. Access it from any screen by clicking "Sistema" (System) in the top navigation bar of the Dashboard.
This tab shows the status of all services Sentinel depends on to function.
| Indicator | Meaning |
|---|---|
| ● Green + "Conectado Xms" | The service is responding correctly |
| ● Red + error message | The service is down or unreachable |
Monitored services:
- Prometheus — collects container and server metrics
- Loki — stores and serves logs
- ChromaDB — agent knowledge base (runbooks)
- Alertmanager — receives and routes automatic alerts
- LangFuse — records AI agent traces for observability
- Supabase — database and authentication
If any indicator turns red, Sentinel can still operate in degraded mode, but some features will be limited. Resolve the issue before a critical on-call session.
The header bar shows the overall system status in real time — no need to open this tab to check.
Shows the 5 sequential steps the AI agent follows to process every incident.
| Lab | Function | Approx. Duration |
|---|---|---|
| 1 — Alert Intake | Classifies the incident type using the LLM | ~5s |
| 2 — Investigation | The specialist agent collects evidence with read-only tools | ~15–30s |
| 3 — Decision & Planning | Proposes a corrective action validated against the command whitelist | ~5s |
| 4 — Action & Verification | Executes the approved action and verifies the result | ~10–20s |
| 5 — Post-Incident | Generates the post-mortem and updates the episodic memory | ~5s |
This view is especially useful for explaining to your team how the system works internally.
Historical summary of system performance over the last 200 incidents.
- Total incidents — how many Sentinel has processed
- Resolved — count and percentage resolution rate
- Average MTTR — mean time from detection to action execution
- Active now — ongoing incidents and how many failed
- By severity — visual distribution of incidents by level
- Most frequent type — the most recurring failure type in the system
Checklist to get Sentinel running for the first time:
- Run
docker compose up -dto start the full stack - Verify all indicators in the Integrations tab are green
- Load runbooks into ChromaDB by running
python Backend/scripts/seed_chromadb.py(only once) - Go to the Dashboard — Sentinel will start detecting incidents automatically
The Dashboard is the central operations screen. It is divided into two panels:
Left panel — Incident list:
| Element | Description |
|---|---|
| Active / All / Resolved | Filters to see only what matters right now |
| Color dot | Incident severity (red = critical, orange = high, yellow = medium, green = low) |
| Severity badge |
Critical, High, Medium, Low
|
| Status badge | Current lifecycle state of the incident |
| ⚡ Approve badge | The agent finished its analysis and is waiting for your decision |
| Header counter |
253 active · 2 critical — real-time summary |
Right panel — Incident detail:
- Appears when you select an incident from the list
- Shows "Select an incident" when nothing is selected
Tip: Use the search bar (shortcut
⌘K) to find incidents by name, target, or type without scrolling.
Incidents with the ⚡ Approve badge require your immediate attention — the agent is waiting for your decision to execute the fix.
Clicking an incident in the list opens the full detail view in the right panel.
The top of the panel shows:
- Title of the incident and affected resource (target)
- Severity and status badges — the status color changes as the agent progresses
-
Progress bar — shows where the incident is in the lifecycle:
Detected → Investigating → Analyzed → Approval → Executing → Verifying → Resolved - "Export incident" button — available at any time to download incident data
- Date and time the incident was created
The right panel displays the AI agent's analysis:
-
Agent identity — which specialist handled the incident (
Docker Agent,PostgreSQL Agent, etc.) and its status (ready) -
Classified type — the incident category detected by the LLM (
app_crash,oom,config_error, etc.) - Similar incidents — how many similar past cases were found in historical memory
-
Tools used — which tools the agent ran to gather evidence (
docker_inspect,docker_logs,docker_stats, etc.) - Root cause — the agent's detailed analysis explaining what happened and why
- Status timeline — the current progress in the incident lifecycle
Shows the metrics of the affected container or database at the time of the incident:
- Memory: current usage vs. configured limit, with a color-coded progress bar (green → amber → red)
- CPU: usage percentage
- Network: inbound and outbound traffic in KB/s
- Restarts: number of restarts in the last hour
If real-time metrics are unavailable (Prometheus has no data), Sentinel displays a historical snapshot taken at detection time, labeled "historical data".
This is the most critical moment in the workflow. When the agent finishes its analysis and proposes an action, a prominent banner appears in the center panel.
What the banner shows:
- "ACCIÓN PROPUESTA POR EL AGENTE" — header indicating your decision is required
-
Exact command in monospace format — for example:
docker restart test-e2e-container - Three response options:
| Button | What it does |
|---|---|
| ✗ Reject | Discards the action. The incident moves to failed status. You will need to intervene manually. |
| ⏱ Postpone 30min | Delays execution by 30 minutes. Useful if there is a change window in progress. |
| ✓ Approve and execute | The system executes the command immediately and verifies the result. |
Before approving: read the full command. Make sure the target container or resource is correct. Sentinel only proposes commands from the approved whitelist (
docker restart,docker logs,podman restart,kubectl rollout restart,pg_cancel_backend, etc.) — never destructive commands.
After clicking "Approve and execute", the status changes to
Executing solutionand within seconds you will see the execution result and whether the service recovered.
If you detect a problem that Prometheus has not yet alerted on, you can create an incident manually.
Steps:
- Click the "+ Nuevo" (New) button at the top of the left panel.
- Fill in the form:
-
Title — briefly describe what is happening (e.g.
Payment service not responding) -
Resource type — select
Manual,Container, orDatabase -
Resource — the exact name of the affected container or service (e.g.
nginx-prod,api-gateway) -
Severity —
Critical,High,Medium, orLow - Description / Context — (optional) paste relevant logs or describe the observed behavior
-
Title — briefly describe what is happening (e.g.
- Click "Crear incidente" (Create incident).
The AI agent will automatically start analyzing the incident within seconds. You will see the status change from
DetectedtoInvestigatingin real time.
The Runbooks tab in the right panel shows relevant historical knowledge for the current incident.
Relevant runbooks:
- The agent searches ChromaDB for the runbooks most similar to the detected incident type
- Each runbook shows: name, type, log signals to look for, and resolution steps
- You can search for additional runbooks using the integrated search in the tab
Similar incidents:
- Below the runbooks, historical cases with the same failure pattern are displayed
- They show which solution was applied and whether it was successful
- Useful for confirming that the agent's proposed action matches what worked in the past
The History tab shows the complete timeline of all status changes for the incident with exact timestamps.
When an incident reaches Resolved status, Sentinel automatically generates a post-mortem document.
To export it:
- Select the resolved incident from the list (filter by the Resolved tab).
- In the right panel, click the Post-Mortem tab.
- Review the generated content — you can edit it directly in the text editor.
- Click "Export .md" to download the document in Markdown format.
- Optionally click "Save" to keep the changes in the system.
The post-mortem includes: identified root cause, collected evidence, executed action, verification result, and MTTR (resolution time).
You can also export any incident (resolved or not) by clicking "Export incident" in the panel header — this downloads the raw data in JSON or Markdown format.
| Status | Color | Meaning |
|---|---|---|
detected |
Blue | Sentinel received the alert |
investigating |
Purple | The agent is collecting evidence |
analyzed |
Amber | Analysis is complete |
awaiting_approval |
Cyan | The agent is waiting for your approval to execute |
executing_solution |
Indigo | The command is being executed |
verifying |
Teal | Sentinel verifies the service recovered |
resolved |
Emerald | Service is operational and the incident is closed |
failed |
Red | The action failed or was rejected |
| Level | Color | When to use |
|---|---|---|
critical |
Red | Production service is completely down |
high |
Orange | Severe degradation affecting users |
medium |
Yellow | Issue detected with no immediate user impact |
low |
Green | Preventive alert or minor incident |
| Runtime | Allowed actions |
|---|---|
| Docker |
docker restart <container>, docker logs <container>
|
| Podman |
podman restart <container>, podman logs <container>
|
| Kubernetes |
kubectl rollout restart deployment/<name>, kubectl delete pod <name>, kubectl scale deployment/<name> --replicas=N
|
| PostgreSQL |
pg_cancel_backend <db>, pg_terminate_backend <db>, pg_stat_activity <db>
|
Sentinel will never execute commands outside this list. The system's guardrails block any unauthorized action before it reaches your approval screen.