Skip to content

User Manual

Nicolas Rico edited this page May 20, 2026 · 7 revisions

User Manual — Sentinel

Sentinel is an AI-powered co-pilot for DevOps/SRE engineers. It automatically detects incidents, analyzes them with an AI agent, proposes a corrective action, and asks for your approval before executing anything. You are always in control.

Target audience: DevOps or SRE engineer responsible for responding to production incidents.


Table of Contents

  1. Logging In
  2. System Configuration & Health Check
  3. Dashboard — Main View
  4. Understanding an Incident
  5. Approving, Rejecting, or Postponing an Action
  6. Creating an Incident Manually
  7. Runbooks & Similar Incidents
  8. Exporting the Post-Mortem
  9. Quick Reference

1. Logging In

2026-05-20_09-16-17

When you open Sentinel, you will see the login screen.

Steps:

  1. Enter your institutional email in the top field.
  2. Enter your password in the bottom field.
  3. Click "Iniciar sesión" (Sign in).

If this is your first time using Sentinel, click the "Ver guía de configuración" (Setup guide) link at the bottom of the screen. It will take you to the configuration page where you can verify all services are active before you begin.

The SoftServe logo at the bottom confirms you are on the correct instance of the system.


2. System Configuration & Health Check

The configuration page (/setup) has four tabs. Access it from any screen by clicking "Sistema" (System) in the top navigation bar of the Dashboard.

2.1 Integrations Tab

image

This tab shows the status of all services Sentinel depends on to function.

Indicator Meaning
● Green + "Conectado Xms" The service is responding correctly
● Red + error message The service is down or unreachable

Monitored services:

  • Prometheus — collects container and server metrics
  • Loki — stores and serves logs
  • ChromaDB — agent knowledge base (runbooks)
  • Alertmanager — receives and routes automatic alerts
  • LangFuse — records AI agent traces for observability
  • Supabase — database and authentication

If any indicator turns red, Sentinel can still operate in degraded mode, but some features will be limited. Resolve the issue before a critical on-call session.

The header bar shows the overall system status in real time — no need to open this tab to check.

2.2 Agent Labs Tab

image

Shows the 5 sequential steps the AI agent follows to process every incident.

Lab Function Approx. Duration
1 — Alert Intake Classifies the incident type using the LLM ~5s
2 — Investigation The specialist agent collects evidence with read-only tools ~15–30s
3 — Decision & Planning Proposes a corrective action validated against the command whitelist ~5s
4 — Action & Verification Executes the approved action and verifies the result ~10–20s
5 — Post-Incident Generates the post-mortem and updates the episodic memory ~5s

This view is especially useful for explaining to your team how the system works internally.

2.3 Metrics Tab

image

Historical summary of system performance over the last 200 incidents.

  • Total incidents — how many Sentinel has processed
  • Resolved — count and percentage resolution rate
  • Average MTTR — mean time from detection to action execution
  • Active now — ongoing incidents and how many failed
  • By severity — visual distribution of incidents by level
  • Most frequent type — the most recurring failure type in the system

2.4 Getting Started Tab

image

Checklist to get Sentinel running for the first time:

  1. Run docker compose up -d to start the full stack
  2. Verify all indicators in the Integrations tab are green
  3. Load runbooks into ChromaDB by running python Backend/scripts/seed_chromadb.py (only once)
  4. Go to the Dashboard — Sentinel will start detecting incidents automatically

3. Dashboard — Main View

image

The Dashboard is the central operations screen. It is divided into two panels:

Left panel — Incident list:

Element Description
Active / All / Resolved Filters to see only what matters right now
Color dot Incident severity (red = critical, orange = high, yellow = medium, green = low)
Severity badge Critical, High, Medium, Low
Status badge Current lifecycle state of the incident
⚡ Approve badge The agent finished its analysis and is waiting for your decision
Header counter 253 active · 2 critical — real-time summary

Right panel — Incident detail:

  • Appears when you select an incident from the list
  • Shows "Select an incident" when nothing is selected

Tip: Use the search bar (shortcut ⌘K) to find incidents by name, target, or type without scrolling.

Incidents with the ⚡ Approve badge require your immediate attention — the agent is waiting for your decision to execute the fix.


4. Understanding an Incident

Clicking an incident in the list opens the full detail view in the right panel.

4.1 Incident Header

image

The top of the panel shows:

  1. Title of the incident and affected resource (target)
  2. Severity and status badges — the status color changes as the agent progresses
  3. Progress bar — shows where the incident is in the lifecycle: Detected → Investigating → Analyzed → Approval → Executing → Verifying → Resolved
  4. "Export incident" button — available at any time to download incident data
  5. Date and time the incident was created

4.2 Agent Panel (Agent tab)

image

The right panel displays the AI agent's analysis:

  • Agent identity — which specialist handled the incident (Docker Agent, PostgreSQL Agent, etc.) and its status (ready)
  • Classified type — the incident category detected by the LLM (app_crash, oom, config_error, etc.)
  • Similar incidents — how many similar past cases were found in historical memory
  • Tools used — which tools the agent ran to gather evidence (docker_inspect, docker_logs, docker_stats, etc.)
  • Root cause — the agent's detailed analysis explaining what happened and why
  • Status timeline — the current progress in the incident lifecycle

4.3 Metrics Tab

image

Shows the metrics of the affected container or database at the time of the incident:

  • Memory: current usage vs. configured limit, with a color-coded progress bar (green → amber → red)
  • CPU: usage percentage
  • Network: inbound and outbound traffic in KB/s
  • Restarts: number of restarts in the last hour

If real-time metrics are unavailable (Prometheus has no data), Sentinel displays a historical snapshot taken at detection time, labeled "historical data".


5. Approving, Rejecting, or Postponing an Action

This is the most critical moment in the workflow. When the agent finishes its analysis and proposes an action, a prominent banner appears in the center panel.

2026-05-20_09-23-16

What the banner shows:

  1. "ACCIÓN PROPUESTA POR EL AGENTE" — header indicating your decision is required
  2. Exact command in monospace format — for example: docker restart test-e2e-container
  3. Three response options:
Button What it does
✗ Reject Discards the action. The incident moves to failed status. You will need to intervene manually.
⏱ Postpone 30min Delays execution by 30 minutes. Useful if there is a change window in progress.
✓ Approve and execute The system executes the command immediately and verifies the result.

Before approving: read the full command. Make sure the target container or resource is correct. Sentinel only proposes commands from the approved whitelist (docker restart, docker logs, podman restart, kubectl rollout restart, pg_cancel_backend, etc.) — never destructive commands.

After clicking "Approve and execute", the status changes to Executing solution and within seconds you will see the execution result and whether the service recovered.


6. Creating an Incident Manually

image

If you detect a problem that Prometheus has not yet alerted on, you can create an incident manually.

Steps:

  1. Click the "+ Nuevo" (New) button at the top of the left panel.
  2. Fill in the form:
    • Title — briefly describe what is happening (e.g. Payment service not responding)
    • Resource type — select Manual, Container, or Database
    • Resource — the exact name of the affected container or service (e.g. nginx-prod, api-gateway)
    • Severity — Critical, High, Medium, or Low
    • Description / Context — (optional) paste relevant logs or describe the observed behavior
  3. Click "Crear incidente" (Create incident).

The AI agent will automatically start analyzing the incident within seconds. You will see the status change from Detected to Investigating in real time.


7. Runbooks & Similar Incidents

image

The Runbooks tab in the right panel shows relevant historical knowledge for the current incident.

Relevant runbooks:

  • The agent searches ChromaDB for the runbooks most similar to the detected incident type
  • Each runbook shows: name, type, log signals to look for, and resolution steps
  • You can search for additional runbooks using the integrated search in the tab

Similar incidents:

  • Below the runbooks, historical cases with the same failure pattern are displayed
  • They show which solution was applied and whether it was successful
  • Useful for confirming that the agent's proposed action matches what worked in the past

The History tab shows the complete timeline of all status changes for the incident with exact timestamps.


8. Exporting the Post-Mortem

image

When an incident reaches Resolved status, Sentinel automatically generates a post-mortem document.

To export it:

  1. Select the resolved incident from the list (filter by the Resolved tab).
  2. In the right panel, click the Post-Mortem tab.
  3. Review the generated content — you can edit it directly in the text editor.
  4. Click "Export .md" to download the document in Markdown format.
  5. Optionally click "Save" to keep the changes in the system.

The post-mortem includes: identified root cause, collected evidence, executed action, verification result, and MTTR (resolution time).

You can also export any incident (resolved or not) by clicking "Export incident" in the panel header — this downloads the raw data in JSON or Markdown format.


9. Quick Reference

Incident Statuses

Status Color Meaning
detected Blue Sentinel received the alert
investigating Purple The agent is collecting evidence
analyzed Amber Analysis is complete
awaiting_approval Cyan The agent is waiting for your approval to execute
executing_solution Indigo The command is being executed
verifying Teal Sentinel verifies the service recovered
resolved Emerald Service is operational and the incident is closed
failed Red The action failed or was rejected

Severity Levels

Level Color When to use
critical Red Production service is completely down
high Orange Severe degradation affecting users
medium Yellow Issue detected with no immediate user impact
low Green Preventive alert or minor incident

Available Agent Actions

Runtime Allowed actions
Docker docker restart <container>, docker logs <container>
Podman podman restart <container>, podman logs <container>
Kubernetes kubectl rollout restart deployment/<name>, kubectl delete pod <name>, kubectl scale deployment/<name> --replicas=N
PostgreSQL pg_cancel_backend <db>, pg_terminate_backend <db>, pg_stat_activity <db>

Sentinel will never execute commands outside this list. The system's guardrails block any unauthorized action before it reaches your approval screen.


Clone this wiki locally