Skip to content

Architecture Design

Nicolas Rico edited this page Mar 13, 2026 · 21 revisions

Table of Contents


1. Scope

Sentinel's MVP focuses on demonstrating the complete incident resolution flow using agentic AI, from automatic detection to learning documentation, in a controlled development environment. Functional Scope of the MVP: Includes:

  • Monitoring of Docker containers as the primary source of incidents
  • All 5 Labs: Alert Intake, Investigation, Decision & Planning, Action & Verification, and Post-Incident
  • Active specialized agents: Cloud Agent, Observability Agent, and Knowledge Agent coordinated by the Orchestrator
  • Initial knowledge base: runbooks embedded in ChromaDB
  • Processing capacity: 1-5 simultaneous incidents, storage of up to 200 historical incidents
  • Functional web dashboard with real-time visualization of agent reasoning
  • Guided execution of solutions with human approval for critical actions

2. System Dimensions and Non-Functional Requirements

The following Non-Functional Requirements define the operational constraints and quality attributes of Sentinel. They describe how the system should perform under expected conditions, focusing on performance, availability, usability, scalability, and cost efficiency. These requirements are aligned with the academic scope of the project and the targeted MVP, ensuring that Sentinel remains functional, realistic, and demonstrable within limited resources.

ID Requirement Description
NFR-01 The system must load the main dashboard in less than 2 seconds in 95% of cases, supporting up to 3 concurrent users.
NFR-02 Critical operations such as incident detection, initial classification, and alert notification must complete in less than 15 seconds under normal load.
NFR-03 Sentinel must support 1 to 3 concurrent users without noticeable performance degradation.
NFR-04 The system must handle up to 10 HTTP requests per second and 5 WebSocket messages per second without data loss.
NFR-05 The user interface must allow users to review an active incident and understand its status in a maximum of 6 interactions.
NFR-06 The system must maintain at least 95% availability during academic usage hours and project demonstrations.
NFR-07 Incident data and agent state must persist across system restarts, allowing recovery from the last saved checkpoint.
NFR-08 The system must manage up to 200 stored incidents and 5 active incidents simultaneously without impacting performance.
NFR-09 Sentinel must function correctly on modern browsers (Chrome, Edge, Firefox, Safari) without requiring additional configuration.
NFR-10 The system must operate in a low-cost environment, supporting full local execution and limiting LLM usage to a maximum of 5 calls per minute.

3. Domain Modeling

The SENTINEL domain model defines the core entities and relationships for intelligent DevOps incident management. At its center is the Incident entity, which progresses sequentially through five Labs (phases): Alert Intake, Investigation, Decision & Planning, Action & Verification, and Post-Incident. Each incident is classified by one or more Contexts (Cloud, CI/CD, Observability, Security, Knowledge), which automatically activate specialized Agents to collect Evidence, propose Hypotheses, and create ActionPlans. The model enforces critical business rules, such as requiring every incident to have at least one technical context, ensuring hypotheses are evidence-based, and mandating that high-risk action plans receive user approval before execution. This structure ensures a controlled, traceable, and systematic approach to incident resolution while enabling multi-agent coordination across different DevOps domains.

SENTINEL DIAG (1)

4. Component Description

Technical Description of Components

The architecture follows a decoupled, service-oriented design to ensure scalability and reliability in DevOps environments. Each layer plays a critical role in the automated triage lifecycle:

A. User Interface Layer (React)

The frontend is built using React, providing a high-performance, reactive dashboard. It serves as the primary interface for DevOps engineers to visualize active incidents, review agent reasoning, and monitor the triage lifecycle in real time.

This layer allows DevOps engineers to:

  • Visualize ongoing incidents and their current phase.

  • Inspect the agent's reasoning process, collected evidence, and recommended actions.

  • Monitor incident timelines and resolution progress in real time.

The frontend communicates with the backend through the API Gateway and receives real-time updates via Supabase Realtime (WebSocket), ensuring the dashboard reflects the latest incident state without manual refresh.

B. Orchestration & Logic Layer (FastAPI & LangGraph)

This layer acts as the system's "nervous system", responsible for managing workflows, state, and communication between components.

FastAPI: Managed as the backend gateway, it provides asynchronous endpoints to handle concurrent triage requests with low latency. It manages incident creation, user interactions, and real-time data exchange. Webhook alerts from Alertmanager are received here and dispatched as background tasks to avoid blocking the HTTP response.

LangGraph: Unlike linear pipelines, LangGraph enables a stateful, cyclic workflow. It manages the AgentState, allowing the system to maintain context as it transitions between classification, research, and resolution nodes. This engine coordinates calls to the AI model, knowledge sources, and monitoring tools, ensuring that each step of the incident resolution process is executed in the correct order and context.

C. AI & Knowledge Layer (OpenAI gpt-4o-mini & ChromaDB)

The AI & Knowledge Layer provides the cognitive capabilities required for automated reasoning and informed decision-making:

OpenAI gpt-4o-mini (Sprint 1): Selected for its low cost and fast inference, making it suitable for academic environments with limited LLM budget. It classifies incident types (Lab 1) and generates root cause analysis with remediation recommendations (Lab 2). The model is configurable via environment variable, allowing a future upgrade to Gemini 2.5 Flash-Lite (Sprint 2) without code changes.

ChromaDB: A high-performance vector database that facilitates Retrieval-Augmented Generation (RAG). It stores vectorized runbooks covering the five incident types (OOM, app crash, config error, dependency failure, unknown), ensuring that the agent's suggestions are grounded in validated operational knowledge rather than general model training alone.

This layer ensures that Sentinel's recommendations are both intelligent and contextually accurate.

D. Observability & Monitoring (LangFuse — self-hosted)

To maintain production-grade standards, LangFuse (self-hosted, v2) provides comprehensive observability for the LangGraph agent. It captures execution traces across every node in the LangGraph workflow, allowing developers to monitor model latency, token usage, and reasoning traces for debugging and auditing purposes.

LangFuse is deployed as a self-hosted instance (port 3001) within the same Docker Compose stack as the rest of the infrastructure. This decision was made based on the Product Owner's recommendation (Leon Jaramillo — SoftServe) for compliance with future client requirements. LangFuse replaces LangSmith as the observability backend.

E. Detection Infrastructure (Prometheus + Alertmanager + Loki + cAdvisor)

Sentinel uses an industry-standard observability stack for automatic incident detection:

  • cAdvisor: Scrapes real-time metrics from all running Docker containers and exposes them to Prometheus.
  • Prometheus: Evaluates alert rules (e.g., container_last_seen stale for > 20 seconds) and triggers alerts.
  • Alertmanager: Routes triggered alerts via webhook to the FastAPI backend (POST /api/alerts).
  • Loki + Promtail: Collects and stores container logs. When an alert fires, the backend queries Loki to retrieve the container's logs from the 5-minute window before the crash and attaches them to the incident.

F. Architectural Summary

In summary, Sentinel's component-based architecture ensures:

  • Clear separation of concerns between interface, orchestration, intelligence, and monitoring.

  • Scalable and maintainable system design.

  • Transparent and auditable AI-driven decision-making.

This design enables Sentinel to function as an intelligent DevOps incident triage copilot while maintaining a simple and intuitive user experience.

G. Technologies Used

Component Technology Sprint 1 Notes
Frontend React 19 + Vite + Tailwind CSS ✅ Real-time dashboard
Backend FastAPI + Uvicorn ✅ API Gateway, async
Agent Orchestration LangGraph ✅ Labs 1 & 2
LLM OpenAI gpt-4o-mini ✅ Sprint 1; upgradeable via .env
Vector DB ChromaDB ✅ RAG over runbooks
Auth & Realtime DB Supabase ✅ JWT auth + WebSocket Realtime
Agent Observability LangFuse v2 (self-hosted) ✅ Replaces LangSmith
Container Metrics cAdvisor ✅ Docker metrics scraping
Metrics & Alerting Prometheus + Alertmanager ✅ Alert rules + webhook
Log Storage Loki + Promtail ✅ Container log collection
Dashboards Grafana ✅ Metrics + log visualization

5. System Component Diagram

graph TD
    subgraph Detection_Infrastructure
        CA[cAdvisor :8080]
        PR[Prometheus :9090]
        AM[Alertmanager :9093]
        LK[Loki :3100]
        PT[Promtail]
    end

    subgraph Frontend_Layer
        UI[React Dashboard :5173]
    end

    subgraph Backend_Orchestration
        API[FastAPI :8000]
        LG[LangGraph Engine]
    end

    subgraph AI_Knowledge
        LLM[OpenAI gpt-4o-mini]
        DB[ChromaDB :8001]
    end

    subgraph Observability
        LF[LangFuse self-hosted :3001]
        GF[Grafana :3000]
    end

    subgraph Database
        SB[Supabase DB + Realtime]
    end

    CA --> PR
    PT --> LK
    PR --> AM
    AM -->|webhook POST /api/alerts| API
    API --> LG
    API -->|query logs| LK
    API --> SB
    LG <-->|Lab 1 & 2| LLM
    LG <-->|RAG runbooks| DB
    LG -->|traces| LF
    LG --> SB
    SB -->|Realtime WebSocket| UI
    UI --> API
    PR --> GF
    LK --> GF
Loading

Clone this wiki locally