Repository navigation
Architecture Design
- 1. Scope
- 2. System Dimensions and Non-Functional Requirements
- 3. Domain Modeling
- 4. Component Description
- 5. System Component Diagram
Sentinel's MVP focuses on demonstrating the complete incident resolution flow using agentic AI, from automatic detection to learning documentation, in a controlled development environment. Functional Scope of the MVP: Includes:
- Monitoring of Docker containers as the primary source of incidents
- All 5 Labs: Alert Intake, Investigation, Decision & Planning, Action & Verification, and Post-Incident
- Active specialized agents: Cloud Agent, Observability Agent, and Knowledge Agent coordinated by the Orchestrator
- Initial knowledge base: runbooks embedded in ChromaDB
- Processing capacity: 1-5 simultaneous incidents, storage of up to 200 historical incidents
- Functional web dashboard with real-time visualization of agent reasoning
- Guided execution of solutions with human approval for critical actions
The following Non-Functional Requirements define the operational constraints and quality attributes of Sentinel. They describe how the system should perform under expected conditions, focusing on performance, availability, usability, scalability, and cost efficiency. These requirements are aligned with the academic scope of the project and the targeted MVP, ensuring that Sentinel remains functional, realistic, and demonstrable within limited resources.
| ID | Requirement Description |
|---|---|
| NFR-01 | The system must load the main dashboard in less than 2 seconds in 95% of cases, supporting up to 3 concurrent users. |
| NFR-02 | Critical operations such as incident detection, initial classification, and alert notification must complete in less than 15 seconds under normal load. |
| NFR-03 | Sentinel must support 1 to 3 concurrent users without noticeable performance degradation. |
| NFR-04 | The system must handle up to 10 HTTP requests per second and 5 WebSocket messages per second without data loss. |
| NFR-05 | The user interface must allow users to review an active incident and understand its status in a maximum of 6 interactions. |
| NFR-06 | The system must maintain at least 95% availability during academic usage hours and project demonstrations. |
| NFR-07 | Incident data and agent state must persist across system restarts, allowing recovery from the last saved checkpoint. |
| NFR-08 | The system must manage up to 200 stored incidents and 5 active incidents simultaneously without impacting performance. |
| NFR-09 | Sentinel must function correctly on modern browsers (Chrome, Edge, Firefox, Safari) without requiring additional configuration. |
| NFR-10 | The system must operate in a low-cost environment, supporting full local execution and limiting LLM usage to a maximum of 5 calls per minute. |
The SENTINEL domain model defines the core entities and relationships for intelligent DevOps incident management. At its center is the Incident entity, which progresses sequentially through five Labs (phases): Alert Intake, Investigation, Decision & Planning, Action & Verification, and Post-Incident. Each incident is classified by one or more Contexts (Cloud, CI/CD, Observability, Security, Knowledge), which automatically activate specialized Agents to collect Evidence, propose Hypotheses, and create ActionPlans. The model enforces critical business rules, such as requiring every incident to have at least one technical context, ensuring hypotheses are evidence-based, and mandating that high-risk action plans receive user approval before execution. This structure ensures a controlled, traceable, and systematic approach to incident resolution while enabling multi-agent coordination across different DevOps domains.
Technical Description of Components
The architecture follows a decoupled, service-oriented design to ensure scalability and reliability in DevOps environments. Each layer plays a critical role in the automated triage lifecycle:
A. User Interface Layer (React)
The frontend is built using React, providing a high-performance, reactive dashboard. It serves as the primary interface for DevOps engineers to visualize active incidents, review agent reasoning, and monitor the triage lifecycle in real time.
This layer allows DevOps engineers to:
-
Visualize ongoing incidents and their current phase.
-
Inspect the agent's reasoning process, collected evidence, and recommended actions.
-
Monitor incident timelines and resolution progress in real time.
The frontend communicates with the backend through the API Gateway and receives real-time updates via Supabase Realtime (WebSocket), ensuring the dashboard reflects the latest incident state without manual refresh.
B. Orchestration & Logic Layer (FastAPI & LangGraph)
This layer acts as the system's "nervous system", responsible for managing workflows, state, and communication between components.
FastAPI: Managed as the backend gateway, it provides asynchronous endpoints to handle concurrent triage requests with low latency. It manages incident creation, user interactions, and real-time data exchange. Webhook alerts from Alertmanager are received here and dispatched as background tasks to avoid blocking the HTTP response.
LangGraph: Unlike linear pipelines, LangGraph enables a stateful, cyclic workflow. It manages the AgentState, allowing the system to maintain context as it transitions between classification, research, and resolution nodes. This engine coordinates calls to the AI model, knowledge sources, and monitoring tools, ensuring that each step of the incident resolution process is executed in the correct order and context.
C. AI & Knowledge Layer (OpenAI gpt-4o-mini & ChromaDB)
The AI & Knowledge Layer provides the cognitive capabilities required for automated reasoning and informed decision-making:
OpenAI gpt-4o-mini (Sprint 1): Selected for its low cost and fast inference, making it suitable for academic environments with limited LLM budget. It classifies incident types (Lab 1) and generates root cause analysis with remediation recommendations (Lab 2). The model is configurable via environment variable, allowing a future upgrade to Gemini 2.5 Flash-Lite (Sprint 2) without code changes.
ChromaDB: A high-performance vector database that facilitates Retrieval-Augmented Generation (RAG). It stores vectorized runbooks covering the five incident types (OOM, app crash, config error, dependency failure, unknown), ensuring that the agent's suggestions are grounded in validated operational knowledge rather than general model training alone.
This layer ensures that Sentinel's recommendations are both intelligent and contextually accurate.
D. Observability & Monitoring (LangFuse — self-hosted)
To maintain production-grade standards, LangFuse (self-hosted, v2) provides comprehensive observability for the LangGraph agent. It captures execution traces across every node in the LangGraph workflow, allowing developers to monitor model latency, token usage, and reasoning traces for debugging and auditing purposes.
LangFuse is deployed as a self-hosted instance (port 3001) within the same Docker Compose stack as the rest of the infrastructure. This decision was made based on the Product Owner's recommendation (Leon Jaramillo — SoftServe) for compliance with future client requirements. LangFuse replaces LangSmith as the observability backend.
E. Detection Infrastructure (Prometheus + Alertmanager + Loki + cAdvisor)
Sentinel uses an industry-standard observability stack for automatic incident detection:
- cAdvisor: Scrapes real-time metrics from all running Docker containers and exposes them to Prometheus.
-
Prometheus: Evaluates alert rules (e.g.,
container_last_seenstale for > 20 seconds) and triggers alerts. -
Alertmanager: Routes triggered alerts via webhook to the FastAPI backend (
POST /api/alerts). - Loki + Promtail: Collects and stores container logs. When an alert fires, the backend queries Loki to retrieve the container's logs from the 5-minute window before the crash and attaches them to the incident.
F. Architectural Summary
In summary, Sentinel's component-based architecture ensures:
-
Clear separation of concerns between interface, orchestration, intelligence, and monitoring.
-
Scalable and maintainable system design.
-
Transparent and auditable AI-driven decision-making.
This design enables Sentinel to function as an intelligent DevOps incident triage copilot while maintaining a simple and intuitive user experience.
G. Technologies Used
| Component | Technology | Sprint 1 | Notes |
|---|---|---|---|
| Frontend | React 19 + Vite + Tailwind CSS | ✅ | Real-time dashboard |
| Backend | FastAPI + Uvicorn | ✅ | API Gateway, async |
| Agent Orchestration | LangGraph | ✅ | Labs 1 & 2 |
| LLM | OpenAI gpt-4o-mini | ✅ | Sprint 1; upgradeable via .env
|
| Vector DB | ChromaDB | ✅ | RAG over runbooks |
| Auth & Realtime DB | Supabase | ✅ | JWT auth + WebSocket Realtime |
| Agent Observability | LangFuse v2 (self-hosted) | ✅ | Replaces LangSmith |
| Container Metrics | cAdvisor | ✅ | Docker metrics scraping |
| Metrics & Alerting | Prometheus + Alertmanager | ✅ | Alert rules + webhook |
| Log Storage | Loki + Promtail | ✅ | Container log collection |
| Dashboards | Grafana | ✅ | Metrics + log visualization |
graph TD
subgraph Detection_Infrastructure
CA[cAdvisor :8080]
PR[Prometheus :9090]
AM[Alertmanager :9093]
LK[Loki :3100]
PT[Promtail]
end
subgraph Frontend_Layer
UI[React Dashboard :5173]
end
subgraph Backend_Orchestration
API[FastAPI :8000]
LG[LangGraph Engine]
end
subgraph AI_Knowledge
LLM[OpenAI gpt-4o-mini]
DB[ChromaDB :8001]
end
subgraph Observability
LF[LangFuse self-hosted :3001]
GF[Grafana :3000]
end
subgraph Database
SB[Supabase DB + Realtime]
end
CA --> PR
PT --> LK
PR --> AM
AM -->|webhook POST /api/alerts| API
API --> LG
API -->|query logs| LK
API --> SB
LG <-->|Lab 1 & 2| LLM
LG <-->|RAG runbooks| DB
LG -->|traces| LF
LG --> SB
SB -->|Realtime WebSocket| UI
UI --> API
PR --> GF
LK --> GF