Repository navigation
Product definition
SoftServe presented the challenge of designing and implementing a DevOps Incident Triage Copilot, aimed at improving how DevOps and Site Reliability Engineering (SRE) teams detect, analyze, and resolve operational incidents.
DevOps and SRE teams face an operational efficiency crisis in incident management that directly impacts service availability and organizational costs.
The average time to identify the root cause of an incident (MTTR - Mean Time To Resolution) is 28-35 minutes according to Datadog (2024), while the average total resolution time reaches 3.8 hours for critical incidents according to Atlassian (2024). Even more concerning, FireHydrant (2023) reports that 43% of this time is spent on repetitive and mechanical tasks: manually reviewing logs, correlating metrics between different tools, and searching for documentation.
The economic impact is considerable. Gartner (2024) estimates that the average cost of downtime is $5,600 USD per minute ($336,000/hour) for business-critical applications. For a team that handles 30-40 incidents per month with an MTTR of 35 minutes, this represents potential losses in the millions.
Sentinel addresses this problem by acting as an intelligent AI-powered co-pilot that assists engineers throughout the entire incident lifecycle. Rather than relying on manually provided inputs or isolated chat interactions, Sentinel integrates directly with the execution environment and observability sources to consume real execution data generated by deployed services. Using this data, the system automatically detects incidents, collects evidence, correlates signals, and guides the analysis through a structured, multi-stage triage process.
Sentinel organizes incident handling into clearly defined phases, maintaining full context across detection, investigation, decision-making, verification, and post-incident analysis. An agentic architecture enables the system to reason over logs, metrics, events, and historical knowledge, while always keeping a human-in-the-loop for validation of critical actions.
All information is presented through a real-time React-based dashboard, where engineers can monitor active incidents, inspect collected evidence, review the agent’s reasoning, visualize incident timelines, and approve or reject proposed remediation steps. Sentinel is not a conversational interface where errors are copied and pasted; it is an integrated operational platform that functions as a continuously available, context-aware SRE assistant embedded directly into the incident workflow.
| Name | GitHub Profile | Role | |
|---|---|---|---|
| Leon Daniel Jaramillo Ramirez | leonjaramillo | ljaram@softserveinc.com | Product Owner |
| Nicolas Rico Montesino | nicolas344 | nricom@eafit.edu.co | Scrum Master, Developer |
| Santiago Alvares Diaz | sxntiagoad | salvarezd3@eafit.edu.co | Developer |
| Jacobo Montes Castaño | jacoEafit | jmontesca@eafit.edu.co | UX/UI |
| Nicol Franchesca Garcia Tabares | nfgarciat | nfgarciat@eafit.edu.co | Tester |
| Thomas Osorio Zambrano | thomaszambrano | tosorioz@eafit.edu.co | Developer |
| Alejandro Rendon Correa | arendonc2 | arendonc2@eafit.edu.co | Architect |
DevOps / Platform Engineers: DevOps engineers are responsible for both responding to production incidents and improving system reliability over time. They use Sentinel to manage incidents triggered by alerts, investigate root causes by reviewing metrics and logs, and monitor the incident timeline through a real-time dashboard.
Additionally, DevOps engineers review the agent’s reasoning and collected evidence, and approve or reject the actions proposed by Sentinel to ensure safe and effective incident resolution. After incidents are resolved, they analyze historical data and recurring issues to improve observability, update runbooks, and reduce the likelihood of future incidents.

The system is designed for a single user type: the DevOps Engineer, who interacts with the platform to manage incidents in a structured and assisted manner. The user flow begins when an incident is received or selected and continues seamlessly through the different system Labs: Alert Intake, Investigation, Decision & Planning, Action & Verification, and Post-Incident. Throughout each Lab, the user remains in control while being supported by a set of domain-specific DevOps agents, coordinated through contextual isolation that separates information by sector (Cloud, CI/CD, Observability, Security, etc.). This approach ensures a smooth and uninterrupted experience, allowing incidents to progress end-to-end within a single, coherent workflow.

| Term | Definition |
|---|---|
| Agentic AI Copilot | An AI assistant that collaborates with the user, proactively executes tasks, and supports decision-making without replacing the human. |
| DevOps Engineer | An engineer responsible for maintaining system reliability, handling incidents, and ensuring critical services remain operational. |
| Triaging Incidents | The process of categorizing and prioritizing incidents to determine urgency and assign the appropriate owner. |
| Stream of Alerts and Logs | A continuous flow of alerts and logs that represent the state and behavior of an application. |
| LangGraph | A Python library built on the LangChain ecosystem for creating multi-step, self-correcting AI agents. |
| ChromaDB | A vector database used to store and retrieve information using embeddings. |
| Embeddings | Numerical representations of data (such as text) that enable semantic similarity search. |
| Observability | The ability to understand the internal state of a system by analyzing metrics, logs, and traces to detect and diagnose issues. |
| Domain-Specific Agents | Specialized AI agents focused on specific technical domains (e.g., Cloud, CI/CD, Observability, Security) that assist users during incident management. |
| Incident | An unplanned event that degrades or interrupts a system’s normal operation and requires action to restore service. |
| Lab | The current working phase of an incident. |
| Id | Description |
|---|---|
| FR-01 | The system must automatically create an incident when an anomalous event is detected in the execution of a service. |
| FR-02 | The system must represent each incident as a single entity that centralizes its status, context, evidence, decisions, and actions. |
| FR-03 | The system must allow visualization of the current incident status, including its phase (Lab), severity, and priority. |
| FR-04 | The system must allow querying past incidents with their full information for subsequent analysis. |
| FR-05 | The system must manage the incident's progress through the defined Labs, ensuring a controlled sequential flow. |
| FR-06 | During the Alert Intake Lab, the system must automatically classify the incident by severity, priority, and technical context. |
| FR-07 | During the Investigation Lab, the system must collect relevant evidence of the incident, including logs, metrics, and events. |
| FR-08 | During the Decision & Planning Lab, the system must analyze the evidence and propose one or more resolution actions. |
| FR-09 | During the Action & Verification Lab, the system must allow executing or simulating actions and validating their impact before marking the incident as resolved. |
| FR-10 | During the Post-Incident Lab, the system must generate a final analysis of the incident, including a timeline and root cause analysis (RCA). |
| FR-11 | The system must associate one or more technical contexts with each incident to organize the analysis. |
| FR-12 | The system must activate only the relevant specialized agents based on the incident's contexts. |
| FR-13 | The system must coordinate the interaction between specialized agents through an orchestrator agent. |
| FR-14 | The system must maintain the incident analysis state throughout its entire lifecycle. |
| FR-15 | The system must allow visualization of the reasoning followed by the agents during the incident analysis. |
| FR-16 | The system must consult runbooks, documentation, and previous cases to support decision-making. |
| FR-17 | The system must record learnings derived from resolved incidents to improve future analyses. |
| FR-18 | The system must allow searching for previously recorded similar incidents. |
| FR-19 | The system must provide a centralized view of incidents without exposing internal complexity. |
| FR-20 | The system must notify the user when a critical incident is detected or changes its status. |
| FR-21 | The system must allow the user to interact with the agent to perform queries about the incident. |
| FR-22 | The system must allow exporting relevant incident information for external analysis or auditing. |
| FR-23 | The system must allow users to log in using their credentials to access the platform securely. |
| FR-24 | The system must allow authenticated users to log out, terminating their active session immediately. |
For the elicitation of the functional requirements described above, the team held several in-class meetings to discuss and define the expected behavior of the system. During these sessions, we collaboratively explored what features and capabilities the system should provide, based on the project description and the problem context.
We conducted a brainstorming process where each team member proposed ideas related to the functionality of the system. These ideas were discussed collectively, evaluated in terms of feasibility and relevance, and refined through group consensus. Only the ideas that aligned with the project objectives and were agreed upon by the team were selected and documented as functional requirements.

As part of the architectural design for our Agentic DevOps Incident Triage Copilot, we conducted a competitive analysis of three industry-leading solutions. This benchmark allows us to position our project within the AIOps ecosystem, justifying our choice of LangGraph, Gemini, and ChromaDB.
Access link: PagerDuty
PagerDuty is the industry standard for incident response. Their platform leverages machine learning to reduce noise and provide "Event Intelligence."
Similar Functionalities: Automated incident classification, noise reduction, and "Past Incidents" context to help responders understand previous similar failures.
Comparison to our Project: While PagerDuty excels at alerting, our project focuses more on the autonomous reasoning layer. We are using LangGraph to build a specific diagnostic flow that mimics a Senior Engineer's decision-making process, rather than just aggregating events.

Access link: Kubiya.ai
Kubiya is an agentic platform that provides a "Virtual Assistant for DevOps." It integrates with Slack to help developers perform self-service operations.
Similar Functionalities: It uses an agentic framework to understand intent, call tools, and retrieve information from technical documentation.
Comparison to our Project: Kubiya is a SaaS (Software as a Service) product. Our project provides a customizable framework where we can define the exact graph of nodes and conditional edges. This level of granular control over the "Triage Logic" is the core value proposition of our university project.


By the end of Sprint 1 , the system must be able to:
Authentication: User can log in and log out Detection: The system automatically detects when a container fails Dashboard: User can view the list of active incidents and their status Classification: The agent classifies the incident by type and severity (Alert Intake Lab) Evidence: The agent automatically collects logs and metrics (Investigation Lab) Transparency: User can see the agent’s reasoning in real time Notifications: User receives alerts for critical incidents Functional demo: A real Docker container incident that is detected, classified, and analyzed automatically in under 2 minutes, with full visualization of the process.