Repository navigation
Project Viability
Sentinel is technically and operationally viable because it is built on technologies that are available, well-documented, and aligned with the needs of a DevOps incident response environment. The project uses a realistic and well-defined stack: React 19, Vite 7, and Tailwind CSS v4 for the frontend; FastAPI and Uvicorn for the backend; LangGraph for agent orchestration; OpenAI gpt-4o-mini for AI reasoning; ChromaDB for RAG over runbooks; Supabase for authentication, persistent storage, and real-time updates; LangFuse v2 self-hosted for agent observability; and cAdvisor, Prometheus, Alertmanager, Loki, Promtail, and Grafana for incident detection, monitoring, logs, and technical dashboards. The system is executed through Docker Compose, which helps integrate the infrastructure services in a controlled and reproducible environment.
From a technical perspective, the project is feasible because each component has a clear responsibility within the architecture. The frontend provides a real-time dashboard for DevOps engineers, the backend coordinates requests and system logic, the agent workflow manages the incident lifecycle, and the knowledge base supports recommendations using runbooks and historical information. This separation of responsibilities makes the system easier to maintain, test, and improve over time.
From an operational perspective, Sentinel is also viable because it addresses a real need in DevOps teams: reducing the time and effort required to analyze and respond to incidents. The system centralizes alerts, evidence, reasoning, recommendations, and incident history in one place, reducing context switching between tools. It also keeps the human-in-the-loop for critical actions, which makes the solution safer and more trustworthy for production environments.
The project has some operational risks, such as dependency on external AI services, possible integration failures, and the need to keep runbooks updated. However, these risks can be managed through mitigation strategies such as fallback modes, human approval, monitoring, secure credential management, and continuous testing.
Overall, Sentinel is viable because it combines a clear business need, a realistic technical architecture, and a strong value proposition for DevOps teams.
Sentinel is a viable project because its benefits justify its development and operational effort. The solution helps reduce manual work, improves incident response time, increases transparency in technical decision-making, and supports long-term system reliability. As detailed in the cost-benefit analysis below, the system operates at approximately $10–12 USD/month, while the problem it addresses — unplanned downtime — costs organizations an average of $5,600 USD/minute (Gartner, 2024). A conservative reduction in MTTR pays for the entire annual infrastructure cost many times over within a single incident.
In conclusion, Sentinel is technically feasible, operationally useful, and strategically aligned with SoftServe’s focus on software engineering, Cloud & DevOps, and AI-based solutions. Its cost-benefit balance is strongly positive for any organization that depends on reliable digital services and needs faster, more structured incident management.
Sentinel is designed to minimize infrastructure spending by combining self-hosted open-source tools with a lean cloud footprint. The table below reflects real 2025/2026 pricing.
| Component | Monthly Cost | Notes |
|---|---|---|
| OpenAI gpt-4o-mini | ~$1–2 USD | $0.15/M input tokens · $0.60/M output tokens. Estimated 1,500–3,000 LLM calls/month at 10–20 incidents/day |
| Supabase | $0 (free tier) | 500 MB DB, 200 Realtime connections — sufficient for project scale |
| LangFuse v2 | $0 | MIT-licensed, self-hosted within Docker Compose |
| ChromaDB | $0 | Apache 2.0, self-hosted within Docker Compose |
| Frontend hosting | $0 | Render Static Sites or Cloudflare Pages |
| Cloud VM (full stack) | $9–12 USD | Hetzner CX22: 2 vCPU, 4 GB RAM — minimum viable to run Prometheus + Loki + FastAPI + ChromaDB + LangFuse simultaneously |
| Total | ~$10–12 USD/month | Scalable to ~$34–36/month with Supabase Pro for production-grade reliability |
The financial case for Sentinel is grounded in industry-reported data on the cost of unplanned downtime:
- Downtime cost: $5,600 USD/minute (Gartner, 2024)
- Industry average MTTR: 28–35 minutes per incident — equivalent to $156,800–$196,000 USD per incident in lost revenue
- 43% of incident response time is spent on manual, repetitive tasks such as correlating logs, checking metrics, and identifying the responsible service (Datadog, 2024)
Sentinel directly targets that 43% by automating evidence collection, root cause analysis, and action proposal, while keeping the engineer in control of the final decision. A conservative 40% reduction in MTTR translates to saving approximately $62,000–$78,000 USD per major incident avoided or accelerated.
At an operational cost of $10–12 USD/month (~$120–144 USD/year)**, the system pays for itself by preventing or accelerating the resolution of a single significant incident per year.
| Risk | Mitigation in Sentinel |
|---|---|
| External AI dependency (OpenAI outage) | Agent fails gracefully; engineer retains full manual control via the dashboard |
| Runbook staleness | ChromaDB collections are updatable; seed scripts are versioned in the repository |
| Proposed action safety | Strict command whitelist + deterministic guardrails block any unsafe action before it reaches the engineer |
| Integration failures (Prometheus, Loki) | Metrics and logs degrade gracefully — incidents can still be created and triaged manually |