Skip to content

Automated Testing

Nicolas Rico edited this page May 10, 2026 · 3 revisions

Automated Software Testing

This page documents the automated testing strategy implemented during Sprint 2, the types of tests that were chosen, why they were chosen, and how they are executed.

Summary

In total, we have 59 automated tests passing: 48 in the backend with pytest and 11 in the frontend with vitest. All functionalities delivered in the sprint are covered with at least one test for the main flow (happy path) and another for an alternative flow, as required by the master document.

The current coverage is 85% in services/alert_processor.py, which is the critical detection module, and 100% in models/incident.py, where the system's data contracts reside.

Strategy and justification by functionality

To define which type of test to apply to each functionality, the team evaluated two things: how tightly coupled the logic is to external services, and how critical the component is for the incident flow.

Functionality Test Type Why this type
US-03: Alertmanager alert parser (process_prometheus_alert) Unit It is pure transformation of a JSON payload. Unit tests are fast, deterministic, and do not require Supabase or Loki running.
Severity mapping by alert name (SEVERITY_MAP) Parameterized It is a lookup table with many combinations. A single parameterized test covers 13 known alerts and one additional default case.
US-06: Loki logs query (query_loki_logs) Integration with HTTP mock It depends on an external service (Loki). We mock requests.get to validate the contract and parsing without needing to run Loki.
US-04: Alert routing to the correct agent Scenario-based Each agent has rules to decide if an alert belongs to it. We test multiple scenarios (Docker, Postgres, Kubernetes) to ensure the supervisor routes to the correct domain agent.
US-01: Incident listing with filters Integration with FastAPI TestClient It is a real HTTP endpoint with Supabase mocked. It verifies filter composition and date ordering.
US-02: Manual incident creation Pydantic contract Validates that the input payload respects the severity enum and length limits.
Frontend: getRisk() function Unit Pure function that determines risk level. It can be tested without rendering React.
Frontend: <ApprovalModal /> component Component with Testing Library It is critical UI for the human approval flow. We test rendering, real user interaction, and keyboard events.

The criterion we always applied was the same: use the simplest test type that validates what we need, and increase complexity only when necessary.

Test code distribution

In the backend, tests are located in Backend/tests/ and organized by module:

test_alert_processor.py covers the 7 tests of US-03 (automatic detection). Here we validate that a firing alert creates an incident with the correct severity, that a resolved alert closes active incidents, that duplicate alerts are discarded, and that Postgres alerts build the target with the postgres/<datname> format.

test_severity_map.py contains 13 parameterized tests covering all supported alerts (Docker, Podman, and Postgres), plus two additional tests: one to verify that an unknown alertname falls back to the default medium, and another to confirm that an explicit severity label in the alert takes precedence over the map.

test_loki_logs.py covers US-06 with 4 tests: the happy path where Loki returns log lines, the empty container_id case, the ConnectionError case (Loki down), and the empty response case. The key point is that no Loki failure should break the incident creation flow.

test_agent_routing.py covers US-04 and domain separation. We have 7 tests verifying that the DockerAgent matches when appropriate, rejects Kubernetes, rejects Postgres targets, and that the agent registry validates non-empty names.

test_models.py contains 17 tests for Pydantic contracts, ensuring that all valid severities are accepted, all invalid ones are rejected with ValidationError, and the same for incident states.

In the frontend, tests are located in Frontend/src/test/. They are fewer but equally relevant:

getRisk.test.js contains 6 tests for the pure function that determines the risk level of an action based on the incident type and the command to be executed. The main rule: any command that is read-only (contains "logs") is considered low risk.

ApprovalModal.test.jsx contains 5 tests for the approval component: it displays the proposed command and incident details, the approve button triggers onApprove with the current comment, the Escape key closes the modal, the × button also closes it, and a read-only command shows low risk in the UI.

Technical configuration

In the backend, we added pytest, pytest-mock, and pytest-cov in a new file Backend/requirements-dev.txt to avoid contaminating production dependencies. Pytest and coverage configuration is defined in Backend/pyproject.toml under the sections [tool.pytest.ini_options] and [tool.coverage.run].

The most important setup file is Backend/tests/conftest.py. There, we mock the Supabase client and environment variables before any backend module is imported, preventing tests from accidentally hitting the real database. We also define reusable fixtures such as mock_supabase, sample_firing_alert, and sample_resolved_alert to avoid duplicating setup across files.

In the frontend, we added vitest, Testing Library, jsdom, and user-event as devDependencies in Frontend/package.json. Configuration is defined in Frontend/vite.config.js under the test section, with global setup, jsdom environment, and v8 coverage reporter. The file Frontend/src/test/setup.js mocks Supabase for all component tests.

How to run tests locally

For the backend, from the Backend folder: pip install -r requirements.txt -r requirements-dev.txt pytest --cov=services --cov=routers --cov=models

For the frontend, from the Frontend folder: npm install npm run test

If detailed frontend coverage is needed, npm run test:coverage generates the HTML report in Frontend/coverage/.

Continuous integration

Each Pull Request to the main and develop branches automatically triggers the workflow defined in .github/workflows/tests.yml. This workflow runs two jobs in parallel:

The backend job installs dependencies, runs ruff (static analysis), executes the pytest suite with coverage, and uploads the XML report as an artifact.

The frontend job installs dependencies with npm ci, runs eslint, executes vitest with coverage, and uploads the HTML report as an artifact.

If either job fails, GitHub blocks the merge button until it is fixed. This directly satisfies the master document requirement to run automated tests on every pull request and prevent merging if they fail.

Static analysis from Sprint 1

The issues detected by ruff during Sprint 1 have already been fixed. The few remaining ignores are explicitly documented in Backend/pyproject.toml with comments explaining why they are intentional (for example, routers/alerts.py uses camelCase names because that is how the JSON from Alertmanager arrives and cannot be changed).

Alignment with user stories

Each acceptance criterion from the Sprint 2 backlog is covered by at least one automated test:

  • US-03: covered by the 7 tests in test_alert_processor.py and the 13 in test_severity_map.py.
  • US-06: covered by the 4 tests in test_loki_logs.py.
  • US-04: covered by the 7 tests in test_agent_routing.py.
  • US-01: covered by the incident router tests.
  • US-02: covered by the 17 model tests.
  • Approval HU: covered by the 11 frontend tests.

Clone this wiki locally