Skip to content

v2.3.0

Latest

Choose a tag to compare

@Aidaho12 Aidaho12 released this 01 Oct 05:06
051508c

Summary

IncidentRelay 2.3 introduces the first production-ready stage of Incident Management v2.

The main architectural change is a clear separation between technical alert processing and operational incident response:

Alert → AlertGroup → Incident

AlertGroup remains responsible for technical signal handling, grouping, notification and escalation, while the new first-class Incident represents the operational investigation and ownership workflow.

Highlights

First-class Incidents

IncidentRelay now provides a dedicated operational Incident entity independent from AlertGroup lifecycle.

Incidents support:

  • independent operational lifecycle;
  • P1–P5 priority;
  • owning team and service;
  • operational assignee;
  • manual assign, reassign and unassign;
  • Assign to me;
  • explicit close and reopen;
  • linking and unlinking AlertGroups;
  • Incident timeline and audit events;
  • optimistic concurrency protection.

An Incident can exist without an AlertGroup, and an AlertGroup can exist without an Incident.

One Incident may also contain multiple independently managed AlertGroups.

AlertGroup and Incident API separation

The public API now reflects the final domain model:

/api/alert-groups → technical AlertGroups
/api/incidents    → operational Incidents

The previous AlertGroup-shaped /api/incidents contract has been removed.

There is no compatibility alias or runtime switch preserving the old API.

Incident IDs and AlertGroup IDs use independent namespaces and may contain the same numeric value.

Explicit creation workflows

The two manual creation workflows are now independent:

Create Alert Group
  → AlertGroup
  → one initial child Alert

and:

Create Incident
  → standalone Incident
  → optional AlertGroup links

Creating an Incident never creates a hidden AlertGroup or Alert.

Creating an AlertGroup never creates a hidden Incident.

Independent ownership

Technical and operational responsibility are now separated.

AlertGroup.assignee represents technical responsibility used by notification and escalation.

Incident.assignee represents ownership of the operational investigation.

Reassigning either resource does not silently modify the other.

AlertGroup reassignment also does not acknowledge the alert or reset its escalation state.

Separate Alerts and Incidents UI

Alerts and Incidents now have separate UI surfaces matching their different responsibilities.

The Incidents page includes:

  • dedicated Incident list and details;
  • status and priority visibility;
  • direct Incident navigation;
  • pagination and per-page controls;
  • operational assignment actions;
  • linked AlertGroup context.

Notification Policy common filters

Notification Policies now expose common routing conditions directly in the rule editor:

  • Priority;
  • Severity;
  • Alert source;
  • Service;
  • Service environment;
  • Service criticality;
  • Service tier.

These controls use the existing matcher engine internally rather than introducing another filtering mechanism.

The existing matcher editor remains available as Advanced matchers for labels, annotations, route fields and custom combinations.

Improved Alert bulk actions

Alert bulk actions now include:

  • Acknowledge;
  • Resolve;
  • Shelve;
  • Unshelve;
  • Merge.

Bulk Shelve reuses the normal shelving lifecycle, including duration and reason.

Bulk action identifiers are now independent from translated UI labels, fixing Resolve failures in non-English locales.

Production and security hardening

2.3 includes additional safeguards around the new Incident model:

  • team-scope enforcement for Incident/AlertGroup links;
  • cross-team RBAC regression coverage;
  • AlertGroup/Incident ID collision coverage;
  • PostgreSQL concurrency coverage for competing Incident links;
  • migration reconciliation reporting;
  • upgrade and rollback regression coverage;
  • audit and timeline coverage for core mutations.

Breaking API change

Before 2.3, technical grouped alerts were exposed through:

/api/incidents

Starting with 2.3 they are available through:

/api/alert-groups

The /api/incidents endpoint now represents only first-class operational Incidents.

External integrations and API clients must be updated before upgrading.

For example:

technical AlertGroup ID → /api/alert-groups/{id}
operational Incident ID → /api/incidents/{id}

Do not assume that an Incident ID and AlertGroup ID identify the same resource.

Database migration

The Incident core migration adds:

incident
incident_event
incident_alert_group_link

Existing AlertGroups and child Alerts are preserved.

IncidentRelay does not automatically create historical Incidents for every existing AlertGroup.

Before upgrading a production installation:

  1. back up the database and configuration;
  2. stop old web/scheduler/worker processes;
  3. update external API clients;
  4. run the 2.3 migrations from a single process;
  5. verify migration status and perform post-upgrade reconciliation.
python manage.py migration-status
python manage.py migrate
python manage.py migration-status

See:

docs/incidents/api-migration-2.3.md

for the full upgrade and rollback procedure.

Warning

Downgrading the Incident core migration removes the new Incident tables.

Once first-class Incident data has been created, rollback should not be treated as data-preserving without an appropriate backup/export.

Upgrade verification

After upgrading, verify that:

  • /api/alert-groups returns technical AlertGroups;
  • /api/incidents returns first-class Incidents;
  • existing AlertGroup history and comments are preserved;
  • creating an Incident does not create an AlertGroup;
  • creating a manual AlertGroup creates exactly one initial child Alert;
  • Incident ↔ AlertGroup link/unlink works correctly;
  • team-scope permissions are enforced.

Run the regular test suite and, for PostgreSQL deployments, the dedicated PostgreSQL tests:

pytest -q
pytest -q tests/postgresql