Skip to content

v2.0.0 — Scalable Runtime, Durable Incident Delivery & Enterprise Operations

Latest

Choose a tag to compare

@Dushyant-rahangdale Dushyant-rahangdale released this 03 Oct 05:02
Immutable release. Only release title and notes can be modified.
097f491

OpsKnight v2.0.0 is the largest platform release to date, delivering a scalable production runtime, an enterprise-grade notification control plane, deep Slack and Microsoft Teams ChatOps, Twilio voice paging, SCIM 2.0 provisioning, and hardened multi-topology deployments.


🚀 Added & Highlights

1. 🏗️ Scalable Production Runtime

  • Independent Tier Scaling: Separate runtime profiles allow Web ingress, Scheduler, General Worker, Critical Worker, Bulk Worker, and Status Projector to scale independently according to workload demand.
  • PgBouncer Connection Pooling: Built-in support for transaction-mode PgBouncer connection pooling to support high-concurrency database workloads with graceful connection lifecycle management.
  • Serialized Schema Migration Ownership: Dedicated migration owner architecture ensures safe, idempotent database migrations without race conditions across multi-replica deployments.
  • Multi-Topology Orchestration: Production-ready deployment configurations and health probes for Docker Compose, Docker Swarm, Kubernetes Helm, and Kustomize.

2. 📟 Durable Notification & Paging Control Plane

  • Multi-Lane Traffic Isolation: Isolated Critical, Transactional, and Bulk delivery lanes with admission control to protect emergency pages from bulk notification saturation.
  • Resilient Retry & Deferred Queueing: Exponential backoff with jitter, dead-letter protection, stale-worker fencing, and automatic recovery.
  • Comprehensive Delivery Evidence: Full observability into notification lifecycle states distinguishing queued, deferred, in-flight, delivered, permanently failed, and superseded notifications.
  • Provider Fallback & Circuit Breakers: Automatic failover between configured notification channels with circuit breaker protection when external providers experience degradation.

3. 💬 Slack & Microsoft Teams ChatOps

  • Microsoft Teams Integration: Native Entra ID application and Azure Bot integration featuring rich Adaptive Cards, incident status synchronization, and meeting collaboration.
  • Multi-Destination Routing: A service can target up to three distinct Slack and three Microsoft Teams destinations simultaneously.
  • Interactive ChatOps Actions: Responder actions directly from chat channels including Acknowledge, Resolve, Add Note, Escalate, and Assign Responder.
  • Incident War Rooms: Dedicated incident collaboration channels with automatic responder and participant synchronization.

4. 📞 Voice Paging & Mobile PWA

  • Twilio Voice Paging: Critical operational voice paging for newly triggered incidents with interactive touch-tone (DTMF) acknowledgement by responders.
  • Signed Webhook Reconciliation: Cryptographically signed callbacks and provider status reconciliation for reliable call tracking.
  • Responder-Grade Mobile PWA: Fully responsive Progressive Web App with fast mobile navigation, offline cache boundaries, and native-like on-call workflows.
  • Web Push Notifications: Real-time push notifications on Android and iOS with per-device subscription tracking and automatic recovery.

5. 🚨 Incident Response & On-Call

  • Hardened Incident Lifecycle: Centralized lifecycle state engine with persistent idempotency and guaranteed side-effect execution.
  • Interactive Postmortems & Action Items: Post-incident review tooling including 5-Whys root cause analysis, action items tracking, and immutable timeline auditing.
  • Flexible On-Call Schedules: Multi-layer timezone-aware rotations, DST-safe handoffs, and temporary responder overrides.
  • Tiered Escalation Policies: Support-hours conditions, priority-based escalation, and automatic fallback rules.
  • Quiet Hours: Personalized timezone-aware alert suppression for non-urgent notifications while guaranteeing critical pages cut through.

6. 🔐 Identity, Security & Compliance

  • SCIM 2.0 Provisioning: Automated enterprise user and group provisioning/deprovisioning with bearer token generation, rotation, and revocation.
  • Auditor Role: Dedicated read-only AUDITOR role providing full compliance and investigation access without operational modification permissions.
  • Enterprise OIDC Single Sign-On: PKCE-hardened SSO support for Microsoft Entra, Okta, Google Workspace, Auth0, and generic OIDC providers with role claim mapping.
  • Canonical Session Registry: Real-time visibility into all active user sessions across devices, browsers, and IP addresses with instant remote revocation.
  • Zero-Downtime Keyring Rotation: Multi-key encryption keyring support (k2:new,k1:old) allowing seamless credential and data re-encryption without service interruption.
  • Dedicated API Key Secret: Separate API_KEY_SECRET isolated from session and database encryption keys.
  • Privacy & Compliance Operations: Built-in DSAR export/erasure workflows, verifiable evidence packages, and retention holds.

7. 📊 Status, Analytics & Operations

  • Status Page V3: Redesigned status page with custom branding, maintenance announcements, subscriber email verification, and uptime history.
  • Operational Analytics & Dashboards: Customizable incident metrics widgets, filtered share links, PDF exports, and fullscreen NOC/TV presentation mode.
  • Administrator Health Center: Real-time diagnostic center monitoring database connectivity, worker backlogs, provider health, and security configuration.
  • Prometheus Metrics & Structured Logs: Native /metrics endpoint and JSON structured audit logging for integration with enterprise observability platforms.

8. 🔌 Integrations

  • ManageEngine ServiceDesk Plus: Native inbound alert ingestion with automated payload parsing and incident deduplication.
  • 28 Certified Inbound Contracts: Rigorously verified inbound integration contracts with HMAC signature verification and replay protection.
  • Jira Cloud Synchronization: Bidirectional lifecycle synchronization with automated issue creation, status transition, and resolution syncing.

9. 📦 Production Deployment

  • Multi-Architecture Container Images: Certified Linux AMD64 and Linux ARM64 container images published to GitHub Container Registry (GHCR).
  • Topology Flexibility: Full support for single-container development, integrated production, and distributed multi-replica architectures.
  • Comprehensive Load Certification: High-throughput validation across Compose, Docker Swarm, and Kubernetes environments.

⚠️ Important 2.0 Boundaries

  • Single Status Page: OpsKnight supports one official status page per instance; multiple independent status pages per deployment are not supported.
  • Service Objectives (SLO) UI Deferred: The SLO user interface is deferred and its route redirects; underlying database schemas are retained for future releases.
  • Installable PWA: Mobile capabilities are delivered via an installable Progressive Web App (PWA) with Web Push, not a native App Store / Google Play binary.
  • Voice Paging Scope: Voice paging is strictly for initial triggered-incident alerting and responder ACK; subsequent lifecycle updates (re-route, resolve) do not trigger phone calls.
  • Self-Hosted Only: OpsKnight is strictly self-hosted; there is no multi-tenant "OpsKnight Cloud" hosted SaaS.
  • Integration Scope: 28 certified integration contracts represents the total number of supported inbound integrations, not 28 newly introduced ones.
  • License: OpsKnight v2.0.0 and subsequent releases are licensed under the GNU Affero General Public License v3.0 (AGPL-3.0-only).

Upgrading from v1.4.0

Warning

Always create a full backup of your PostgreSQL database before upgrading.
Ensure existing ENCRYPTION_KEY is preserved, and generate a new dedicated API_KEY_SECRET before starting v2.0.0.
If running PgBouncer or connection poolers, ensure DIRECT_DATABASE_URL is configured to point directly to PostgreSQL port 5432 so database migrations run successfully.

Docker / Compose

# 1. Pull the certified v2.0.0 release image
docker pull ghcr.io/opsknight-labs/opsknight:2.0.0

# 2. Update .env with new required variables
# Generate API_KEY_SECRET: openssl rand -base64 32
# If migrating encryption keys: ENCRYPTION_KEYS="k1:<current_ENCRYPTION_KEY>"

# 3. Restart the stack
docker compose down && docker compose up -d

Kubernetes / Helm

# 1. Update repository
helm repo update

# 2. Upgrade release with v2.0.0 image tag
helm upgrade opsknight opsknight/opsknight \
  --set image.tag=2.0.0 \
  --reuse-values

Docker Swarm

# 1. Pull release image across nodes
docker pull ghcr.io/opsknight-labs/opsknight:2.0.0

# 2. Deploy updated stack
OPSKNIGHT_IMAGE=ghcr.io/opsknight-labs/opsknight:2.0.0 ./deploy/swarm/scripts/deploy.sh

Post-Upgrade Verification Checklist

  1. Verify Health Center: Open Administration → Health Center and confirm all database, worker, and scheduler indicators report green.
  2. Synthetic Incident Test: Create a test incident manually or via API to verify the new incident engine.
  3. Lifecycle Confirmation: Acknowledge and resolve the incident to verify state machine transitions and audit log entries.
  4. Notification Delivery: Verify alerts are dispatched successfully to configured channels (Email, Slack, Teams, Voice, Web Push).
  5. Integration Review: Confirm external monitoring webhooks and alert ingestions continue to process normally.
  6. Mobile PWA Verification: If using Web Push, open the PWA on responder devices to confirm subscription registration and notification delivery.

Full Changelog · Documentation · Upgrade Guide

Release certification

  • Release image index digest: sha256:099023cbd024fac4c53435aaf2bff669f0e8bf04ba96564d3a3d62d6bdfe0d62
  • PgBouncer image digest: sha256:20bb362e68d351c0eddce594c23fadacbb423b68f5875ef33c0ff85f04623afd