Skip to content
Draft. This page is in editorial review and is not indexed yet.

DevOps and Platform Engineering

Incident Management and On-Call: from Pagers to AI agents

Incident management moved from pagers and Nagios checks to SaaS paging with PagerDuty and Opsgenie, then to AIOps noise reduction. The agentic shift: AI SRE agents now triage alerts, investigate across telemetry and deploys, and name a likely root cause before a human opens a laptop. Fixes still wait for human approval.

Verified Updated Oct 9, 2026By Deepak Gupta
2.5

Autonomy level

2.5 of 5 · Supervised agents

launch score

Projected 3.7 by 2031

Tasks automated
2.5
Approval load
2.0
Production maturity
3.0

Overall is the mean of the three sub-scores. How scores work

Autonomy by era: On-prem 0.2, SaaS and cloud 0.8, AI-assisted 1.5, Agentic 2.5, Next 5 years 3.7.

The same job, five eras

Drag across the eras to see who did the work, with what, and what broke.

Era 1 · 1996-2009

On-prem

0.2
Who did the work
A NOC or operations team watched dashboards and escalated by phone; system administrators carried pagers and fixed things by hand. Developers were rarely on call for their own code.
Tools
Nagios (NetSaint)HP OpenViewIBM TivoliBMC PatrolNumeric and alphanumeric pagersConference bridges and phone trees
Representative product
Nagios (NetSaint)
What broke
Static thresholds paged on symptoms, so most alerts were noise. Escalation depended on printed rotas and someone answering a phone.

Autonomy 0.2/5 · Manual

Era 2 · 2009-2019

SaaS and cloud

0.8
Who did the work
Developers joined on-call rotations alongside SRE and platform teams. Incident commanders emerged as a defined role for major incidents.
Tools
PagerDutyOpsgenieVictorOpsStatuspageDatadogNew Relic
Representative product
PagerDuty
What broke
Alert fatigue from hundreds of uncorrelated monitors. Per-responder pricing made broad on-call coverage expensive.

Autonomy 0.8/5 · Tool-assisted

Era 3 · 2019-2024

AI-assisted

1.5
Who did the work
SRE and platform teams owned tooling and process; developers stayed on call. Incident commander and communications lead became standard roles on major incidents.
Tools
BigPandaMoogsoftPagerDuty AIOpsincident.ioRootlyFireHydrant
Representative product
BigPanda
What broke
Correlation models needed months of tuning and still missed novel failures. Copilots summarized incidents but could not investigate them.

Autonomy 1.5/5 · Copilot

Era 4 · 2025-2026

The Agentic Shift

2.5
Who did the work
On-call engineers now review an agent's findings instead of starting from a blank dashboard. Platform teams spend new time on agent access, runbooks, and telemetry quality.
Tools
PagerDuty SRE AgentDatadog Bits AI SREAWS DevOps Agentincident.io AI SREResolve AITraversal
Representative product
PagerDuty SRE Agent
What broke
Agents recommend fixes but rarely execute them, so humans still own mitigation. Root cause accuracy depends on telemetry coverage the team may not have.

Autonomy 2.5/5 · Supervised agents

Era 5 · 2026-2031

Next 5 years

3.7
Who did the work
Smaller on-call rotations that handle escalations, plus reliability engineers who design runbooks agents can execute and review agent actions after the fact.
Tools
Self-remediating AI SRE agentsAgent-executable runbooksProgressive delivery with automatic rollbackAgent identity and audit tooling
Representative product
Self-remediating AI SRE agents
What broke
Trust has to be earned per failure class, which is slow. An agent acting on a misdiagnosis can widen an outage.

Autonomy 3.7/5 · Agents with approvals

What job does incident management and on-call software do?

Incident management exists to shrink the time between something breaking in production and customers no longer feeling it. The job has four parts: notice the problem, get the right person awake and in the room, find and fix the cause, and learn enough that the same failure does not come back.

Every era of tooling has attacked a different part of that chain. Pagers and Nagios solved noticing. PagerDuty and Opsgenie solved waking the right person. AIOps tried to cut the noise. AI SRE agents are the first tools aimed at the expensive middle: the investigation that used to burn an hour of senior engineering time at 3 AM.

Era 1 · Before SaaS · 1996-2009

How did incident management and on-call work before SaaS?

Verified

On-call meant a pager clipped to a belt and a monitoring server in the corner of the data center. Nagios, first released as NetSaint in 1999 and renamed in 2002, ran plugin checks against hosts and services and sent email or pager notifications when a check failed. HP OpenView, IBM Tivoli, and BMC Patrol did the same job for enterprises with bigger budgets.

Larger companies staffed a NOC: a room of operators watching wall screens, running scripted triage, and calling engineers off a printed phone tree. Escalation was a laminated sheet. If the first name did not answer, the operator dialed the next one.

The tools could tell you a check had failed. They could not tell you why. Diagnosis happened in a war room or a conference bridge, with people tailing logs over SSH on individual servers. Knowledge lived in heads, so the incident lasted as long as it took to find the one person who knew the system.

What broke was signal quality. Thresholds were static, every disk at 90 percent paged someone, and teams learned to ignore alerts. The postmortem, when there was one, often ended with a name instead of a fix.

  1. NetSaint, the project that became Nagios, ships its first releasesource
  2. NetSaint is renamed Nagiossource
  3. Google starts its Site Reliability Engineering team under Ben Treynor Slosssource

Era 2 · The Cloud Move · 2009-2019

What changed when incident management and on-call moved to the cloud?

Verified

PagerDuty, founded in 2009, turned on-call into a cloud service. Monitoring tools sent events to an API, and the service handled schedules, escalation policies, and delivery by push, SMS, and phone call. Opsgenie and VictorOps followed with the same model, and per-user pricing became the norm.

The bigger change was cultural. Cloud and continuous deployment put developers on call for their own services, the you-build-it-you-run-it model. Google published its Site Reliability Engineering book in 2016, which gave the industry a shared vocabulary: SLOs, error budgets, toil, and blameless postmortems.

Communication became a product too. Statuspage, bought by Atlassian in 2016, gave companies a public place to say what was broken. Consolidation followed: Splunk bought VictorOps in 2018, and Atlassian bought Opsgenie that same year.

What broke was volume. Microservices multiplied the number of things that could alert, and every team wired its monitors straight into the pager. Waking the right person got easy. Waking them for the right reason did not. I wrote about why the same incidents keep recurring in why incidents keep happening: the cause is usually people and process, rarely just the server.

  1. PagerDuty is foundedsource
  2. Atlassian acquires Statuspagesource
  3. Google's Site Reliability Engineering book is publishedsource
  4. Splunk agrees to acquire VictorOps for about $120 millionsource
  5. Atlassian agrees to acquire Opsgenie and launches Jira Opssource

Era 3 · The Copilot Years · 2019-2024

What did AI copilots change in incident management and on-call?

Verified

The first AI in incident management was aimed at noise. AIOps platforms such as BigPanda and Moogsoft ingested alerts from every monitor, clustered related events into one incident, and suppressed duplicates. PagerDuty, by then the category leader, went public in April 2019. Dell bought Moogsoft in 2023 to fold AIOps into its operations portfolio.

Response moved into chat. incident.io, Rootly, and FireHydrant built Slack-native incident bots that open a channel, page responders, assign roles, keep a timeline, and draft the postmortem from the conversation. The bot did the paperwork; humans did the thinking.

Generative AI arrived late in this era as a copilot: summarize this incident, draft a status update, explain this log line. Useful, but the human still drove every step of the investigation.

This was also the period when the platform side matured. At LoginRadius, we moved from informal on-call rotations to an SRE practice with SLOs per service, error budgets, and post-incident reviews that produced structural fixes, which I described in our 2022 resilience review. The lesson I took: tooling helps, but the practice has to exist first.

  1. PagerDuty lists on the New York Stock Exchangesource
  2. Dell Technologies completes its acquisition of Moogsoftsource

Era 4 · The Agentic Shift · 2025-2026

The Agentic Shift: What do AI agents do in incident management and on-call today?

Verified

The AI SRE agent is the new unit of incident tooling. When an alert fires, the agent starts investigating before anyone acknowledges the page. It pulls recent deploys, queries metrics, logs, and traces, checks past incidents, tests several hypotheses in parallel, and posts a ranked root cause with its evidence into the incident channel.

The products are real and shipping. PagerDuty's SRE Agent reached general availability in October 2025 and classifies incidents, surfaces context from past incidents, and recommends remediation. Datadog launched Bits AI SRE in December 2025. AWS made its DevOps Agent generally available in March 2026. incident.io, Rootly, Resolve AI, Traversal, and Cleric all sell investigation agents, and the open-source HolmesGPT is a CNCF sandbox project.

The limit is action. Almost every vendor stops at a recommendation or a pull request. incident.io says the only change its agent can make is a pull request you review and merge. Rootly says every change requires explicit human sign-off. Cleric and HolmesGPT are read-only by default. A few, such as Resolve AI and Traversal, describe automated mitigation, but customers gate it carefully.

The second limit is context. An agent is only as good as the telemetry and access it gets, which is the same lesson as everyone buying AI observability while nobody owns agent behavior. Meanwhile Atlassian stopped selling Opsgenie in June 2025, which pushed a wave of teams to re-platform right as agents arrived.

  1. Atlassian stops selling Opsgeniesource
  2. Traversal launches its AI SRE with a $48 million seed and Series Asource
  3. incident.io opens early access to its AI SREsource
  4. PagerDuty SRE Agent becomes generally availablesource
  5. Datadog launches Bits AI SREsource
  6. Freshworks closes its acquisition of FireHydrantsource
  7. AWS DevOps Agent becomes generally availablesource

Era 5 · The Next Five Years · 2026-2031

What will incident management and on-call look like by 2031?

Verified

My bet is that incident management splits in two. Known failure classes, the ones with a runbook, a clear signal, and a reversible fix, get handled end to end by agents: a bad deploy rolled back, a saturated pool scaled, a stuck job restarted, a certificate renewed. The human sees a summary in the morning, not a page at 3 AM.

Novel failures stay human. Cascading outages, data corruption, security incidents, and anything that needs a judgment about customer impact will still pull in an incident commander. The agent becomes the fastest investigator in the room, not the decision maker.

The pattern mirrors the autonomous SOC: agents absorb the tier-1 queue, and people move up to judgment, design, and prevention. On-call rotations get smaller and quieter, but they do not disappear.

Three things gate this. Rollback and mitigation have to be safe by construction, agents need scoped identities with audit trails, and teams need evidence that agent actions are right often enough to remove the approval step for a given failure class.

My prediction · by 2031 · medium confidence

By 2031, AI SRE agents will investigate nearly every production alert and remediate most known failure classes without a human page, while novel and high-impact incidents still run under a human incident commander.

What has to be true

  • Rollback and mitigation paths are safe and reversible by design, so an agent mistake is cheap
  • Ops agents get their own scoped, auditable identities instead of shared admin credentials
  • Vendors publish measured root cause accuracy per failure class, not just MTTR anecdotes
  • Teams build executable runbooks and telemetry coverage that agents can reason over

Projected autonomy 3.7 of 5

  1. Opsgenie support ends and access shuts offsource

Then vs now: who does each step?

The job broken into its steps, and who or what does each one in each era.

Who or what does each step of Incident Management and On-Call, by era
Job stepOn-premSaaS and cloudAI-assistedAgenticNext 5 years
Detect the problemNagios check on a static thresholdCloud monitor fires into PagerDutyAIOps clusters and dedupes alertsAgent triages the alert and dismisses noiseAgent catches regressions before they page
Wake the right personNOC operator dials a phone treeEscalation policy pages the on-call engineerChat bot opens a channel and pages rolesPage arrives with the agent's findings attachedHuman paged only for novel or high-impact incidents
Investigate the causeSSH into servers and tail logsEngineer jumps between dashboardsCopilot summarizes; engineer still queriesAgent tests hypotheses and ranks root causesAgent investigates; human checks the evidence
Mitigate and fixAdmin restarts the service by handEngineer runs a runbook or rolls backEngineer runs automation scriptsAgent drafts a fix PR; human approvesAgent fixes known failure classes on its own
Communicate statusEmail to a distribution listManual Statuspage updateBot drafts updates for a human to sendAgent drafts updates from live findingsRoutine updates sent automatically, sensitive ones by a human
Learn and preventPostmortem often skipped or blame-drivenBlameless postmortem written by handBot drafts the timeline and postmortemAgent drafts postmortem and suggests runbook updatesAgent turns repeat incidents into prevention work

How does the incident management and on-call team change?

The on-call engineer's job is moving from investigator to reviewer. Instead of opening five dashboards at 3 AM, they read an agent's ranked hypotheses and evidence, approve or reject a proposed fix, and step in when the agent is wrong. That cuts the most expensive part of an incident: senior attention spent on routine diagnosis.

The platform team gains new work. Someone has to decide what each agent can read and change, keep runbooks executable, and make sure telemetry covers the systems agents are asked to reason about. When I described how LoginRadius DevOps grew into a platform team, the shift was from tickets to self-service. The next shift is from self-service for humans to safe, scoped self-service for agents.

Roles that shrink

  • Tier-1 NOC operators watching dashboards
  • Manual alert triage on every page
  • Hand-written incident timelines and postmortems
  • Large follow-the-sun rotations for routine failures

Roles that appear

  • Agent supervisor who reviews AI SRE findings and actions
  • Runbook engineer who writes procedures agents can safely execute
  • Reliability workflow designer who sets which failure classes agents may fix alone
  • Agent access owner who manages scopes, credentials, and audit for ops agents

Skills to learn

  • Judging an agent's evidence quickly and spotting a confident wrong answer
  • Writing machine-executable runbooks with clear rollback steps
  • Designing SLOs and telemetry that agents can reason over
  • Scoping least-privilege access for non-human identities
  • Incident command and stakeholder communication for novel failures

What gets easier for the humans?

BeforeAfter
Paged at 3 AM to start an investigation from a blank dashboardPaged with a ranked root cause and evidence already in the incident channel
Hundreds of uncorrelated alerts per incidentOne triaged incident with noise dismissed and reasons stated
Hunting through deploy history to find what changedAgent links the alert to the likely deploy or config change
Writing the postmortem timeline by hand from chat scrollbackReviewing a drafted timeline and postmortem built from the incident record
Status updates written from memory during the firefightDrafted updates based on live findings, sent after a quick human check

Decisions that stay human

  • Declaring a major incident and setting its severity when customer impact is unclear
  • Approving changes to production data, security controls, or anything irreversible
  • Deciding what to tell customers, regulators, and executives
  • Leading response to security incidents where the attacker may be steering the evidence
  • Choosing which failure classes an agent is trusted to fix alone

Where should agents not act alone?

Risks and failure modes, through a security and identity lens.

  1. 01

    Over-privileged ops agents

    An AI SRE agent with broad production write access is a high-value identity. Give each agent its own credentials, least-privilege scopes per action, and short-lived tokens, the same discipline I argue for in [AI agents don't have passwords](/ai-agents-dont-have-passwords/). Read-only by default should be the starting point.

  2. 02

    Prompt injection through telemetry

    Agents read logs, tickets, and chat. An attacker who can write a crafted log line or ticket comment can try to steer the agent's conclusion or its proposed fix. Treat everything the agent reads as untrusted input and never let text in telemetry authorize an action.

  3. 03

    Confident wrong root cause

    An agent that names the wrong cause with high confidence can send responders down the wrong path or trigger the wrong rollback. Require cited evidence for every finding and keep a human approval gate on mitigation until accuracy is measured per failure class.

  4. 04

    Missing audit trail

    If an agent restarts a service or rolls back a deploy, the incident record must show which agent, under which identity, with what evidence, and who approved. Without that, postmortems and compliance reviews cannot reconstruct what happened.

  5. 05

    Security incidents handled as ops incidents

    An outage caused by an intrusion looks like an ops failure at first. An agent that restores service by redeploying can destroy forensic evidence. Route anything with a security signal to humans and the SOC before an agent acts.

  6. 06

    Skill atrophy on call

    If agents handle every routine incident, engineers lose the hands-on practice they need for the novel ones. Keep game days and chaos exercises so people still know the systems.

Who is building agentic incident management and on-call?

Incumbents adding agents vs agent-native entrants. Capability lines are checked against each vendor's own site.

Incumbents

  • SRE Agent classifies incidents, surfaces context from past incidents, and recommends remediation steps, in Slack and the Operations Console.

    Checked Oct 9, 2026Compare

  • AI SRE investigates from the moment an incident is declared and names the likely cause with sources; its only change to systems is a pull request a human merges.

    Checked Oct 9, 2026Compare

  • AI SRE ranks likely root causes with confidence scores and suggests fixes; every change requires explicit human sign-off.

    Checked Oct 9, 2026Compare

  • Bits AI SRE investigates alerts using telemetry and runbooks and reports a likely root cause to collaboration tools.

    Checked Oct 9, 2026Compare

  • Handles initial incident triage across AWS, other clouds, and on-premises systems and recommends changes to prevent repeat outages.

    Checked Oct 9, 2026

  • Agentic ITOps platform that uses AI agents for L1 detection, triage, and major incident coordination.

    Checked Oct 9, 2026

Agent-native

  • AI SRE that goes on call, screens alerts, applies mitigations, and escalates to engineers when needed.

    Checked Oct 9, 2026

  • AI SRE that tests many hypotheses in parallel to isolate root cause and offers automated remediation; read-only by default.

    Checked Oct 9, 2026

  • Investigates alerts and production changes, dismisses false positives, and opens fix pull requests for human approval.

    Checked Oct 9, 2026

Open source

  • Apache 2.0 CNCF sandbox SRE agent that investigates incidents with read-only access by default; optional toolsets can apply fixes.

    Checked Oct 9, 2026

Side-by-side comparisons: Top 6 Incident Management and On-Call Platforms for 2026: incident.io vs PagerDuty vs Rootly vs Grafana IRM vs Better Stack vs FireHydrant, Top 5 Observability Platforms of 2026: Datadog vs Grafana vs the Rest.

Questions people ask

How is AI changing incident management?

AI SRE agents now start investigating the moment an alert fires. They pull recent deploys, query telemetry, compare past incidents, and post a ranked root cause with evidence. PagerDuty, Datadog, AWS, incident.io, and Rootly all ship this. Most still stop at a recommendation or a pull request for a human to approve.

What is an AI SRE?

An AI SRE is an agent that does the investigation work of a site reliability engineer during an incident: triaging alerts, correlating signals across logs, metrics, traces, and code changes, and proposing a root cause and fix. Vendors include Resolve AI, Traversal, Cleric, incident.io, and Rootly, plus the open-source HolmesGPT.

Will AI agents replace on-call engineers?

Not in the next five years. Agents will absorb routine triage and fixes for known failure classes, so rotations get smaller and quieter. Novel outages, security incidents, and decisions about customer impact still need a human incident commander. The on-call role shifts from investigator to reviewer.

Can AI SRE agents fix incidents automatically?

Some can, but most are configured not to. incident.io limits its agent to pull requests a human merges, and Rootly requires human sign-off on every change. Resolve AI and Traversal describe automated mitigation. In practice, teams allow auto-remediation only for reversible actions on well-understood failure classes.

What happened to Opsgenie?

Atlassian stopped selling Opsgenie on 4 June 2025 and ends support on 5 April 2027, when access shuts off and unmigrated data is deleted. Atlassian points customers to Jira Service Management or Compass; many teams are using the deadline to move to PagerDuty, incident.io, Rootly, or similar tools.

Is it safe to give an AI agent access to production?

Only with guardrails. Give the agent its own identity, least-privilege scopes, and short-lived credentials. Start read-only, log every action with the evidence behind it, and require human approval for anything irreversible. Treat logs and tickets the agent reads as untrusted input because they can carry prompt injection.

What is the difference between AIOps and an AI SRE agent?

AIOps, from tools like BigPanda and Moogsoft, groups and deduplicates alerts to cut noise. An AI SRE agent goes further: it investigates the incident, reasons across telemetry and code changes, and proposes a cause and fix. AIOps tells you what is related; an agent tells you why it broke.

Sources

  1. Nagios, accessed Oct 9, 2026
  2. History of Nagios, accessed Oct 9, 2026
  3. Site reliability engineering, accessed Oct 9, 2026
  4. PagerDuty, accessed Oct 9, 2026
  5. PagerDuty Announces Closing of Initial Public Offering, accessed Oct 9, 2026
  6. Atlassian Acquires Status and Incident Communication Platform StatusPage, accessed Oct 9, 2026
  7. Splunk Agrees to Acquire VictorOps, accessed Oct 9, 2026
  8. Jira Ops: incident management platform, accessed Oct 9, 2026
  9. Dell Technologies Acquires Moogsoft, accessed Oct 9, 2026
  10. Opsgenie pricing and end of sale FAQ, accessed Oct 9, 2026
  11. SRE Agent is Generally Available, accessed Oct 9, 2026
  12. PagerDuty Launches Industry's First End to End AI Agent Suite, accessed Oct 9, 2026
  13. AI SRE has entered the chat, accessed Oct 9, 2026
  14. incident.io AI SRE, accessed Oct 9, 2026
  15. Rootly AI SRE, accessed Oct 9, 2026
  16. Datadog Launches Bits AI SRE Agent to Resolve Incidents Faster, accessed Oct 9, 2026
  17. AWS DevOps Agent is now generally available, accessed Oct 9, 2026
  18. Announcing our Seed and Series A from Sequoia and Kleiner Perkins to Launch the AI SRE for the Enterprise, accessed Oct 9, 2026
  19. Traversal, accessed Oct 9, 2026
  20. Resolve AI, accessed Oct 9, 2026
  21. Ex-Splunk execs' startup Resolve AI hits $1 billion valuation with Series A, accessed Oct 9, 2026
  22. Cleric, accessed Oct 9, 2026
  23. HolmesGPT, accessed Oct 9, 2026
  24. BigPanda, accessed Oct 9, 2026
  25. Delivering the future of ServiceOps today (Freshworks closes FireHydrant acquisition), accessed Oct 9, 2026

Published Oct 9, 2026. Last verified Oct 9, 2026. Eras 4 and 5, vendors, and scores are re-checked every six to eight weeks; see the changelog and methodology.