DevOps and Platform Engineering
Incident Management and On-Call: from Pagers to AI agents
Incident management moved from pagers and Nagios checks to SaaS paging with PagerDuty and Opsgenie, then to AIOps noise reduction. The agentic shift: AI SRE agents now triage alerts, investigate across telemetry and deploys, and name a likely root cause before a human opens a laptop. Fixes still wait for human approval.
Autonomy level
2.5 of 5 · Supervised agents
launch score
Projected 3.7 by 2031
- Tasks automated
- 2.5
- Approval load
- 2.0
- Production maturity
- 3.0
Overall is the mean of the three sub-scores. How scores work
The same job, five eras
Drag across the eras to see who did the work, with what, and what broke.
What job does incident management and on-call software do?
Incident management exists to shrink the time between something breaking in production and customers no longer feeling it. The job has four parts: notice the problem, get the right person awake and in the room, find and fix the cause, and learn enough that the same failure does not come back.
Every era of tooling has attacked a different part of that chain. Pagers and Nagios solved noticing. PagerDuty and Opsgenie solved waking the right person. AIOps tried to cut the noise. AI SRE agents are the first tools aimed at the expensive middle: the investigation that used to burn an hour of senior engineering time at 3 AM.
Era 1 · Before SaaS · 1996-2009
How did incident management and on-call work before SaaS?
On-call meant a pager clipped to a belt and a monitoring server in the corner of the data center. Nagios, first released as NetSaint in 1999 and renamed in 2002, ran plugin checks against hosts and services and sent email or pager notifications when a check failed. HP OpenView, IBM Tivoli, and BMC Patrol did the same job for enterprises with bigger budgets.
Larger companies staffed a NOC: a room of operators watching wall screens, running scripted triage, and calling engineers off a printed phone tree. Escalation was a laminated sheet. If the first name did not answer, the operator dialed the next one.
The tools could tell you a check had failed. They could not tell you why. Diagnosis happened in a war room or a conference bridge, with people tailing logs over SSH on individual servers. Knowledge lived in heads, so the incident lasted as long as it took to find the one person who knew the system.
What broke was signal quality. Thresholds were static, every disk at 90 percent paged someone, and teams learned to ignore alerts. The postmortem, when there was one, often ended with a name instead of a fix.
Era 2 · The Cloud Move · 2009-2019
What changed when incident management and on-call moved to the cloud?
PagerDuty, founded in 2009, turned on-call into a cloud service. Monitoring tools sent events to an API, and the service handled schedules, escalation policies, and delivery by push, SMS, and phone call. Opsgenie and VictorOps followed with the same model, and per-user pricing became the norm.
The bigger change was cultural. Cloud and continuous deployment put developers on call for their own services, the you-build-it-you-run-it model. Google published its Site Reliability Engineering book in 2016, which gave the industry a shared vocabulary: SLOs, error budgets, toil, and blameless postmortems.
Communication became a product too. Statuspage, bought by Atlassian in 2016, gave companies a public place to say what was broken. Consolidation followed: Splunk bought VictorOps in 2018, and Atlassian bought Opsgenie that same year.
What broke was volume. Microservices multiplied the number of things that could alert, and every team wired its monitors straight into the pager. Waking the right person got easy. Waking them for the right reason did not. I wrote about why the same incidents keep recurring in why incidents keep happening: the cause is usually people and process, rarely just the server.
Era 3 · The Copilot Years · 2019-2024
What did AI copilots change in incident management and on-call?
The first AI in incident management was aimed at noise. AIOps platforms such as BigPanda and Moogsoft ingested alerts from every monitor, clustered related events into one incident, and suppressed duplicates. PagerDuty, by then the category leader, went public in April 2019. Dell bought Moogsoft in 2023 to fold AIOps into its operations portfolio.
Response moved into chat. incident.io, Rootly, and FireHydrant built Slack-native incident bots that open a channel, page responders, assign roles, keep a timeline, and draft the postmortem from the conversation. The bot did the paperwork; humans did the thinking.
Generative AI arrived late in this era as a copilot: summarize this incident, draft a status update, explain this log line. Useful, but the human still drove every step of the investigation.
This was also the period when the platform side matured. At LoginRadius, we moved from informal on-call rotations to an SRE practice with SLOs per service, error budgets, and post-incident reviews that produced structural fixes, which I described in our 2022 resilience review. The lesson I took: tooling helps, but the practice has to exist first.
Era 4 · The Agentic Shift · 2025-2026
The Agentic Shift: What do AI agents do in incident management and on-call today?
The AI SRE agent is the new unit of incident tooling. When an alert fires, the agent starts investigating before anyone acknowledges the page. It pulls recent deploys, queries metrics, logs, and traces, checks past incidents, tests several hypotheses in parallel, and posts a ranked root cause with its evidence into the incident channel.
The products are real and shipping. PagerDuty's SRE Agent reached general availability in October 2025 and classifies incidents, surfaces context from past incidents, and recommends remediation. Datadog launched Bits AI SRE in December 2025. AWS made its DevOps Agent generally available in March 2026. incident.io, Rootly, Resolve AI, Traversal, and Cleric all sell investigation agents, and the open-source HolmesGPT is a CNCF sandbox project.
The limit is action. Almost every vendor stops at a recommendation or a pull request. incident.io says the only change its agent can make is a pull request you review and merge. Rootly says every change requires explicit human sign-off. Cleric and HolmesGPT are read-only by default. A few, such as Resolve AI and Traversal, describe automated mitigation, but customers gate it carefully.
The second limit is context. An agent is only as good as the telemetry and access it gets, which is the same lesson as everyone buying AI observability while nobody owns agent behavior. Meanwhile Atlassian stopped selling Opsgenie in June 2025, which pushed a wave of teams to re-platform right as agents arrived.
- Atlassian stops selling Opsgeniesource
- Traversal launches its AI SRE with a $48 million seed and Series Asource
- incident.io opens early access to its AI SREsource
- PagerDuty SRE Agent becomes generally availablesource
- Datadog launches Bits AI SREsource
- Freshworks closes its acquisition of FireHydrantsource
- AWS DevOps Agent becomes generally availablesource
Era 5 · The Next Five Years · 2026-2031
What will incident management and on-call look like by 2031?
My bet is that incident management splits in two. Known failure classes, the ones with a runbook, a clear signal, and a reversible fix, get handled end to end by agents: a bad deploy rolled back, a saturated pool scaled, a stuck job restarted, a certificate renewed. The human sees a summary in the morning, not a page at 3 AM.
Novel failures stay human. Cascading outages, data corruption, security incidents, and anything that needs a judgment about customer impact will still pull in an incident commander. The agent becomes the fastest investigator in the room, not the decision maker.
The pattern mirrors the autonomous SOC: agents absorb the tier-1 queue, and people move up to judgment, design, and prevention. On-call rotations get smaller and quieter, but they do not disappear.
Three things gate this. Rollback and mitigation have to be safe by construction, agents need scoped identities with audit trails, and teams need evidence that agent actions are right often enough to remove the approval step for a given failure class.
My prediction · by 2031 · medium confidence
By 2031, AI SRE agents will investigate nearly every production alert and remediate most known failure classes without a human page, while novel and high-impact incidents still run under a human incident commander.
What has to be true
- Rollback and mitigation paths are safe and reversible by design, so an agent mistake is cheap
- Ops agents get their own scoped, auditable identities instead of shared admin credentials
- Vendors publish measured root cause accuracy per failure class, not just MTTR anecdotes
- Teams build executable runbooks and telemetry coverage that agents can reason over
Projected autonomy 3.7 of 5
- Opsgenie support ends and access shuts offsource
Then vs now: who does each step?
The job broken into its steps, and who or what does each one in each era.
| Job step | On-prem | SaaS and cloud | AI-assisted | Agentic | Next 5 years |
|---|---|---|---|---|---|
| Detect the problem | Nagios check on a static threshold | Cloud monitor fires into PagerDuty | AIOps clusters and dedupes alerts | Agent triages the alert and dismisses noise | Agent catches regressions before they page |
| Wake the right person | NOC operator dials a phone tree | Escalation policy pages the on-call engineer | Chat bot opens a channel and pages roles | Page arrives with the agent's findings attached | Human paged only for novel or high-impact incidents |
| Investigate the cause | SSH into servers and tail logs | Engineer jumps between dashboards | Copilot summarizes; engineer still queries | Agent tests hypotheses and ranks root causes | Agent investigates; human checks the evidence |
| Mitigate and fix | Admin restarts the service by hand | Engineer runs a runbook or rolls back | Engineer runs automation scripts | Agent drafts a fix PR; human approves | Agent fixes known failure classes on its own |
| Communicate status | Email to a distribution list | Manual Statuspage update | Bot drafts updates for a human to send | Agent drafts updates from live findings | Routine updates sent automatically, sensitive ones by a human |
| Learn and prevent | Postmortem often skipped or blame-driven | Blameless postmortem written by hand | Bot drafts the timeline and postmortem | Agent drafts postmortem and suggests runbook updates | Agent turns repeat incidents into prevention work |
How does the incident management and on-call team change?
The on-call engineer's job is moving from investigator to reviewer. Instead of opening five dashboards at 3 AM, they read an agent's ranked hypotheses and evidence, approve or reject a proposed fix, and step in when the agent is wrong. That cuts the most expensive part of an incident: senior attention spent on routine diagnosis.
The platform team gains new work. Someone has to decide what each agent can read and change, keep runbooks executable, and make sure telemetry covers the systems agents are asked to reason about. When I described how LoginRadius DevOps grew into a platform team, the shift was from tickets to self-service. The next shift is from self-service for humans to safe, scoped self-service for agents.
Roles that shrink
- Tier-1 NOC operators watching dashboards
- Manual alert triage on every page
- Hand-written incident timelines and postmortems
- Large follow-the-sun rotations for routine failures
Roles that appear
- Agent supervisor who reviews AI SRE findings and actions
- Runbook engineer who writes procedures agents can safely execute
- Reliability workflow designer who sets which failure classes agents may fix alone
- Agent access owner who manages scopes, credentials, and audit for ops agents
Skills to learn
- Judging an agent's evidence quickly and spotting a confident wrong answer
- Writing machine-executable runbooks with clear rollback steps
- Designing SLOs and telemetry that agents can reason over
- Scoping least-privilege access for non-human identities
- Incident command and stakeholder communication for novel failures
What gets easier for the humans?
| Before | After |
|---|---|
| Paged at 3 AM to start an investigation from a blank dashboard | Paged with a ranked root cause and evidence already in the incident channel |
| Hundreds of uncorrelated alerts per incident | One triaged incident with noise dismissed and reasons stated |
| Hunting through deploy history to find what changed | Agent links the alert to the likely deploy or config change |
| Writing the postmortem timeline by hand from chat scrollback | Reviewing a drafted timeline and postmortem built from the incident record |
| Status updates written from memory during the firefight | Drafted updates based on live findings, sent after a quick human check |
Decisions that stay human
- Declaring a major incident and setting its severity when customer impact is unclear
- Approving changes to production data, security controls, or anything irreversible
- Deciding what to tell customers, regulators, and executives
- Leading response to security incidents where the attacker may be steering the evidence
- Choosing which failure classes an agent is trusted to fix alone
Where should agents not act alone?
Risks and failure modes, through a security and identity lens.
- 01
Over-privileged ops agents
An AI SRE agent with broad production write access is a high-value identity. Give each agent its own credentials, least-privilege scopes per action, and short-lived tokens, the same discipline I argue for in [AI agents don't have passwords](/ai-agents-dont-have-passwords/). Read-only by default should be the starting point.
- 02
Prompt injection through telemetry
Agents read logs, tickets, and chat. An attacker who can write a crafted log line or ticket comment can try to steer the agent's conclusion or its proposed fix. Treat everything the agent reads as untrusted input and never let text in telemetry authorize an action.
- 03
Confident wrong root cause
An agent that names the wrong cause with high confidence can send responders down the wrong path or trigger the wrong rollback. Require cited evidence for every finding and keep a human approval gate on mitigation until accuracy is measured per failure class.
- 04
Missing audit trail
If an agent restarts a service or rolls back a deploy, the incident record must show which agent, under which identity, with what evidence, and who approved. Without that, postmortems and compliance reviews cannot reconstruct what happened.
- 05
Security incidents handled as ops incidents
An outage caused by an intrusion looks like an ops failure at first. An agent that restores service by redeploying can destroy forensic evidence. Route anything with a security signal to humans and the SOC before an agent acts.
- 06
Skill atrophy on call
If agents handle every routine incident, engineers lose the hands-on practice they need for the novel ones. Keep game days and chaos exercises so people still know the systems.
Who is building agentic incident management and on-call?
Incumbents adding agents vs agent-native entrants. Capability lines are checked against each vendor's own site.
Incumbents
SRE Agent classifies incidents, surfaces context from past incidents, and recommends remediation steps, in Slack and the Operations Console.
Checked Oct 9, 2026Compare
AI SRE investigates from the moment an incident is declared and names the likely cause with sources; its only change to systems is a pull request a human merges.
Checked Oct 9, 2026Compare
AI SRE ranks likely root causes with confidence scores and suggests fixes; every change requires explicit human sign-off.
Checked Oct 9, 2026Compare
Bits AI SRE investigates alerts using telemetry and runbooks and reports a likely root cause to collaboration tools.
Checked Oct 9, 2026Compare
Handles initial incident triage across AWS, other clouds, and on-premises systems and recommends changes to prevent repeat outages.
Checked Oct 9, 2026
Agentic ITOps platform that uses AI agents for L1 detection, triage, and major incident coordination.
Checked Oct 9, 2026
Agent-native
AI SRE that goes on call, screens alerts, applies mitigations, and escalates to engineers when needed.
Checked Oct 9, 2026
AI SRE that tests many hypotheses in parallel to isolate root cause and offers automated remediation; read-only by default.
Checked Oct 9, 2026
Investigates alerts and production changes, dismisses false positives, and opens fix pull requests for human approval.
Checked Oct 9, 2026
Open source
Apache 2.0 CNCF sandbox SRE agent that investigates incidents with read-only access by default; optional toolsets can apply fixes.
Checked Oct 9, 2026
Side-by-side comparisons: Top 6 Incident Management and On-Call Platforms for 2026: incident.io vs PagerDuty vs Rootly vs Grafana IRM vs Better Stack vs FireHydrant, Top 5 Observability Platforms of 2026: Datadog vs Grafana vs the Rest.
Questions people ask
How is AI changing incident management?
AI SRE agents now start investigating the moment an alert fires. They pull recent deploys, query telemetry, compare past incidents, and post a ranked root cause with evidence. PagerDuty, Datadog, AWS, incident.io, and Rootly all ship this. Most still stop at a recommendation or a pull request for a human to approve.
What is an AI SRE?
An AI SRE is an agent that does the investigation work of a site reliability engineer during an incident: triaging alerts, correlating signals across logs, metrics, traces, and code changes, and proposing a root cause and fix. Vendors include Resolve AI, Traversal, Cleric, incident.io, and Rootly, plus the open-source HolmesGPT.
Will AI agents replace on-call engineers?
Not in the next five years. Agents will absorb routine triage and fixes for known failure classes, so rotations get smaller and quieter. Novel outages, security incidents, and decisions about customer impact still need a human incident commander. The on-call role shifts from investigator to reviewer.
Can AI SRE agents fix incidents automatically?
Some can, but most are configured not to. incident.io limits its agent to pull requests a human merges, and Rootly requires human sign-off on every change. Resolve AI and Traversal describe automated mitigation. In practice, teams allow auto-remediation only for reversible actions on well-understood failure classes.
What happened to Opsgenie?
Atlassian stopped selling Opsgenie on 4 June 2025 and ends support on 5 April 2027, when access shuts off and unmigrated data is deleted. Atlassian points customers to Jira Service Management or Compass; many teams are using the deadline to move to PagerDuty, incident.io, Rootly, or similar tools.
Is it safe to give an AI agent access to production?
Only with guardrails. Give the agent its own identity, least-privilege scopes, and short-lived credentials. Start read-only, log every action with the evidence behind it, and require human approval for anything irreversible. Treat logs and tickets the agent reads as untrusted input because they can carry prompt injection.
What is the difference between AIOps and an AI SRE agent?
AIOps, from tools like BigPanda and Moogsoft, groups and deduplicates alerts to cut noise. An AI SRE agent goes further: it investigates the incident, reasons across telemetry and code changes, and proposes a cause and fix. AIOps tells you what is related; an agent tells you why it broke.
Sources
- Nagios, accessed Oct 9, 2026
- History of Nagios, accessed Oct 9, 2026
- Site reliability engineering, accessed Oct 9, 2026
- PagerDuty, accessed Oct 9, 2026
- PagerDuty Announces Closing of Initial Public Offering, accessed Oct 9, 2026
- Atlassian Acquires Status and Incident Communication Platform StatusPage, accessed Oct 9, 2026
- Splunk Agrees to Acquire VictorOps, accessed Oct 9, 2026
- Jira Ops: incident management platform, accessed Oct 9, 2026
- Dell Technologies Acquires Moogsoft, accessed Oct 9, 2026
- Opsgenie pricing and end of sale FAQ, accessed Oct 9, 2026
- SRE Agent is Generally Available, accessed Oct 9, 2026
- PagerDuty Launches Industry's First End to End AI Agent Suite, accessed Oct 9, 2026
- AI SRE has entered the chat, accessed Oct 9, 2026
- incident.io AI SRE, accessed Oct 9, 2026
- Rootly AI SRE, accessed Oct 9, 2026
- Datadog Launches Bits AI SRE Agent to Resolve Incidents Faster, accessed Oct 9, 2026
- AWS DevOps Agent is now generally available, accessed Oct 9, 2026
- Announcing our Seed and Series A from Sequoia and Kleiner Perkins to Launch the AI SRE for the Enterprise, accessed Oct 9, 2026
- Traversal, accessed Oct 9, 2026
- Resolve AI, accessed Oct 9, 2026
- Ex-Splunk execs' startup Resolve AI hits $1 billion valuation with Series A, accessed Oct 9, 2026
- Cleric, accessed Oct 9, 2026
- HolmesGPT, accessed Oct 9, 2026
- BigPanda, accessed Oct 9, 2026
- Delivering the future of ServiceOps today (Freshworks closes FireHydrant acquisition), accessed Oct 9, 2026
Published Oct 9, 2026. Last verified Oct 9, 2026. Eras 4 and 5, vendors, and scores are re-checked every six to eight weeks; see the changelog and methodology.
Keep reading
Essays and analysis
- Why Incidents Keep Happening (And It's Usually Not What You Think)
- Leveraging AI in DevOps for Non-Linear Scaleup
- How LoginRadius's DevOps Delivered in 2021-2022
- 2022: A Year of Engineering Resilience at LoginRadius
- Everyone Bought AI Observability. Nobody Owns Agent Behavior.
- AI Agents Don't Have Passwords. Your Auth Stack Assumes Everyone Does.
- Seven Patterns of Technology Failure That Repeat Across Decades
- The Post-Agentic Organization: What Becomes Scarce When Intelligence Is Abundant