Build cluster · SRE, platform engineering
Technology Rewired: DevOps and Platform Engineering
DevOps went from racked servers, cron jobs, and pagers to cloud infrastructure defined in code and observed through SaaS dashboards. AI added anomaly detection and noise reduction. Agents now triage incidents, correlate signals, and draft fixes, while engineers keep approval over changes to production. Fully self-healing infrastructure is still narrow.
The shift: Self-healing infra; agents run incident triage and remediation.
Supervised agents today
3.7 in five years
How has the devops and platform engineering team changed across five eras?
Era 1 · On-prem
Before 2005
0.2Operations was a separate team that racked servers, patched them, and carried pagers. Developers threw releases over the wall, and change advisory boards approved deployments in weekly meetings. Monitoring meant threshold alerts that woke someone up.
Era 2 · SaaS and cloud
2005 to 2020
0.8Cloud and infrastructure as code turned operations into software work. DevOps and SRE teams owned CI/CD pipelines, observability stacks, and on-call rotations shared with developers. Platform engineering emerged to build internal paved roads so product teams could ship without filing tickets.
Era 3 · AI-assisted
2020 to 2024
1.5Observability platforms added anomaly detection, alert grouping, and log summaries. Fewer pages fired, but an engineer still investigated every incident, wrote the timeline, and ran the fix.
Era 4 · Agentic
2024 onward
2.5AI SRE agents now pick up an alert, pull metrics, logs, traces, and recent deploys, propose a likely cause, and draft a remediation or rollback for a human to approve. Platform teams are adding agent workflows to their paved roads. The open problem is accountability: everyone bought AI observability, nobody owns agent behavior.
Incidents still come mostly from change, process, and ownership gaps, which I cover in why incidents keep happening. Agents speed up diagnosis; they do not fix an unclear ownership model.
Era 5 · Next 5 years
2026 to 2031
3.7My bet: agents handle most first response and known-pattern remediation, including rollbacks, scaling, and failing-build fixes, under policies that platform teams write and audit. On-call becomes supervising the agent's decisions and handling novel failures. Teams stay small and get more senior.
Which devops and platform engineering software is being rewired?
- Incident Management and On-Call
Incident management moved from pagers and Nagios checks to SaaS paging with PagerDuty and Opsgenie, then to AIOps noise reduction. The agentic shift: AI SRE agents now triage alerts, investigate across telemetry and deploys, and name a likely root cause before a human opens a laptop. Fixes still wait for human approval.
Coming next
- CI/CD and releasewave 2
- Observability and monitoringwave 2
- Infrastructure as code
- Cloud cost management (FinOps)
- Logging
- Kubernetes management
- Internal developer portals
- Database operations
- Backup and disaster recovery
Also wired into this category: Coding Assistants and IDEs, Firewall and Network Security, SIEM and SOC Platforms.
Keep reading
Essays and analysis
- Everyone Bought AI Observability. Nobody Owns Agent Behavior.
- Why Incidents Keep Happening (And It's Usually Not What You Think)
- Leveraging AI in DevOps for Non-Linear Scaleup
- Critical Controls for DevOps: Best Practices for Continuous Security
- The Post-Agentic Organization: What Becomes Scarce When Intelligence Is Abundant
What died (Tech Graveyard)
Comparisons
- Top 6 Incident Management and On-Call Platforms for 2026: incident.io vs PagerDuty vs Rootly vs Grafana IRM vs Better Stack vs FireHydrant
- Top 5 Observability Platforms of 2026: Datadog vs Grafana vs the Rest
- Top 5 CI/CD Platforms for 2026: GitHub Actions vs GitLab vs CircleCI vs Jenkins vs Buildkite
- Top 6 Internal Developer Portals for 2026: Backstage vs Port vs Cortex vs OpsLevel vs Spotify Portal vs Humanitec