Skip to content
Draft. This page is in editorial review and is not indexed yet.

Build cluster · SRE, platform engineering

Technology Rewired: DevOps and Platform Engineering

DevOps went from racked servers, cron jobs, and pagers to cloud infrastructure defined in code and observed through SaaS dashboards. AI added anomaly detection and noise reduction. Agents now triage incidents, correlate signals, and draft fixes, while engineers keep approval over changes to production. Fully self-healing infrastructure is still narrow.

The shift: Self-healing infra; agents run incident triage and remediation.

Verified
2.5

Supervised agents today
3.7 in five years

How has the devops and platform engineering team changed across five eras?

  1. Era 1 · On-prem

    Before 2005

    0.2

    Operations was a separate team that racked servers, patched them, and carried pagers. Developers threw releases over the wall, and change advisory boards approved deployments in weekly meetings. Monitoring meant threshold alerts that woke someone up.

  2. Era 2 · SaaS and cloud

    2005 to 2020

    0.8

    Cloud and infrastructure as code turned operations into software work. DevOps and SRE teams owned CI/CD pipelines, observability stacks, and on-call rotations shared with developers. Platform engineering emerged to build internal paved roads so product teams could ship without filing tickets.

  3. Era 3 · AI-assisted

    2020 to 2024

    1.5

    Observability platforms added anomaly detection, alert grouping, and log summaries. Fewer pages fired, but an engineer still investigated every incident, wrote the timeline, and ran the fix.

  4. Era 4 · Agentic

    2024 onward

    2.5

    AI SRE agents now pick up an alert, pull metrics, logs, traces, and recent deploys, propose a likely cause, and draft a remediation or rollback for a human to approve. Platform teams are adding agent workflows to their paved roads. The open problem is accountability: everyone bought AI observability, nobody owns agent behavior.

    Incidents still come mostly from change, process, and ownership gaps, which I cover in why incidents keep happening. Agents speed up diagnosis; they do not fix an unclear ownership model.

  5. Era 5 · Next 5 years

    2026 to 2031

    3.7

    My bet: agents handle most first response and known-pattern remediation, including rollbacks, scaling, and failing-build fixes, under policies that platform teams write and audit. On-call becomes supervising the agent's decisions and handling novel failures. Teams stay small and get more senior.

Which devops and platform engineering software is being rewired?

  • 2.5
    Incident Management and On-Call

    Incident management moved from pagers and Nagios checks to SaaS paging with PagerDuty and Opsgenie, then to AIOps noise reduction. The agentic shift: AI SRE agents now triage alerts, investigate across telemetry and deploys, and name a likely root cause before a human opens a laptop. Fixes still wait for human approval.

Coming next

  • CI/CD and releasewave 2
  • Observability and monitoringwave 2
  • Infrastructure as code
  • Cloud cost management (FinOps)
  • Logging
  • Kubernetes management
  • Internal developer portals
  • Database operations
  • Backup and disaster recovery

Also wired into this category: Coding Assistants and IDEs, Firewall and Network Security, SIEM and SOC Platforms.