On 5 April 2026 Microsoft published, in its official blog (the Customer Zero series), the results of Azure SRE Agent, an AI agent it uses in its own cloud operations.
What makes this case worth reading is not how large the numbers are, but the conditions behind them and where human approval sits. Microsoft says the agent autonomously handled more than 35,000 incidents in the last nine months, yet it does not define what "handled autonomously" means. This article separates what the official sources state from what they do not.
This article covers only what Microsoft's official blog and documentation say as of 6 October 2026. It contains no guessed architecture.
What is running
In the official description Azure SRE Agent is "an AI-powered operations agent that serves as an always-on SRE partner for engineers". It continuously observes production environments to detect and investigate incidents, and it reasons across logs, metrics, code changes and deployment records to analyse root causes. It supports engineers from triage to resolution, at autonomy levels from assistive investigation to automated remediation proposals.
It reached general availability on 10 March 2026 (per the GA announcement of that day). Microsoft uses it in its own operations and shares the lessons in the Customer Zero series.
The numbers Microsoft published
- ▸In the last nine months (as of the April 2026 blog) Azure SRE Agent autonomously handled more than 35,000 incidents
- ▸More than 50,000 developer hours were saved (by reducing manual investigation and response work)
- ▸Azure App Service live-site incidents reached time-to-mitigation of 3 minutes, against a human-only average of 40.5 hours
- ▸Azure Container Apps engineers gave overwhelmingly positive (89%) responses to the agent's root-cause results, covering over 90% of incidents
- ▸Microsoft also says on-call burden fell and time-to-mitigation during incidents improved
These figures are as stated in the official blog, but Microsoft's own announcements differ in numbers and wording by date.
The GA announcement of 10 March 2026 says that across its own services Microsoft has "1,300+ agents deployed, 35,000+ incidents mitigated, and over 20,000 engineering hours saved".
The Customer Zero blog of 5 April says 35,000+ incidents "handled autonomously" and 50,000+ developer hours saved.
The incident count stays at "35,000+" while the hours move from 20,000+ to 50,000+. The official materials do not say whether the cut-off date or the scope differs.
How to read the numbers (this site's view)
- ▸"Handled autonomously" is not defined. The blog does not say whether it means incidents resolved end to end or incidents where the agent investigated or proposed a fix on its own. The same "35,000+" is called "mitigated" in March and "handled autonomously" in April, so it cannot be confirmed that the same thing is being counted
- ▸The comparison behind "3 minutes" versus "40.5 hours" is not given either. There is no statement on whether the incidents are the same kind, whether it is a mean or a median, or what period the human-only figure covers. Taking "40.5 hours to 3 minutes" as an improvement you will get yourself is risky
- ▸Broad figures and single-team figures sit side by side. 3 minutes is App Service and 89% is Container Apps, each a per-product figure
- ▸This is the vendor reporting on its own operations. It is not an independent evaluation. If you adopt it, measure again on your own incidents
Architecture: what can be confirmed
What Microsoft publishes is the product's structure and usage; the internal implementation (which models, the internal structure) is not published. The following is what the official blog and documentation confirm.
The five extension points and integrations
- ▸Skills: runbooks and Azure CLI scripts that extend what the agent can do without custom code
- ▸Custom agents: purpose-built agents for operational domains. Some ship ready to use and you can build your own
- ▸Python tools: custom logic and API integrations for cases that need code rather than configuration
- ▸MCP servers: connect observability platforms such as Datadog, Splunk, New Relic, Dynatrace and Elasticsearch, or any custom tool through the Model Context Protocol
- ▸Agent hooks: event-triggered automations at lifecycle points, such as after a tool runs or when the agent stops. Used to enforce policies, emit telemetry or integrate with external approval workflows
- ▸Integrations: Azure Monitor, Application Insights and Log Analytics; PagerDuty and ServiceNow; GitHub and Azure DevOps; Azure Data Explorer; Teams and Outlook, among others
Every tool call the agent proposes passes through governance controls before it runs. The documentation says your team defines the boundaries even for fully automated workflows.
Where human approval sits
- ▸You choose the level of autonomy: from assisting an investigation to proposing fixes and, if configured, acting autonomously. Whether it mitigates automatically or waits for approval depends on the configured "run mode"
- ▸Review mode: an administrator reviews the investigation summary with runbook context and approves actions that need approval
- ▸Per-tool permissions: each tool can be set to allow, ask or deny. Admins set global guardrails, team leads customise each custom agent, and users approve tools within their conversation
- ▸Identity and access: it authenticates with a managed identity and operates under Azure role-based access control (RBAC)
- ▸Network: virtual network (VNet) integration can route the agent's outbound traffic through your own network. It was announced as a preview at Build 2026 in June and became generally available on 25 August 2026
- ▸Infrastructure as Code: the agent, its network configuration, identity and tool policies can be deployed with Bicep templates
Microsoft's own lesson pairs autonomy with approval: autonomy "had to be balanced with clear approval boundaries, role-based access, and safety checks to build trust".
Building agents with agents
The blog also explains that Azure SRE Agent itself was built using agentic workflows: specialised agents at each stage of the software development lifecycle.
- ▸Plan and code: spec-driven development, with agents helping draft requirement documents, prototypes and check code into staging
- ▸Verify, test and deploy: agents for code review, security, evaluation and deployment shift quality and security issues left
- ▸Operate and optimise: Azure SRE Agent handles alert investigation, remediation help and some autonomous resolution. It also has its own specialised instance to maintain itself
The lessons Microsoft lists
- ▸Building agents with agents is essential to scaling. Manual development quickly became a bottleneck
- ▸A generic agent with rich context, memory and learning improves with experience. Alongside it, specialised agents for well-defined categories of incidents encode proven workflows and safeguards for repeatability
- ▸Integrate deeply with existing systems. Add intelligence on top of existing telemetry, workflows and platforms rather than replacing them
- ▸Keep humans in the loop. Balancing autonomy with approval boundaries, access control and safety checks was key to trust
- ▸Invest in continuous feedback and evaluation. Measure where automation added value and where human judgement should stay central
What could not be confirmed
- ▸The models used and Microsoft's internal structure. Neither the blog nor the documentation states them
- ▸The definition of "handled autonomously", the basis of the "3 minutes" versus "40.5 hours" comparison, and a breakdown by incident type
- ▸Updates to the figures after April 2026. The June Build 2026 post says only that the footprint grew fast and gives no new numbers for the same metrics
- ▸Cost, and the operating burden of the agent itself. The official materials do not give them
What to take away (this site's view)
- ▸Decide autonomy and approval in the same design. Here every proposed tool call passes per-tool allow, ask or deny controls, and Review mode can insert approval. Widening automation and setting the stopping boundary are designed together
- ▸Take the numbers with their conditions. If you borrow "35,000" or "3 minutes", also check what was counted and what it was compared with. The official blog does not say
- ▸Plug into existing operations tools. It connects alerts, tickets, repositories and observability platforms, adding to operating procedures rather than replacing them
- ▸Measure again yourself. This is the vendor's own report. In your first month, start with a few kinds of incidents and compare against human time
Sources
- Microsoft Tech Community: How we build and use Azure SRE Agent with agentic workflows (Customer Zero series; published 5 April 2026, updated 9 April; the source of the figures)↗
- Microsoft Learn: Overview of Azure SRE Agent (extension points, integrations, security and governance, run modes)↗
- Microsoft Tech Community: Announcing general availability for the Azure SRE Agent (10 March 2026; general availability and the internal figures at that date)↗
- Microsoft Tech Community: Azure SRE Agent at Microsoft Build 2026 (2 June 2026; five enterprise releases, VNet integration in preview)↗
- Microsoft Tech Community: Azure SRE Agent VNet integration is now generally available (25 August 2026)↗
- Microsoft Azure: Azure SRE Agent (product page)↗
