AWS Cloud Operations Blog
Category: DevOps
Best practices for writing AWS DevOps Agent Skills
When an incident hits at 2 AM, the on-call engineer’s effectiveness depends on what they know about the system, which metrics to check first, what “normal” looks like for this service, and where to find the deployment history. That knowledge often lives in runbooks, internal wikis, and the heads of senior engineers who built the […]
Automate RCA across ServiceNow, Dynatrace and Slack with AWS DevOps Agent
If you manage production incidents, you know the drill. A ServiceNow ticket fires at 2 AM. The on-call engineer wakes up, logs in to Dynatrace, pulls traces and metrics across multiple dashboards, cross-references change records, forms a hypothesis, updates the ticket, and posts findings to Slack. The investigation takes one to three hours, and that’s […]
Introducing Amazon CloudWatch Omni: Observability for the AI Era
An AI-powered observability experience for AI agents and applications, built on open standards and delivered outside of the AWS Console. Organizations are handing agents the keys to their day-to-day operations, from resolving support tickets and managing infrastructure to approving expenses, shipping production code, and increasingly the long tail of workflows that keep the business running. […]
Reduce MTTR with AI-driven RCA using AWS DevOps Agent and Splunk
Modern cloud-native applications generate rich telemetry across metrics, logs, and deployment histories. When performance degrades, operations teams have the data, but the challenge is correlating signals across multiple tools quickly enough to minimize customer impact. Root cause analysis remains a largely manual process dependent on institutional knowledge and operator experience. This post shows how AWS […]
Use AWS DevOps Agent to triage and route AWS Health event impact
Triaging the impact of AWS Health events is one of the most repetitive jobs in cloud operations, and it is exactly the kind of work AWS DevOps Agent can take on. Scheduled maintenance, operational issues, and Trust & Safety notifications (alerts about resources that may violate the AWS Acceptable Use Policy) land in your inbox […]
This Month in AWS Observability: July 2026
Introduction July was a busy month for AWS Observability. We launched features to make the telemetry you already collect more actionable, and take the operational work of collecting it off your hands. Log analytics moved closer to action with alarms that run straight from a log query and enrichment that happens at ingestion. Application-level observability […]
Use CloudWatch syslog and Log Alarms to give AWS DevOps Agent on-premises visibility
Your on-premises firewalls, routers, and switches emit syslog that record device events such as denied connections, tunnel state changes, and routing changes. Network devices send their logs over syslog rather than the Amazon CloudWatch Logs API, so bringing that data into AWS takes extra components. A common approach has been to run a collection tier […]
Build bespoke operational workflows with AWS DevOps Agent custom SRE agents
The production standards that matter most are often the ones no general purpose tool is built to check. For example, the read replica behind your customer dashboard can’t lag more than five seconds behind the primary, but the analytics replica can tolerate lag during peak ingestion. The nightly extract, transform, and load (ETL) job should […]
This Month in AWS Observability: June 2026
Introduction Welcome to the latest edition of This Month in AWS Observability, featuring what’s new across Amazon CloudWatch and AI-driven operations this June! Native OpenTelemetry metrics with PromQL querying is now generally available in CloudWatch, 23 new Logs Insights commands launched for deeper statistical and structured analysis, Session Replay now in CloudWatch RUM, and AWS […]
Announcing General Availability of AWS DevOps Agent
Today, we’re announcing the general availability of AWS DevOps Agent. AWS DevOps Agent is your always-available operations teammate. It resolves and proactively prevents incidents, optimizes application reliability and performance, and handles on-demand SRE tasks across AWS, multicloud, and on-premises environments. Operations teams spend countless hours investigating incidents, correlating data across multiple tools, and manually triaging […]








