AWS Cloud Operations Blog
Category: Best Practices
Share Systems Manager documents across your AWS Organization with AWS RAM
Introduction If you operate AWS Systems Manager (SSM) documents, such as Automation runbooks and Run Command documents, across an AWS Organization, you must decide which accounts can run them. Until now, that meant sharing each document with a list of individual account IDs through the ModifyDocumentPermission API, which accepts up to 1,000 account IDs per […]
Best practices for writing AWS DevOps Agent Skills
When an incident hits at 2 AM, the on-call engineer’s effectiveness depends on what they know about the system, which metrics to check first, what “normal” looks like for this service, and where to find the deployment history. That knowledge often lives in runbooks, internal wikis, and the heads of senior engineers who built the […]
Automate RCA across ServiceNow, Dynatrace and Slack with AWS DevOps Agent
If you manage production incidents, you know the drill. A ServiceNow ticket fires at 2 AM. The on-call engineer wakes up, logs in to Dynatrace, pulls traces and metrics across multiple dashboards, cross-references change records, forms a hypothesis, updates the ticket, and posts findings to Slack. The investigation takes one to three hours, and that’s […]
A Practical Guide to Amazon CloudWatch Logs Cost Optimization
Introduction As organizations centralize logs from AWS services, on-premises infrastructure, and third-party sources into CloudWatch Logs, storage becomes the dominant cost driver. Ingestion is one-time, but storage charges accumulate for the lifetime of the data, often years when compliance mandates like PCI-DSS or HIPAA apply. This post shows you how to use built-in CloudWatch Logs […]
How Amazon Achieved Full Stack Observability Across 400 Offices with Amazon OpenSearch Serverless
Summary Amazon Corporate Infrastructure Services (CIS) supports 330,000+ employees across 400 offices in 50+ countries. Every day, these employees depend on critical services including video conferencing (Zoom, Webex, Teams), collaboration tools (Slack, Microsoft M365), audiovisual systems, guest Wi-Fi, and ServiceNow for seamless productivity. Before we implemented the Full Stack Observability (FSO) platform, individual teams had […]
Deploy OpenTelemetry Gateway on AWS: Monitoring Your Observability Pipeline
Deploy an OpenTelemetry gateway on Amazon EKS, export metrics to CloudWatch over native OTLP, and monitor the pipeline’s own health with PromQL dashboards and alarms.
Turn Your Amazon CloudWatch Alarms into Actionable Signals
Your alarm fires at 2 AM. You grab your phone, squint at the notification, and see: “ALARM: my-service-alarm has transitioned to ALARM state.” No context. No application. No hint about which of your 200 instances is the problem, or whether it even matters. I’ve been there. We’ve all been there. Alarm frustration often comes from […]
GPU Cost Attribution in Amazon EKS Using Amazon Managed Service for Prometheus, Amazon Managed Grafana, and OpenTelemetry
As organizations scale their AI and machine learning workloads on Amazon Elastic Kubernetes Service (Amazon EKS), GPU instances often represent the largest portion of compute costs. Without granular visibility into how these resources are consumed, teams struggle to attribute costs accurately. Consider a shared EKS cluster where Team A (Research) runs experimental ML models on […]
Shift-Left Tag Compliance using AWS Organizations and Terraform
In this post you will learn about AWS Organizations tag policies, the tag_policy_compliance Terraform provider setting, a reusable tagging module that automatically applies required tags, and a test-driven approach that dynamically validates against your organizational policies.
Adaptive sampling with AWS X-Ray to capture critical spans
Introduction Enterprise applications using AWS X-Ray generate large volumes of distributed tracing data across multiple services. Static sampling strategies keep costs down by capturing a fixed percentage of traffic. However, they frequently miss critical data during intermittent failures or sudden latency spikes. Tracing every request for maximum visibility at scale may increase sampling costs for […]








