AWS Cloud Operations Blog

Category: Best Practices

Share Systems Manager documents across your AWS Organization with AWS RAM

Introduction If you operate AWS Systems Manager (SSM) documents, such as Automation runbooks and Run Command documents, across an AWS Organization, you must decide which accounts can run them. Until now, that meant sharing each document with a list of individual account IDs through the ModifyDocumentPermission API, which accepts up to 1,000 account IDs per […]

Best practices for writing AWS DevOps Agent Skills

When an incident hits at 2 AM, the on-call engineer’s effectiveness depends on what they know about the system, which metrics to check first, what “normal” looks like for this service, and where to find the deployment history. That knowledge often lives in runbooks, internal wikis, and the heads of senior engineers who built the […]

Automate RCA across ServiceNow, Dynatrace and Slack with AWS DevOps Agent

If you manage production incidents, you know the drill. A ServiceNow ticket fires at 2 AM. The on-call engineer wakes up, logs in to Dynatrace, pulls traces and metrics across multiple dashboards, cross-references change records, forms a hypothesis, updates the ticket, and posts findings to Slack. The investigation takes one to three hours, and that’s […]

A Practical Guide to Amazon CloudWatch Logs Cost Optimization

Introduction As organizations centralize logs from AWS services, on-premises infrastructure, and third-party sources into CloudWatch Logs, storage becomes the dominant cost driver. Ingestion is one-time, but storage charges accumulate for the lifetime of the data, often years when compliance mandates like PCI-DSS or HIPAA apply. This post shows you how to use built-in CloudWatch Logs […]

How Amazon Achieved Full Stack Observability Across 400 Offices with Amazon OpenSearch Serverless

Summary Amazon Corporate Infrastructure Services (CIS) supports 330,000+ employees across 400 offices in 50+ countries. Every day, these employees depend on critical services including video conferencing (Zoom, Webex, Teams), collaboration tools (Slack, Microsoft M365), audiovisual systems, guest Wi-Fi, and ServiceNow for seamless productivity. Before we implemented the Full Stack Observability (FSO) platform, individual teams had […]

Turn Your Amazon CloudWatch Alarms into Actionable Signals

Your alarm fires at 2 AM. You grab your phone, squint at the notification, and see: “ALARM: my-service-alarm has transitioned to ALARM state.” No context. No application. No hint about which of your 200 instances is the problem, or whether it even matters. I’ve been there. We’ve all been there. Alarm frustration often comes from […]

GPU Cost Attribution in Amazon EKS Using Amazon Managed Service for Prometheus, Amazon Managed Grafana, and OpenTelemetry

As organizations scale their AI and machine learning workloads on Amazon Elastic Kubernetes Service (Amazon EKS), GPU instances often represent the largest portion of compute costs. Without granular visibility into how these resources are consumed, teams struggle to attribute costs accurately. Consider a shared EKS cluster where Team A (Research) runs experimental ML models on […]

Shift-Left Tag Compliance using AWS Organizations and Terraform

In this post you will learn about AWS Organizations tag policies, the tag_policy_compliance Terraform provider setting, a reusable tagging module that automatically applies required tags, and a test-driven approach that dynamically validates against your organizational policies.

Adaptive sampling with AWS X-Ray to capture critical spans

Introduction Enterprise applications using AWS X-Ray generate large volumes of distributed tracing data across multiple services. Static sampling strategies keep costs down by capturing a fixed percentage of traffic. However, they frequently miss critical data during intermittent failures or sudden latency spikes. Tracing every request for maximum visibility at scale may increase sampling costs for […]