AWS Storage Blog

Category: Management Tools

AWS Transfer Family Featured Image

Build AI-powered file classification with AWS Transfer Family

Organizations that receive files from external partners through SFTP face a persistent operational challenge: routing each file to the correct downstream system. Invoices, contracts, images, CSVs, and reports all arrive in a single landing zone, and each requires a different destination. The traditional approach—pattern-matching on file names with regular expressions—is inherently fragile. It relies on […]

Amazon S3 Object Lock

Flexibly control Amazon S3 Object Lock retention based on real business events

Immutability is a foundational data protection control, protecting data in place against unintended changes and deletions by authorized users, and changes by unauthorized users. In many cases, however, it isn’t clear at the time data is written how long it needs to stay immutable, or when that period should begin. A signed contract might need […]

Amazon S3 Express One Zone thumbnail

Run Spark 31% faster and optimize compute costs with Amazon S3 Express One Zone on Amazon EMR

As your Spark datasets grow, storage latency often becomes the constraint, impeding application performance. Query runtimes stretch, and the bottleneck shifts from compute to how fast each node can read from Amazon Simple Storage Service (Amazon S3). We benchmarked this directly on Amazon EMR with TPC-DS at 3 TB scale. On an 8-node Graviton4 cluster […]

Amazon S3 Storage Lens featured image

How WeatherBug reduced storage costs by 80% using Amazon S3 Storage Lens and Kiro CLI

WeatherBug is the third largest weather intelligence company in the US, delivering real-time forecasts, radar, lightning alerts, and interactive maps to over 10 million users. As their data footprint has grown across hundreds of Amazon Simple Storage Service (Amazon S3) buckets in a multi-account AWS environment, their storage costs rose steadily with no clear visibility […]

s3-annotations-header-image

Track healthcare data lineage in real time with Amazon S3 annotations

Organizations that store sensitive data need to answer a simple question: where did this data come from, and what happened to it? In healthcare, this isn’t optional. Regulations like HIPAA, GDPR, and SOX require organizations to produce this information on demand. When an auditor asks, for example, “Show me every transformation that touched patient record […]

Amazon S3 Replication

Accelerate Amazon S3 Replication with automated S3 Batch Operations parallelization

As data volumes grow, organizations must move large datasets between storage locations to meet compliance requirements, optimize performance, and help meet data sovereignty requirements. However, migrating petabytes of data presents significant challenges: lengthy transfer times, complex coordination of parallel operations, data integrity verification, and substantial engineering overhead. Automation reduces resource consumption and operational complexity during […]

EBS feature image

Optimize Amazon EBS volumes to get the right performance at the right time

Business-critical applications depend on reliable storage performance. However, workloads aren’t static and can vary by time of day, day of week, or business cycle. Matching storage performance to these shifting demands is essential for maintaining application reliability without overspending. Under-provisioning can affect your applications when demand peaks, while provisioning too much for too long means […]

Zero-downtime Amazon S3 Versioning: Architectural patterns for mission-critical workloads

Organizations delivering content on a global scale rely on distributed edge networks to cache and serve billions of requests daily. These architectures depend on highly aggressive Time-To-Live (TTL) configurations to maximize performance and minimize origin load. On a cache miss, the network falls through to the origin to retrieve the requested content. At this scale, […]

Figure 1: Solution architecture diagram

Replicate Amazon S3 bucket configurations across AWS Regions with AWS Step Functions

Many organizations operate thousands of Amazon S3 buckets in a single AWS Region, each with its own configuration accumulated over the years. Some were created manually in the AWS Management Console and others by scripts that are no longer actively maintained, provisioned by different business units with their own policies, lifecycle rules, encryption, and tags. […]

Amazon S3 Tables

Query Amazon S3 access logs instantly with CloudWatch and S3 Tables

Knowing who accessed your data, when, and how is the foundation for security investigations, compliance audits, cost attribution, and performance troubleshooting. Detailed access logs capture every request: who made it, which resource was accessed, and what response was returned. In practice, though, they arrive as semi-structured records spread across different locations. Turning them into actionable […]