AWS Storage Blog
Category: Compute
Build AI-powered file classification with AWS Transfer Family
Organizations that receive files from external partners through SFTP face a persistent operational challenge: routing each file to the correct downstream system. Invoices, contracts, images, CSVs, and reports all arrive in a single landing zone, and each requires a different destination. The traditional approach—pattern-matching on file names with regular expressions—is inherently fragile. It relies on […]
How Tubular Labs reclaimed 50% of engineering capacity by rebuilding their 70TB pipeline on Apache Iceberg and Amazon S3 Tables
Customer Story | Amazon S3 Tables – Learn how Tubular Labs (part of Chartbeat Inc.) reclaimed 50% of engineering capacity by replacing fragile, file-based data pipelines with a Common Pipeline Runtime (CPR) built on Apache Iceberg and Amazon S3 Tables. By Tubular Labs Engineering (part of Chartbeat Inc.), in collaboration with the AWS Solution Architecture team […]
Run Spark 31% faster and optimize compute costs with Amazon S3 Express One Zone on Amazon EMR
As your Spark datasets grow, storage latency often becomes the constraint, impeding application performance. Query runtimes stretch, and the bottleneck shifts from compute to how fast each node can read from Amazon Simple Storage Service (Amazon S3). We benchmarked this directly on Amazon EMR with TPC-DS at 3 TB scale. On an 8-node Graviton4 cluster […]
How Nearmap built continental-scale aerial search using Amazon S3 Vectors
Nearmap captures high-resolution aerial imagery across populated areas of the United States, Canada, Australia, and New Zealand several times a year, at resolutions as fine as 1.5 inches. More than 10,000 customers globally use these images to assess insurance portfolios, size solar arrays, and track construction sites. Since 2007, Nearmap has completed more than 35,000 […]
Orchestrating multi-agent AI architectures with Amazon S3 Files
Organizations are moving beyond single-model AI toward multi-agent architectures. In these systems, agents offload intermediate results to files rather than carrying everything in the prompt, because a large prompt inflates cost and degrades quality. A model’s context window is finite, so files become working memory that persists after a session ends. In multi-agent systems, a […]
Hybrid ML inferencing on Amazon EKS with Amazon FSx for NetApp ONTAP and on-premises NetApp
Machine learning (ML) models used for inference on Kubernetes are often several gigabytes in size. When these models are embedded in container images, images become oversized and pod scheduling slows. More critically, inference pods are inherently stateful. Model weights, tokenizer files, compiled GPU kernels, and runtime caches must persist across pod restarts, node failures, and […]
Optimize your self-managed PostgreSQL data warehouse with Amazon FSx for OpenZFS
A data warehouse is the analytical backbone of a modern enterprise, consolidating data from disparate sources into a single, authoritative view that enables complex queries, trend analysis, and confident decision-making. In financial services, this means sharper regulatory reporting, faster fraud detection, and deeper customer understanding. The operational reality is demanding. Enterprises juggle multiple source databases […]
Track healthcare data lineage in real time with Amazon S3 annotations
Organizations that store sensitive data need to answer a simple question: where did this data come from, and what happened to it? In healthcare, this isn’t optional. Regulations like HIPAA, GDPR, and SOX require organizations to produce this information on demand. When an auditor asks, for example, “Show me every transformation that touched patient record […]
Accelerate Amazon S3 Replication with automated S3 Batch Operations parallelization
As data volumes grow, organizations must move large datasets between storage locations to meet compliance requirements, optimize performance, and help meet data sovereignty requirements. However, migrating petabytes of data presents significant challenges: lengthy transfer times, complex coordination of parallel operations, data integrity verification, and substantial engineering overhead. Automation reduces resource consumption and operational complexity during […]
Simplify compliance-driven Amazon S3 data movement with multi-criteria filtering
Organizations routinely need to move a subset of their stored data from one location to another, such as a compliance audit that requires all PDFs from a specific quarter, a company reorganization that splits one tenant’s records into an isolated bucket, or a regulatory mandate that demands financial documents older than 7 years be archived […]





