Skip to main content

AWS for Data

Legacy data formats weren't built for AI.

Make your data agent-ready with AWS support for Apache Iceberg across the entire data & AI stack

color-palette-1_aws-library_graphics_gradient_P1-S1-F_L_1200

Give your agents a single, open source of truth

AI agents need governed access to data spread across data lakes, databases, and applications. Without a unified data foundation in place, data teams end up managing copies, building custom pipelines, and slowing down the AI initiatives that depend on trusted data.

That's why organizations like Comcast, Indeed, and Vanguard
turn to AWS to bring reliability and interoperability to all of their data. With the broadest native Apache Iceberg support of any major cloud provider — spanning storage, streaming, ETL, catalog, and analytics — AWS lets every engine and agent reach the same trusted data so teams can simplify their architecture and ship AI faster.

Why Apache Iceberg?

    Features like ACID transactions, schema evolution, snapshot isolation, and built-in time travel ensure every query sees a complete, consistent view of your data.

    Iceberg allows you to flexibly evolve schemas and change partitioning strategies as your needs change. And because Iceberg is a 100% open-source, community-driven table format, it’s not tied to a single vendor’s roadmap. 

    Consolidate workloads on a single copy of governed data, eliminate redundant ETL pipelines, and reduce storage costs by removing duplicate datasets across siloed systems. 

    With native support across popular frameworks like Apache Spark, Apache Flink, and Presto, Iceberg lets you choose the most performant tool for each task. 

Why choose AWS for Apache Iceberg?

    Amazon S3 stores data in your own account in open formats accessible by any compatible engine. You know exactly where your data lives, can audit access patterns through detailed request-level logging, and retain complete ownership. Amazon S3 delivers industry-leading durability, availability, security, and performance while you retain full control and visibility.

    Learn more about building data lakes on Amazon S3

    AWS Glue Data Catalog and Amazon S3 Tables expose Iceberg REST catalog endpoints, so Iceberg tables on S3 work with any compatible catalog. Catalog federation in AWS Glue lets AWS engines like Amazon Redshift, Amazon Athena, and Amazon EMR query tables cataloged in remote Iceberg catalogs without copying data, giving you a unified view of all your data regardless of where it lives.

    Learn more about catalog federation to remote Iceberg catalogs

Support for every stage of your open data journey

Whether your modernizing from Apache Hive or deploying AI agents at scale, AWS offers support every use case

Foundation for AI

Give AI models and agents governed, versioned access to production data with lineage, fine-grained access control, and time travel built in. Teams can reproduce training runs, audit agent behavior, and roll back when needed, without building custom infrastructure.

Colorful, stacked blocks

Lakehouse modernization

Migrate from legacy Hive tables or proprietary formats to AI-ready Iceberg data lakes and lakehouse architectures without downtime or data rewrites. AWS provides in-place migration paths so your teams can modernize incrementally while workloads keep running.

How Yelp modernized its data infrastructure with a streaming lakehouse on AWS

Enterprise scale in-place migration to Apache Iceberg: Implementation guide 

Abstract digital illustration of flowing vertical lines with network nodes in blue and cyan, fanning outward against a gradient background of pink, purple, and gold, representing data flow or neural network connectivity.

Cost reduction

Stop managing metadata across disconnected catalogs. Centralize critical metadata for a trusted, comprehensive view of data assets and eliminate redundant data copies and complex ETL systems.   

Access Databricks Unity Catalog data using catalog federation in the AWS Glue Data Catalog

Access Snowflake Horizon Catalog data using catalog federation in the AWS Glue Data Catalog

Abstract illustration of iridescent rainbow-gradient blocks arranged on concentric circular platforms.

Petabyte-scale analytics

Run interactive queries across petabytes of Iceberg data without moving it into a warehouse. With multiple engines all safely reading from the same tables, analysts and data engineers choose the right engine for each workload without copying data. 

Abstract illustration of colorful geometric shapes on thin stems rising above layered wavy gradient hills on a dark navy background.

How it works

Notice

Apache and Apache project trademarks are trademarks of The Apache Software Foundation.

Did you find what you were looking for today?

Let us know so we can improve the quality of the content on our pages