Legacy data formats weren't built for AI.
Make your data agent-ready with AWS support for Apache Iceberg across the entire data & AI stack
Give your agents a single, open source of truth
AI agents need governed access to data spread across data lakes, databases, and applications. Without a unified data foundation in place, data teams end up managing copies, building custom pipelines, and slowing down the AI initiatives that depend on trusted data.
That's why organizations like Comcast, Indeed, and Vanguard
turn to AWS to bring reliability and interoperability to all of their data. With the broadest native Apache Iceberg support of any major cloud provider — spanning storage, streaming, ETL, catalog, and analytics — AWS lets every engine and agent reach the same trusted data so teams can simplify their architecture and ship AI faster.
Why Apache Iceberg?
Features like ACID transactions, schema evolution, snapshot isolation, and built-in time travel ensure every query sees a complete, consistent view of your data.
Iceberg allows you to flexibly evolve schemas and change partitioning strategies as your needs change. And because Iceberg is a 100% open-source, community-driven table format, it’s not tied to a single vendor’s roadmap.
Consolidate workloads on a single copy of governed data, eliminate redundant ETL pipelines, and reduce storage costs by removing duplicate datasets across siloed systems.
With native support across popular frameworks like Apache Spark, Apache Flink, and Presto, Iceberg lets you choose the most performant tool for each task.
Why choose AWS for Apache Iceberg?
Amazon S3 stores data in your own account in open formats accessible by any compatible engine. You know exactly where your data lives, can audit access patterns through detailed request-level logging, and retain complete ownership. Amazon S3 delivers industry-leading durability, availability, security, and performance while you retain full control and visibility.
Apache Iceberg brings reliability and simplicity to your data lakes, but managing Iceberg tables can take hours of engineering time. Amazon S3 Tables offer purpose-built Iceberg storage that streamlines governance and automates routine tasks like compaction and snapshot management, returning hours of productivity to your team.
Learn more about Amazon S3 Tables automated maintenance for Apache Iceberg
With AWS engines, you never sacrifice performance for cost. Iceberg materialized views in AWS Glue accelerate query performance from Apache Spark up to 8x, Amazon EMR runs Apache Iceberg write jobs over 2x faster than open-source equivalents, and Amazon Redshift RG instances run Apache Iceberg workloads up to 2.4x faster than previous generations at up to 30% lower cost.
Learn more about working with Apache Iceberg using AWS engines and services
Openness does not come at the expense of governance. AWS allows you to give every engine and user a consistent view of what tables exist, where they live, and how they're structured. With AWS Lake Formation, define fine-grained access policies once and they're enforced consistently across AWS services.
Learn more about AWS Lake Formation for Apache Iceberg tables
AWS Glue Data Catalog and Amazon S3 Tables expose Iceberg REST catalog endpoints, so Iceberg tables on S3 work with any compatible catalog. Catalog federation in AWS Glue lets AWS engines like Amazon Redshift, Amazon Athena, and Amazon EMR query tables cataloged in remote Iceberg catalogs without copying data, giving you a unified view of all your data regardless of where it lives.
Learn more about catalog federation to remote Iceberg catalogs
Support for every stage of your open data journey
Whether your modernizing from Apache Hive or deploying AI agents at scale, AWS offers support every use case
Foundation for AI
Give AI models and agents governed, versioned access to production data with lineage, fine-grained access control, and time travel built in. Teams can reproduce training runs, audit agent behavior, and roll back when needed, without building custom infrastructure.
Lakehouse modernization
Migrate from legacy Hive tables or proprietary formats to AI-ready Iceberg data lakes and lakehouse architectures without downtime or data rewrites. AWS provides in-place migration paths so your teams can modernize incrementally while workloads keep running.
How Yelp modernized its data infrastructure with a streaming lakehouse on AWS
Enterprise scale in-place migration to Apache Iceberg: Implementation guide
Cost reduction
Stop managing metadata across disconnected catalogs. Centralize critical metadata for a trusted, comprehensive view of data assets and eliminate redundant data copies and complex ETL systems.
Access Databricks Unity Catalog data using catalog federation in the AWS Glue Data Catalog
Access Snowflake Horizon Catalog data using catalog federation in the AWS Glue Data Catalog
Petabyte-scale analytics
Run interactive queries across petabytes of Iceberg data without moving it into a warehouse. With multiple engines all safely reading from the same tables, analysts and data engineers choose the right engine for each workload without copying data.
Real customer impact
How it works
AWS data and analytics services for Apache Iceberg
Next steps
Step-by-step guidance for your Iceberg architecture
Notice
Apache and Apache project trademarks are trademarks of The Apache Software Foundation.
Did you find what you were looking for today?
Let us know so we can improve the quality of the content on our pages