AWS Big Data Blog
Category: Announcements
Query across accounts and table formats with multi-catalog in Amazon EMR
Amazon EMR 8.1.0 adds multi-catalog support through the RedirectingSessionCatalog, so you can query Iceberg, Delta Lake, Hudi, and Hive tables through a single catalog and join data across AWS accounts without copying it. This post shows how to put these capabilities into practice on Amazon EMR Serverless.
Materialize once, query anywhere: Introducing Iceberg materialized views in Amazon Redshift
Amazon Redshift now supports Iceberg materialized views. Compute an aggregation once in Amazon Redshift and store the result as a standard Apache Iceberg table in Amazon S3, queryable by Amazon Athena, Apache Spark, Amazon SageMaker, and AWS Glue. This post covers use cases, incremental refresh, and a step-by-step getting-started guide.
Announcing Spark Connect on Amazon EMR on EKS: Interactive PySpark development, anywhere
Announcing Spark Connect on Amazon EMR on EKS: build, test, and debug Spark applications from VS Code, PyCharm, Jupyter notebooks, Amazon SageMaker Unified Studio, or dbt, while running full-scale Spark operations on your existing Amazon EKS clusters.
Announcing Spark Connect on Amazon EMR on EC2: Interactive PySpark anywhere
Amazon EMR on EC2 now supports Spark Connect, so you can develop and debug PySpark interactively from Amazon SageMaker Unified Studio Data Notebooks or your own IDE while Spark runs on your cluster. This post shows you how to get started from both a Data Notebook and a local IDE.
Connect Amazon SageMaker Unified Studio to Microsoft Power BI – Part 1: IAM Identity Center (IDC)-based domains
Connect Microsoft Power BI directly to governed data in Amazon SageMaker Unified Studio using new authentication modes in the Amazon Athena ODBC driver, with no third-party ODBC-JDBC bridge. Part 1 covers IAM Identity Center (IDC)-based domains with both DSN-based and DSN-less connection methods.
Connect Amazon SageMaker Unified Studio to Microsoft Power BI – Part 2: IAM-based domains
Connect Microsoft Power BI directly to governed data in Amazon SageMaker Unified Studio using the Amazon Athena ODBC driver. Part 2 covers IAM-based domains with SageMakerIam authentication, including AWS IAM Identity Center administrator setup, for both DSN-based and DSN-less connection methods.
Scale down Kinesis Data Streams on-demand capacity with ODA warm throughput
Amazon Kinesis Data Streams now supports scaling down ingest capacity for on-demand Advantage streams with warm throughput. Learn how the scale-down works, how to monitor stream behavior with Amazon CloudWatch, and best practices for releasing excess capacity after transient traffic bursts.
Every team is a data team — bring Amazon Redshift analytics to ChatGPT Work
AWS is announcing the AWS Data Analytics plugin for the new Data agent in ChatGPT Work. Teams can ask questions in natural language, analyze governed data across their Amazon Redshift data warehouse and data lakes, and build shareable dashboards, all from a conversation in ChatGPT Work.
Deliver real-time data to streaming tables for Apache Iceberg with Amazon Kinesis Data Streams
Amazon Kinesis Data Streams now supports streaming tables, a fully managed capability that continuously delivers your streaming data as queryable Apache Iceberg tables on Amazon S3 Tables. Streaming tables reduce data delivery costs to S3 Tables by up to 50% compared to self-managed alternatives and reduce downstream query costs by up to 30% through intelligent inline compaction that eliminates the small file problem. You need no custom applications, no self-managed compute, and no operational overhead.
Integrate Amazon Redshift and IAM Identity Center with enhanced VPC routing
Amazon Redshift now supports AWS IAM Identity Center authentication on clusters and workgroups that use enhanced VPC routing. Create two interface VPC endpoints to give your users single sign-on with their corporate credentials while keeping all authentication traffic on the AWS private network.









