AWS Big Data Blog
Category: Announcements
GPU-accelerated Apache Spark with Amazon EMR and NVIDIA RTX PRO 4500 on Amazon EC2 G7 instances runs up to 3.7x faster
Amazon EMR on EKS now runs Apache Spark up to 3.7x faster on Amazon EC2 G7 instances with NVIDIA RTX PRO 4500 Blackwell GPUs than on comparable CPU instances, with no changes to existing Spark code. See the TPC-DS benchmark results, the cost comparison, and how to get started.
Introducing AWS Glue 6.0 for faster and more cost-effective data integration
AWS Glue 6.0 is now available, lowering AWS Glue pricing by 30%, adding an AWS optimized build of Apache Spark 4.1, and introducing Apache Iceberg V3 capabilities suitable for enterprise adoption. This post covers the key capabilities and performance benefits, with code examples to help you get started.
Upgrade AWS Glue jobs to Glue 6.0 with AI-powered Spark upgrades
Walk through upgrading a PySpark ETL job from AWS Glue 5.1 to AWS Glue 6.0 using the generative AI upgrades for Apache Spark. The upgrade analysis automatically detects incompatibilities, applies fixes, and validates results with data quality checks.
Long-term system tables retention in Amazon Redshift with Amazon S3 Tables
Amazon Redshift system table integration with Amazon S3 Tables automatically delivers your system table logs to Amazon S3 Tables in Apache Iceberg format. You can retain this data well beyond the 7-day limit for compliance, auditing, and cross-warehouse observability, without custom ETL pipelines or cluster resource consumption.
Amazon MSK simplifies configuring custom domain names
With Amazon MSK, you can now configure custom domain names for provisioned clusters using a single configuration property that works identically on ZooKeeper and KRaft. Define the domain once and Amazon MSK applies it across every broker, so custom domain names keep working as the cluster scales.
OAuth 2.0, LDAP, and HTTP auth for Amazon MQ for RabbitMQ
Amazon MQ for RabbitMQ supports OAuth 2.0, LDAP, and HTTP-based authentication backends so you can connect your broker to the identity infrastructure you already use. This post explains how each approach works and helps you decide which one fits your use case.
Mutual TLS and SSL certificate authentication for Amazon MQ for RabbitMQ
Learn how to add certificate-based identity verification to Amazon MQ for RabbitMQ. This post explains SSL certificate authentication for passwordless login through the EXTERNAL SASL mechanism and mutual TLS (mTLS) for two-way certificate verification, highlights the key rabbitmq.conf settings, and helps you decide which approach fits your compliance requirements.
Authentication and authorization options for Amazon MQ for RabbitMQ
Amazon MQ for RabbitMQ supports multiple authentication and authorization methods, so you can connect your broker to the identity infrastructure you already use. This post introduces the available options and helps you choose the right one for your use case.
Amazon OpenSearch Service extends version lifecycle support timelines
In November 2024, we announced Standard and Extended Support dates for legacy Elasticsearch and OpenSearch versions on Amazon OpenSearch Service. We are extending security and operating system patch coverage for these versions by 12 months, through November 7, 2027, and announcing support dates for additional Elasticsearch and OpenSearch versions.
Introducing Apache Spark troubleshooting agent for Amazon EMR on EKS
In this post, we show you how to set up the agent for Amazon EMR on EKS and walk through troubleshooting a failed job run. We demonstrate the workflow from both the Amazon EMR console and an AI assistant that supports the Model Context Protocol (MCP), an open standard for connecting AI assistants to external tools and data.









