AWS Big Data Blog
Category: Technical How-to
Streamline Apache Kafka cluster operations and migrations with Agent Skills for Amazon MSK
Agent Skills for Amazon MSK bring broker-type-aware expertise to operating and migrating Apache Kafka clusters. In this post, we walk through installing the managing-amazon-msk and migrate-to-msk skills and demonstrate how they diagnose performance issues, size clusters with cost breakdowns, and plan migrations from self-managed Kafka to Amazon MSK.
Upgrade Amazon Redshift DC2 clusters to the new Amazon Redshift RG
Upgrading your Amazon Redshift DC2 clusters to AWS Graviton-based RG instances unlocks managed storage, data sharing, zero-ETL, and an integrated data lake engine. This post covers the new features, node-mapping guidance for sizing, the available upgrade methods, and how to validate your target configuration with Amazon Redshift Test Drive.
Optimizing costs and performance with Advanced Managed Scaling on Amazon EMR on EC2
In this post, we discuss the benefits of Advanced Scaling for Amazon EMR on Amazon EC2 and demonstrate how it works through some example scenarios. You’ll learn when to prioritize utilization optimized settings for cost savings with conservative scaling, balanced approaches for mixed workloads, or performance optimized configurations for SLA-sensitive jobs requiring aggressive scaling.
Lowering AWS KMS decrypt API costs in EMR Spark jobs
Processing encrypted data in Amazon S3 with Amazon EMR and Apache Spark can drive up AWS KMS decrypt API costs as the number of objects grows. This post shows three techniques to reduce those costs without compromising encryption: optimizing file formats (including Apache Iceberg), aggregating data, and using AWS Glue Data Catalog partition indexes.
Automate Spark Scala migration to 4.x with AWS Spark Upgrade Agent
Learn how to automate Apache Spark 3.x to 4.0 Scala migration on Amazon EMR using the AWS Spark Upgrade Agent. This post covers API deprecations, behavioral changes, build configuration updates, and job validation, turning months of manual effort into hours.
Build a contract compliance search system with Amazon OpenSearch
In this post, you build a contract compliance search system that combines semantic search with semantic highlighting in Amazon OpenSearch Service. You deploy the solution using two AWS CloudFormation stacks, test it with synthetic contract documents, and see how a single query surfaces both the right contracts and the right clauses within them.
Automate creating AWS Glue Data Catalog views with AWS SDK for data mesh use case
This post shows you how to use the Catalog objects API CreateTable() to programmatically create ATHENA and SPARK dialects using cross-account IAM definer roles, and how to add the ATHENA dialect programmatically for the views that were created earlier with only SPARK dialect.
Efficient log management with Amazon OpenSearch Service data streams
In this post, we show you how to implement data streams with Index State Management (ISM) in Amazon OpenSearch Service. This approach automatically manages your time series data lifecycle and optimizes both performance and costs. Data streams distribute incoming data across multiple backing indices, helping to reduce single-index bottlenecks, while ISM policies automate rollover, retention, and storage tiering to help manage costs.
Govern Amazon Redshift Data Warehouses Data Across Accounts using Amazon SageMaker Unified Studio
In this post, we show you how to use Amazon SageMaker Unified Studio to implement cross-account data sharing in Amazon Redshift using data mesh principles. We demonstrate how to build a scalable data mesh architecture that supports secure, auditable data sharing across AWS accounts while reducing operational burden.
Patch perfect: Automating Amazon Redshift patch testing
In this post, we demonstrate an automated test suite that validates your Amazon Redshift cluster automatically after any patch, reboot, or modification. It uses standard drivers against real workload patterns to provide a verified gate between a patch landing and that patch reaching production.









