AWS Big Data Blog
Category: Advanced (300)
Centralized CloudTrail monitoring across 100+ AWS accounts
Learn how to build a centralized AWS CloudTrail monitoring solution on Amazon OpenSearch Service, with Terraform managing the full stack. It handles 200 GB/day of logs across 100+ accounts, provides automated threat detection, delivers on-demand SOC 2, PCI DSS, and HIPAA compliance reporting, and gives four teams isolated access.
Scaling fine-grained access control for enterprise lakehouse using SageMaker Unified Studio and AWS Lake Formation
As enterprise lakehouses grow to thousands of tables across business domains and regions, fine-grained access control becomes a governance bottleneck. This post shows how to combine AWS IAM Identity Center, AWS Lake Formation tag-based access control, and trusted identity propagation in Amazon SageMaker Unified Studio for automated, auditable, least-privilege access.
Event-driven pipeline orchestration with Amazon MWAA and Airflow 3.0
Data engineering teams running Apache Airflow across multiple AWS accounts have no built-in way to coordinate workflows between separate Amazon MWAA environments. With Airflow 3.0 on Amazon MWAA, you can use asset-based scheduling and Asset Watchers with Amazon SQS to build event-driven, cross-account orchestration that replaces polling with near real-time triggers.
Amazon Redshift multi-Region disaster recovery
In this post, we walk through the core concepts of cross-Region disaster recovery, introduce a framework for assessing your requirements, and then dive deep into three primary DR strategies for Amazon Redshift: Active-Passive, Active-Active, and a Hybrid approach. For each strategy, we cover architecture, trade-offs, implementation guidance, and cost considerations so you can make an informed decision for your workload.
Upgrade Amazon Redshift DC2 clusters to the new Amazon Redshift RG
Upgrading your Amazon Redshift DC2 clusters to AWS Graviton-based RG instances unlocks managed storage, data sharing, zero-ETL, and an integrated data lake engine. This post covers the new features, node-mapping guidance for sizing, the available upgrade methods, and how to validate your target configuration with Amazon Redshift Test Drive.
Optimizing costs and performance with Advanced Managed Scaling on Amazon EMR on EC2
In this post, we discuss the benefits of Advanced Scaling for Amazon EMR on Amazon EC2 and demonstrate how it works through some example scenarios. You’ll learn when to prioritize utilization optimized settings for cost savings with conservative scaling, balanced approaches for mixed workloads, or performance optimized configurations for SLA-sensitive jobs requiring aggressive scaling.
Lowering AWS KMS decrypt API costs in EMR Spark jobs
Processing encrypted data in Amazon S3 with Amazon EMR and Apache Spark can drive up AWS KMS decrypt API costs as the number of objects grows. This post shows three techniques to reduce those costs without compromising encryption: optimizing file formats (including Apache Iceberg), aggregating data, and using AWS Glue Data Catalog partition indexes.
Automate Spark Scala migration to 4.x with AWS Spark Upgrade Agent
Learn how to automate Apache Spark 3.x to 4.0 Scala migration on Amazon EMR using the AWS Spark Upgrade Agent. This post covers API deprecations, behavioral changes, build configuration updates, and job validation, turning months of manual effort into hours.
Building a scalable personalized recommendation system on AWS: From batch to real-time
Learn how the Everyday Essentials team built a scalable personalized recommendation platform on AWS using a batch-first architecture with Amazon MWAA for orchestration, Amazon SageMaker for training and vector search, and AWS Lake Formation for governed data access, then extended it to real-time with Amazon MemoryDB.
Automate creating AWS Glue Data Catalog views with AWS SDK for data mesh use case
This post shows you how to use the Catalog objects API CreateTable() to programmatically create ATHENA and SPARK dialects using cross-account IAM definer roles, and how to add the ATHENA dialect programmatically for the views that were created earlier with only SPARK dialect.









