AWS Big Data Blog
Category: Technical How-to
Monitoring MWAA-orchestrated ETL pipelines with Amazon OpenSearch Service
Troubleshooting a multi-service, MWAA-orchestrated ETL pipeline often means hunting across scattered Amazon CloudWatch log groups. This post shows how to centralize your ETL logs in Amazon OpenSearch Service and query them in plain language through an MCP server on Amazon Bedrock AgentCore, so you find root causes faster.
Cost-effective ETL with DuckDB and Amazon S3 Tables on AWS Glue
Learn how to pair DuckDB with AWS Glue 6.0 to run SQL-centric ETL on a single worker, reading Parquet from Amazon S3 and writing Apache Iceberg tables to Amazon S3 Tables. This post walks through a complete, runnable example and compares measured cost and runtime against an equivalent Apache Spark job on the same Glue runtime.
Run an automated operational review with the Amazon Redshift MCP server
Amazon Redshift automates much of its own tuning, but periodic operational reviews still pay off. Learn how the review_cluster tool in the open-source Amazon Redshift MCP server runs a full cluster diagnostic from one natural-language request, evaluating 12 diagnostic areas and returning prioritized, documentation-linked recommendations in minutes.
Optimize consumer rebalancing on Amazon MSK with next generation protocol
The KIP-848 consumer protocol in Apache Kafka 4.0 redesigns consumer group rebalancing to eliminate stop-the-world pauses. This post explains how the consumer protocol works on Amazon MSK, how to enable it on MSK Standard and Express brokers, and how to diagnose and resolve slow rebalancing issues.
Building an LLM-powered DAG failure analysis plugin for Amazon MWAA
Debugging Apache Airflow DAG failures across services like AWS Glue, Amazon EMR, and Amazon Athena is slow and manual. In this post, we show you how to build a custom Airflow plugin that integrates with Amazon Bedrock to automatically analyze DAG task failures and deliver on-demand root cause analysis on Amazon MWAA.
Aurora PostgreSQL zero-ETL integration with Amazon SageMaker
Amazon Aurora PostgreSQL zero-ETL integration with Amazon SageMaker replicates your operational data to a lakehouse in near real time, without building custom ETL pipelines. Learn the architecture and change data capture mechanics, then set up the integration and query your data in Amazon SageMaker.
Getting started with Apache Iceberg write support in Amazon Redshift – Part 3
Amazon Redshift now supports evolving Apache Iceberg table schemas and partition layouts through ALTER statements, with no data rewrites or pipeline rebuilds. In this final post of the series, you rename, add, drop, and widen columns, evolve partitions, and create AWS Lake Formation resource links for governed cross-engine access to Amazon S3 Tables.
Configure domain-level VPC networking in Amazon SageMaker Unified Studio
Configuring VPC networking per project across a SageMaker Unified Studio domain creates inconsistent, hard-to-audit networks. This post shows administrators how to configure domain-level VPC networking once, so every new project automatically inherits consistent, private network isolation, then update existing projects and validate connectivity.
Query unstructured data in Amazon SageMaker Catalog using generative AI
In Part 2 of this series, sign in as a data consumer, subscribe to enriched unstructured data assets in Amazon SageMaker Catalog, and query them using natural language through a no-code Amazon Bedrock chat agent and Amazon Bedrock model inference.
Discover and govern Snowflake data using SageMaker Unified Studio
Connect Snowflake to Amazon SageMaker Unified Studio to build a unified data catalog. Query federated Snowflake tables without moving data, publish enriched assets to SageMaker Catalog, and validate data quality with AWS Glue Data Quality, all while keeping data in Snowflake.









