AWS Big Data Blog

Category: Technical How-to

Enable cross-cloud analytics with Amazon S3 Tables and Google BigQuery, Part 1: IAM-based access control

Enable cross-cloud analytics with Amazon S3 Tables and Google BigQuery, Part 1: IAM-based access control

Your Google BigQuery users need to query data that lives in Amazon S3 Tables on AWS without copying it across clouds. This post shows how to connect BigQuery to Amazon S3 Tables through the AWS Glue Iceberg REST Catalog using IAM-based access control, so you keep one governed dataset and query it live from BigQuery.

Enable cross-cloud analytics with Amazon S3 Tables and Google BigQuery, Part 2: access control with Lake Formation

Enable cross-cloud analytics with Amazon S3 Tables and Google BigQuery, Part 2: access control with Lake Formation

In Part 2 of this series, connect Google BigQuery to Amazon S3 Tables using AWS Lake Formation credential vending. Lake Formation manages fine-grained permissions and issues short-lived, scoped credentials to external engines, so you can centrally govern which teams and query engines read your Iceberg tables on AWS without managing IAM policies for every consumer.

Track SageMaker Unified Studio project costs with custom tags and AWS CUR

Track SageMaker Unified Studio project costs with custom tags and AWS CUR

Learn how to track Amazon SageMaker Unified Studio project costs by custom tags. This serverless solution enriches AWS Cost and Usage Report (CUR) data with custom project tags and visualizes cost by CostCenter, Team, or Environment in an Amazon Quick Sight dashboard.

Secure SageMaker Unified Studio access with SAML and conditional policies

Secure SageMaker Unified Studio access with SAML and conditional policies

Learn how to secure Amazon SageMaker Unified Studio by integrating it with an external SAML identity provider such as Okta. This post shows you how to apply conditional access policies that enforce device compliance, IP-based restrictions, and multi-factor authentication for your data and AI workloads.

IAM authentication with OAuth 2.0 for Amazon MQ for RabbitMQ

IAM authentication with OAuth 2.0 lets clients connect to Amazon MQ for RabbitMQ using their existing IAM identity instead of static broker-local credentials. This post covers the key rabbitmq.conf configuration for using AWS IAM as an OAuth 2.0 provider and shows a multi-tenant example with vhost isolation enforced by IAM roles and broker scope aliases.

Centralized CloudTrail monitoring across 100+ AWS accounts

Centralized CloudTrail monitoring across 100+ AWS accounts

Learn how to build a centralized AWS CloudTrail monitoring solution on Amazon OpenSearch Service, with Terraform managing the full stack. It handles 200 GB/day of logs across 100+ accounts, provides automated threat detection, delivers on-demand SOC 2, PCI DSS, and HIPAA compliance reporting, and gives four teams isolated access.

AI-powered cost optimization agent for Amazon Kinesis Data Streams

AI-powered cost optimization agent for Amazon Kinesis Data Streams

Learn how to deploy an open-source, AI-powered agent built on Amazon Bedrock that automatically analyzes every Amazon Kinesis Data Streams stream in your account, compares costs across the three capacity modes, and recommends the optimal mode to help you save over 60% on streaming costs on a schedule you choose.

Streamline Apache Kafka cluster operations and migrations with Agent Skills for Amazon MSK

Streamline Apache Kafka cluster operations and migrations with Agent Skills for Amazon MSK

Agent Skills for Amazon MSK bring broker-type-aware expertise to operating and migrating Apache Kafka clusters. In this post, we walk through installing the managing-amazon-msk and migrate-to-msk skills and demonstrate how they diagnose performance issues, size clusters with cost breakdowns, and plan migrations from self-managed Kafka to Amazon MSK.

Upgrade Amazon Redshift DC2 clusters to the new Amazon Redshift RG

Upgrade Amazon Redshift DC2 clusters to the new Amazon Redshift RG

Upgrading your Amazon Redshift DC2 clusters to AWS Graviton-based RG instances unlocks managed storage, data sharing, zero-ETL, and an integrated data lake engine. This post covers the new features, node-mapping guidance for sizing, the available upgrade methods, and how to validate your target configuration with Amazon Redshift Test Drive.

Optimizing costs and performance with Advanced Managed Scaling on Amazon EMR on EC2

In this post, we discuss the benefits of Advanced Scaling for Amazon EMR on Amazon EC2 and demonstrate how it works through some example scenarios. You’ll learn when to prioritize utilization optimized settings for cost savings with conservative scaling, balanced approaches for mixed workloads, or performance optimized configurations for SLA-sensitive jobs requiring aggressive scaling.