AWS Big Data Blog

Category: Advanced (300)

Building an LLM-powered DAG failure analysis plugin for Amazon MWAA

Building an LLM-powered DAG failure analysis plugin for Amazon MWAA

Debugging Apache Airflow DAG failures across services like AWS Glue, Amazon EMR, and Amazon Athena is slow and manual. In this post, we show you how to build a custom Airflow plugin that integrates with Amazon Bedrock to automatically analyze DAG task failures and deliver on-demand root cause analysis on Amazon MWAA.

Aurora PostgreSQL zero-ETL integration with Amazon SageMaker

Aurora PostgreSQL zero-ETL integration with Amazon SageMaker

Amazon Aurora PostgreSQL zero-ETL integration with Amazon SageMaker replicates your operational data to a lakehouse in near real time, without building custom ETL pipelines. Learn the architecture and change data capture mechanics, then set up the integration and query your data in Amazon SageMaker.

Getting started with Apache Iceberg write support in Amazon Redshift – Part 3

Getting started with Apache Iceberg write support in Amazon Redshift – Part 3

Amazon Redshift now supports evolving Apache Iceberg table schemas and partition layouts through ALTER statements, with no data rewrites or pipeline rebuilds. In this final post of the series, you rename, add, drop, and widen columns, evolve partitions, and create AWS Lake Formation resource links for governed cross-engine access to Amazon S3 Tables.

Enforce IAM permissions boundaries for Amazon SageMaker Unified Studio Tooling blueprints

Enforce IAM permissions boundaries for Amazon SageMaker Unified Studio Tooling blueprints

Amazon SageMaker Unified Studio now supports custom permissions boundaries for the IAM roles its Tooling blueprint creates. Learn how to create a permissions boundary that restricts AI agent capabilities, configure it on the Tooling blueprint with the AWS CLI, and validate that every provisioned role carries the boundary.

Configure domain-level VPC networking in Amazon SageMaker Unified Studio

Configure domain-level VPC networking in Amazon SageMaker Unified Studio

Configuring VPC networking per project across a SageMaker Unified Studio domain creates inconsistent, hard-to-audit networks. This post shows administrators how to configure domain-level VPC networking once, so every new project automatically inherits consistent, private network isolation, then update existing projects and validate connectivity.

Announcing Spark Connect on Amazon EMR on EC2: Interactive PySpark anywhere

Announcing Spark Connect on Amazon EMR on EC2: Interactive PySpark anywhere

Amazon EMR on EC2 now supports Spark Connect, so you can develop and debug PySpark interactively from Amazon SageMaker Unified Studio Data Notebooks or your own IDE while Spark runs on your cluster. This post shows you how to get started from both a Data Notebook and a local IDE.

Scale down Kinesis Data Streams on-demand capacity with ODA warm throughput

Scale down Kinesis Data Streams on-demand capacity with ODA warm throughput

Amazon Kinesis Data Streams now supports scaling down ingest capacity for on-demand Advantage streams with warm throughput. Learn how the scale-down works, how to monitor stream behavior with Amazon CloudWatch, and best practices for releasing excess capacity after transient traffic bursts.