AWS Big Data Blog
Category: Technical How-to
Building an LLM-powered DAG failure analysis plugin for Amazon MWAA
Debugging Apache Airflow DAG failures across services like AWS Glue, Amazon EMR, and Amazon Athena is slow and manual. In this post, we show you how to build a custom Airflow plugin that integrates with Amazon Bedrock to automatically analyze DAG task failures and deliver on-demand root cause analysis on Amazon MWAA.
Aurora PostgreSQL zero-ETL integration with Amazon SageMaker
Amazon Aurora PostgreSQL zero-ETL integration with Amazon SageMaker replicates your operational data to a lakehouse in near real time, without building custom ETL pipelines. Learn the architecture and change data capture mechanics, then set up the integration and query your data in Amazon SageMaker.
Getting started with Apache Iceberg write support in Amazon Redshift – Part 3
Amazon Redshift now supports evolving Apache Iceberg table schemas and partition layouts through ALTER statements, with no data rewrites or pipeline rebuilds. In this final post of the series, you rename, add, drop, and widen columns, evolve partitions, and create AWS Lake Formation resource links for governed cross-engine access to Amazon S3 Tables.
Configure domain-level VPC networking in Amazon SageMaker Unified Studio
Configuring VPC networking per project across a SageMaker Unified Studio domain creates inconsistent, hard-to-audit networks. This post shows administrators how to configure domain-level VPC networking once, so every new project automatically inherits consistent, private network isolation, then update existing projects and validate connectivity.
Query unstructured data in Amazon SageMaker Catalog using generative AI
In Part 2 of this series, sign in as a data consumer, subscribe to enriched unstructured data assets in Amazon SageMaker Catalog, and query them using natural language through a no-code Amazon Bedrock chat agent and Amazon Bedrock model inference.
Discover and govern Snowflake data using SageMaker Unified Studio
Connect Snowflake to Amazon SageMaker Unified Studio to build a unified data catalog. Query federated Snowflake tables without moving data, publish enriched assets to SageMaker Catalog, and validate data quality with AWS Glue Data Quality, all while keeping data in Snowflake.
How to migrate from Amazon CloudSearch to Amazon OpenSearch Serverless
Learn how to migrate an Amazon CloudSearch domain to Amazon OpenSearch Serverless: assess your configuration, create a collection with explicit index mappings, convert your documents and queries to the OpenSearch query DSL, configure security, load data with Amazon OpenSearch Ingestion, and validate before cutover.
Accelerating Spark queries with Iceberg materialized views
Accelerate slow, repetitive Apache Spark analytical queries on Apache Iceberg tables without rewriting any SQL. This post shows how automatic query rewrite in Amazon EMR and AWS Glue uses Iceberg materialized views in the AWS Glue Data Catalog to transparently substitute matching query plans, and how to design materialized views for the best speedup.
Build declarative ETL pipelines with AWS Glue 6.0
AWS Glue 6.0 introduces Spark Declarative Pipelines. In this post, you build a single declarative AWS Glue 6.0 job that turns raw order records into validated, aggregated tables through a bronze, silver, and gold sequence, without writing any orchestration logic.
From silos to insights: Federated data access patterns for AI agents
AI agents can reach enterprise data where it lives instead of routing every question through data engineers. This post presents three reference patterns for federated data access using Model Context Protocol (MCP) servers and Amazon Bedrock AgentCore: catalog-first, direct source, and hybrid access.









