AWS Big Data Blog
AWS and DuckLabs: Building the future of analytics together
Today we are announcing that Amazon has signed a definitive agreement to acquire DuckLabs, the Amsterdam-based company behind the open-source analytical database DuckDB. We expect the transaction to close shortly, subject to customary closing conditions. Hannes Mühleisen and Mark Raasveldt, who created DuckDB and co-founded DuckLabs, will continue leading the team and the open-source project’s technical direction as part of AWS. The DuckDB open-source project will also continue to be driven by the DuckLabs team, remain open source under the independent Foundation (the non-profit that oversees DuckDB), and available under the MIT license as it does today.
Building an LLM-powered DAG failure analysis plugin for Amazon MWAA
Debugging Apache Airflow DAG failures across services like AWS Glue, Amazon EMR, and Amazon Athena is slow and manual. In this post, we show you how to build a custom Airflow plugin that integrates with Amazon Bedrock to automatically analyze DAG task failures and deliver on-demand root cause analysis on Amazon MWAA.
Announcing Spark Connect on Amazon EMR on EKS: Interactive PySpark development, anywhere
Announcing Spark Connect on Amazon EMR on EKS: build, test, and debug Spark applications from VS Code, PyCharm, Jupyter notebooks, Amazon SageMaker Unified Studio, or dbt, while running full-scale Spark operations on your existing Amazon EKS clusters.
Aurora PostgreSQL zero-ETL integration with Amazon SageMaker
Amazon Aurora PostgreSQL zero-ETL integration with Amazon SageMaker replicates your operational data to a lakehouse in near real time, without building custom ETL pipelines. Learn the architecture and change data capture mechanics, then set up the integration and query your data in Amazon SageMaker.
Getting started with Apache Iceberg write support in Amazon Redshift – Part 3
Amazon Redshift now supports evolving Apache Iceberg table schemas and partition layouts through ALTER statements, with no data rewrites or pipeline rebuilds. In this final post of the series, you rename, add, drop, and widen columns, evolve partitions, and create AWS Lake Formation resource links for governed cross-engine access to Amazon S3 Tables.
Enforce IAM permissions boundaries for Amazon SageMaker Unified Studio Tooling blueprints
Amazon SageMaker Unified Studio now supports custom permissions boundaries for the IAM roles its Tooling blueprint creates. Learn how to create a permissions boundary that restricts AI agent capabilities, configure it on the Tooling blueprint with the AWS CLI, and validate that every provisioned role carries the boundary.
How Delivery Hero rebuilt real-time ad measurement with Apache Flink
Learn how Delivery Hero rebuilt its ad measurement pipeline from hourly batch processing to real time on Amazon Managed Service for Apache Flink, cutting the gap between an event and its recording from 61 minutes to 1.2 seconds and reducing monthly operational costs by about 57%.
Configure domain-level VPC networking in Amazon SageMaker Unified Studio
Configuring VPC networking per project across a SageMaker Unified Studio domain creates inconsistent, hard-to-audit networks. This post shows administrators how to configure domain-level VPC networking once, so every new project automatically inherits consistent, private network isolation, then update existing projects and validate connectivity.
Announcing Spark Connect on Amazon EMR on EC2: Interactive PySpark anywhere
Amazon EMR on EC2 now supports Spark Connect, so you can develop and debug PySpark interactively from Amazon SageMaker Unified Studio Data Notebooks or your own IDE while Spark runs on your cluster. This post shows you how to get started from both a Data Notebook and a local IDE.
Query unstructured data in Amazon SageMaker Catalog using generative AI
In Part 2 of this series, sign in as a data consumer, subscribe to enriched unstructured data assets in Amazon SageMaker Catalog, and query them using natural language through a no-code Amazon Bedrock chat agent and Amazon Bedrock model inference.
Best practices for scaling large consumer groups on Amazon MSK
As consumer groups on Amazon MSK scale to thousands of members, the metadata record Kafka writes during rebalances can exceed the 1 MB limit and stall the group. Learn how to estimate metadata size, raise the topic-level limit safely, plan capacity, and apply complementary strategies for scaling large consumer groups.










