AWS Cloud Operations Blog
Share Systems Manager documents across your AWS Organization with AWS RAM
Introduction If you operate AWS Systems Manager (SSM) documents, such as Automation runbooks and Run Command documents, across an AWS Organization, you must decide which accounts can run them. Until now, that meant sharing each document with a list of individual account IDs through the ModifyDocumentPermission API, which accepts up to 1,000 account IDs per […]
Best practices for writing AWS DevOps Agent Skills
When an incident hits at 2 AM, the on-call engineer’s effectiveness depends on what they know about the system, which metrics to check first, what “normal” looks like for this service, and where to find the deployment history. That knowledge often lives in runbooks, internal wikis, and the heads of senior engineers who built the […]
Automate RCA across ServiceNow, Dynatrace and Slack with AWS DevOps Agent
If you manage production incidents, you know the drill. A ServiceNow ticket fires at 2 AM. The on-call engineer wakes up, logs in to Dynatrace, pulls traces and metrics across multiple dashboards, cross-references change records, forms a hypothesis, updates the ticket, and posts findings to Slack. The investigation takes one to three hours, and that’s […]
From 2 AM alarm to answer: Security Triage with AWS DevOps Agent
Security incident triage usually starts cold. An Amazon CloudWatch alarm fires at 2 a.m. because an AWS Identity and Access Management (IAM) access key just created another access key. The responder spends the first 30 minutes assembling evidence from AWS CloudTrail, Amazon CloudWatch Logs, and Amazon Virtual Private Cloud (Amazon VPC) Flow Logs before deciding […]
Accelerate troubleshooting with AWS Observability as a Kiro power
Troubleshooting a distributed application means correlating signals across alarms, traces, logs, and deployments, usually across several consoles while the clock is running. Imagine your on-call engineer gets paged at 2 AM. P99 latency on the checkout API has spiked past the SLO threshold. What follows is a familiar scramble: open Application Signals to check service […]
This Month in AWS Observability: August – September 2026
Introduction August and September brought the general availability of Amazon CloudWatch Omni, an AI-first, app-centric observability experience that brings together telemetry across AWS accounts, Regions, and Azure workloads in a single space. Alarms gained warm-up periods and wall clock evaluation windows, cutting the noise that comes from startup gaps and rolling-window edge cases. Database observability […]
Introducing Amazon CloudWatch Omni: Observability for the AI Era
An AI-powered observability experience for AI agents and applications, built on open standards and delivered outside of the AWS Console. Organizations are handing agents the keys to their day-to-day operations, from resolving support tickets and managing infrastructure to approving expenses, shipping production code, and increasingly the long tail of workflows that keep the business running. […]
Root Cause Analysis with Amazon Managed Service for Prometheus and AWS DevOps Agent
Introduction Teams running Prometheus-instrumented workloads, whether on Kubernetes, Amazon EC2, containers, or on-premises servers, face a growing challenge: alert fatigue from static threshold monitoring and hours spent manually investigating performance degradations. These teams can significantly reduce time spent investigating false positive alerts and manually correlating metrics by implementing automated root cause analysis. When real issues […]
Investigate your AWS account activity in plain language with Amazon Q
AWS CloudTrail now integrates with Amazon Q in the AWS Management Console, letting you investigate your AWS account activity using plain language. CloudTrail records API activity across your AWS account for security auditing, compliance, and operational troubleshooting. Getting insights from this data has traditionally meant writing queries in Amazon Athena or Amazon CloudWatch Logs Insights, […]
Reduce MTTR with AI-driven RCA using AWS DevOps Agent and Splunk
Modern cloud-native applications generate rich telemetry across metrics, logs, and deployment histories. When performance degrades, operations teams have the data, but the challenge is correlating signals across multiple tools quickly enough to minimize customer impact. Root cause analysis remains a largely manual process dependent on institutional knowledge and operator experience. This post shows how AWS […]








