A recent outage in the us-east-1 region of Amazon Web Services (AWS) caused significant complications for many global services, including communication platforms and thousands of web applications. This incident, which lasted a full 14 hours, deserves a deeper analysis to understand its roots and lessons learned.
The primary cause of this outage was a DNS service failure for the DynamoDB database. DynamoDB is a NoSQL database designed for high availability and resilience, with a guaranteed 99.99% availability (SLA) when configured with replication across multiple Availability Zones (AZs). It is a popular choice for many applications and a key component for numerous internal AWS services themselves. During the incident, however, the address dynamodb.us-east-1.amazonaws.com began returning an empty DNS record, creating the impression that DynamoDB had completely disappeared from the region.
To understand the failure mechanism, it is essential to look at how DNS is managed for DynamoDB. The system uses two key components: the DNS Planner and the DNS Enactor. The DNS Planner monitors the state of load balancers and creates plans for load distribution. The DNS Enactor is then responsible for updating these routes in the internal Route 53 DNS service. To ensure resilience, an instance of the DNS Enactor runs in each Availability Zone, which in the us-east-1 region meant three instances running in parallel.
Although race conditions are expected in such a parallel environment, and the system is designed to handle them using the principle of eventual consistency, a combination of several unfortunate events led to the failure. One of the DNS Enactor instances (Enactor #1) began processing updates unusually slowly. Simultaneously, the DNS Planner started generating new plans much faster, and another Enactor instance (Enactor #2) processed these new plans at high speed. Once Enactor #2 finished writing the latest plans to Route 53, it returned to the DNS Planner and deleted the old plans.
This sequence of events plunged the system into an inconsistent state. The slowly working Enactor #1 unknowingly used an old plan that had already been deleted by Enactor #2. When Enactor #2 detected the use of the old plan, its cleanup mechanism was triggered, which in this case cleared all IP addresses for regional endpoints in Route 53, effectively nullifying the DNS record for DynamoDB in us-east-1. This state caused all services attempting to connect to DynamoDB in that region to experience DNS failures.
The DynamoDB outage subsequently manifested in Amazon EC2 services as well. The DropletWorkflow Manager (DWFM) subsystem, which manages physical servers for EC2 instances (called "droplets"), records the "leasing" of each server to determine if it is occupied or can be allocated to a customer. The results of server health checks are stored in DynamoDB. Due to the unavailability of DynamoDB, droplet leases began to expire, and DWFM started returning "insufficient capacity" errors, as it believed the servers were unavailable.
Even after DynamoDB was restored, the situation did not immediately improve. Attempts to re-establish leases on a vast number of droplets took so long that the work could not be completed before further timeouts expired. DWFM entered a state of "congestive collapse," unable to effectively progress in lease restoration. AWS engineers had to spend additional hours developing and applying mitigation measures to get EC2 instance allocation working again.

Haben Sie eine Idee für eine Website oder App?
Lassen Sie uns unverbindlich darüber sprechen – ohne Druck, mit einem konkreten nächsten Schritt.
However, the problems were not limited to EC2. Errors in network propagation followed. Even though EC2 instances appeared healthy internally, they could not communicate with the outside world due to congestion in the Network Manager system, which processes network state changes. This complication lasted another five hours before the load on the Network Manager was reduced and network connectivity restoration accelerated.
The overall system recovery was completed after 14 hours, when request throttles, which had been introduced to reduce the load on individual EC2 subsystems, were gradually removed. Throughout this challenging period, the AWS engineering team worked intensively to resolve complications that had a widespread impact on other services as well, such as Network Load Balancer (NLB), Lambda functions, ECS, EKS, and many others.
This incident highlighted the complexity and interconnectedness of modern cloud infrastructures. It showed that even systems designed for high availability can fail due to an unexpected combination of factors within their distributed components. It is crucial to continuously review and improve resilience mechanisms, especially in critical areas such as DNS and resource management.
