Amazon AWS Outage: What Happened, Why It Keeps Happening and How to Protect Your Business

Table of Contents

amazon aws outage

Amazon AWS Outage: What Happened, Why It Keeps Happening and How to Protect Your Business

When a single cloud region stumbles, thousands of websites, apps and payment flows stumble with it. The Amazon AWS outage that struck the US-EAST-1 region in October 2025 was a vivid reminder of that fact, and a second US-EAST-1 failure in May 2026 proved the lesson was not a one-off. If your business runs a website, a store, a booking flow or a customer portal, an AWS outage is a risk inside your own stack, whether you host on AWS directly or simply depend on a vendor that does.

This guide explains what happened during the major AWS outage events, what actually caused them, who felt the impact and what the AWS downtime record tells us about planning. It also lays out practical steps any small or mid-sized business can take to reduce the damage from the next cloud outage. Every date, duration and figure comes from published incident reports, monitoring firms and reputable news coverage, and where a number is an estimate we say so.

What Is an Amazon AWS Outage and Why Does It Matter?

Amazon Web Services is the largest cloud provider in the world. It rents out computing power, storage, databases and networking to companies of every size. An AWS outage happens when one or more of those services stops working correctly in a region. Because so many businesses build on the same foundation, a fault in one place can look like the whole internet going down.

AWS is organized into regions, and each region contains several availability zones, which are separate groups of data centers. The design promise is that if one zone fails, the others keep working. In practice, the outages that make headlines usually happen when a shared control layer breaks or when the failure is bigger than the redundancy was built to absorb.

Why US-EAST-1 keeps making the news

US-EAST-1 in Northern Virginia is one of the oldest and most heavily used AWS regions. Engineers often call it the problem child of the network, not because it is careless, but because so much traffic and so many internal dependencies concentrate there. When something goes wrong in US-EAST-1, the blast radius is larger than almost anywhere else. A fair question for your own business is how much of your stack lives in, or depends on, a single region.

The October 2025 AWS Outage: Timeline and Root Cause

The most widely covered incident of recent years began in the early hours of October 20, 2025 (Pacific time). One network monitoring analysis put the disruption in US-EAST-1 at more than 15 hours for some customers, and tracking sites recorded more than 50,000 outage reports at the peak as consumer apps, games, banking tools and workplace software failed together.

StageWhat happened
Start of disruptionErrors began in US-EAST-1 in the early hours of October 20, 2025 (Pacific time).
DiagnosisThe problem was tied to DNS resolution for the regional DynamoDB endpoint.
First fixThe DNS issue was corrected and DynamoDB connectivity began to recover.
Long tailEC2 launches, Network Load Balancer health checks and dependent services needed hours more to stabilize.
Post-event reportAWS published a detailed summary within days, much faster than after the large 2023 event.

The root cause: a DNS race condition in DynamoDB

AWS’s post-event summary said the root cause was a latent race condition in the DynamoDB DNS management system. It produced an incorrect, empty DNS record for the regional endpoint dynamodb.us-east-1.amazonaws.com, and the automation meant to repair it failed to do so. Once DNS stopped returning valid addresses, applications could no longer find DynamoDB in that region.

Engineers who reviewed the report noted that DNS was the first symptom rather than the deepest cause. The deeper problem was a subtle bug in an internal automation service. The failure then cascaded. New EC2 instances were created, but their network configuration never finished. That delayed network state propagation, caused false health check failures and took down Network Load Balancer nodes, producing connection errors across dependent services.

This is why a database fault looked like a total internet failure. DynamoDB is a serverless NoSQL database that many AWS services rely on internally, so its loss spread sideways into services customers never thought of as related. AWS said it disabled the DynamoDB DNS automation worldwide while it fixed the race condition, and it committed to velocity controls for Network Load Balancers and new recovery testing for EC2.

Who Was Affected by the AWS Outage?

The October 2025 event hit consumer and business software alike. Coverage named Snapchat, Roblox, Fortnite, Venmo, Robinhood, Reddit, Slack and Atlassian products, and Amazon’s own Alexa and Ring services were disrupted as well. Universities and public bodies posted their own incident notices.

The financial cost is hard to pin down because cloud providers do not publish customer impact figures. An insurance analytics firm published preliminary insured loss estimates between 38 million and 581 million dollars and estimated roughly 70,000 organizations were affected. Larger figures floated by commentators should be read as rough opinion rather than measured fact.

The hidden victims: small businesses on shared platforms

A local service company might not use AWS directly, yet its booking widget, payment processor, email tool, CRM or hosting provider does. When those tools fail together, the owner sees lost leads and failed checkouts with no obvious explanation. This is the same dependency chain we describe when explaining how a website, CRM and SEO effort should work together in a connected digital ecosystem. Every extra dependency is another place an outage can reach you.

Why people search for reviews and complaints during an outage

There is a consumer side effect most outage write-ups skip. When an app refuses to log in or a subscription page throws an error, people go looking for answers, and searches such as finelo reviews and finelo reviews and complaints are a good example. Finelo is an investing education app. Its Trustpilot profile shows a very high volume of reviews with a strong overall rating, while some App Store and Better Business Bureau accounts describe billing and refund frustrations. Nothing in that record links Finelo to AWS. The point is behavioral: a login or payment failure caused by infrastructure can be misread as a company problem, so users turn to review sites. If your business depends on subscriptions or logins, an outage can become a reputation event even when the cause was upstream.

The May 2026 AWS Outage: A Thermal Event in US-EAST-1

If October 2025 was a software failure, May 2026 was a physical one. On May 7, 2026, a data center in the use1-az4 availability zone overheated after cooling capacity dropped. Power loss followed, and EC2 instances and EBS volumes on affected racks went down. According to AWS builder content, the impairment lasted about 28 hours while engineers restored cooling and brought hardware back safely.

The impact was concentrated but severe. Reporting named Coinbase, FanDuel and CME Group, with Coinbase offline for roughly seven hours. One database provider noted that multi-AZ high availability did not protect Coinbase, because its latency-sensitive matching engine ran inside a single zone by design. Redundancy only helps if it covers the component that fails.

ComparisonOctober 2025May 2026
RegionUS-EAST-1US-EAST-1 (zone use1-az4)
Nature of failureSoftware: DNS race conditionPhysical: cooling and power loss
ScopeRegion-wide service disruptionSingle availability zone
Approximate durationAround 15 hours for many customersAbout 28 hours of impairment
Notable services hitSnapchat, Roblox, Slack, Atlassian, Venmo, RedditCoinbase, FanDuel, CME Group

Two other events belong in the record. In March 2026, AWS facilities in the United Arab Emirates suffered a physical event involving fires and emergency power shutdowns, and the disruption spread to Bahrain. In September 2026, AWS said damage in the Bahrain region exceeded what its multi-AZ services were designed to withstand and that it could not restore some resources hosted only there. Earlier, a 2017 S3 outage ran about four hours and a December 2021 US-EAST-1 failure lasted more than five hours. Each had a different trigger, and each exposed the same weakness: heavy dependence on a few regions.

How to Check If AWS Is Down Right Now

When something breaks, first work out whether the fault is yours, your vendor’s or the cloud provider’s. Use this short checklist before changing any settings.

  1. Open the official AWS Health Dashboard and look for active events in the regions you use.
  2. Check an independent tracker such as Downdetector to see whether other sites report the same problem.
  3. Test your own site from a second network, browser and device to rule out a local issue.
  4. Check the status pages of your payment processor, email platform, CRM and host, because they may run on AWS even if you do not.

If the official dashboard shows nothing but users still report failures, the problem may be local, so avoid assuming a cloud outage without evidence.

What an AWS Outage Really Costs a Business

The cost of downtime is more than the hours your site is offline. Direct losses include missed sales and abandoned checkouts. Indirect losses include support tickets, refunds, staff time and lost trust. AWS service credits typically return only a small percentage of the affected service’s monthly bill, and commentary on the 2026 events notes they do not cover lost revenue or customer confidence.

To size your exposure, estimate the revenue that passes through your website per hour, multiply it by a realistic outage length and add the cost of manual recovery. Even a rough figure changes the question from whether resilience is worth paying for to how much you can afford. Slow sites feel this even outside outages, which is why the steps in our post on how website speed affects client acquisition belong in the same budget conversation.

How to Protect Your Business from the Next AWS Outage

You cannot prevent an AWS outage, but you can control how much of one reaches your customers. Start with an honest answer to how much downtime you can tolerate, then build up from there.

1. Back up everything, and store copies outside one region

The Bahrain statement is the clearest argument for off-region backups. If your only copy of critical data sits beside your live system, a large enough event can remove both. Automated, tested backups stored in a second region turn a catastrophe into an inconvenience. A managed approach such as website backup and security services covers the routine work, including restore testing, which most teams skip.

2. Plan for multi-region recovery where the stakes justify it

Multi-AZ setups protect against the loss of one building, but the 2025 and 2026 events showed their limits. For revenue-critical systems such as checkout, login and lead capture, a multi-region or warm standby design lets you shift traffic when a region degrades. It costs more, so reserve it for the parts that would hurt most.

3. Reduce hidden dependencies

List every third-party tool your website and sales flow rely on, find out where each is hosted and decide what happens when it fails. A form that silently drops leads when a CRM is unreachable is a bigger risk than most owners realize. Well-planned custom CRM automation can include retry logic, so a brief outage delays a lead instead of losing it.

4. Keep your website fast, lean and easy to move

A lightweight site is easier to cache, easier to serve from a content delivery network and easier to migrate. Caching static pages means visitors can still read your content while the origin struggles. Investing in speed and performance optimization pays off twice, and if your store runs on WordPress or Shopify, dedicated ecommerce support services help keep checkout and inventory sync working when an integration fails.

5. Monitor from the outside and alert early

Set up uptime monitoring that checks your site from several locations and alerts your team by phone or chat. The sooner you know, the sooner you can post a status message and pause ad spend. If phone leads matter, call tracking software can also reveal when inbound calls suddenly vanish, which is a useful early signal.

6. Write an incident plan and keep your platform maintained

A one page plan naming who decides, who communicates and what the customer message says beats a perfect plan that never gets written. Many small business outages are not caused by AWS at all. They come from outdated plugins, expired certificates or failed updates. Ongoing website maintenance and support keeps the basics healthy, and WordPress development services in Orange County can help with caching layers, staging environments and safe update workflows.

Do Small Businesses Need a Multi-Cloud Strategy?

Most small businesses do not need identical systems on two clouds. That approach is expensive, complex and can introduce new failure modes. A more realistic goal is graceful degradation. Decide which functions must stay online, which can be slow and which can wait, then make sure your backups, DNS and communication channels do not depend on the same provider as your main application.

If you are unsure how to weigh cost against protection, a short technical consultation can map your dependencies and identify the two or three changes that reduce risk most for the least spend. For a broader view, our overview of technology solutions for business is a useful companion read.

What the AWS Outage Means for SEO and Online Visibility

If search crawlers hit errors repeatedly, they may slow their crawl rate, and pages that fail for long periods can drop from results until they return. Short outages rarely cause lasting harm, but repeated failures can. Return proper server response codes during incidents, keep your XML sitemap accessible and check Search Console for crawl errors afterward.

Publishing a clear status page and a plain explanation also builds trust and captures demand that would otherwise go to news sites. Our guide to SEO content strategy explains how to map topics to search intent, and our SEO services page describes how technical health and content work together.

Automation, AI and the Growing Cost of Cloud Concentration

Businesses are automating more of their sales, support and operations, which makes cloud concentration a bigger issue over time. During the March 2026 disruption, reporting noted that several AI platforms hosted on AWS suffered consumer-facing outages. If your operations rely on agents, chatbots or automated workflows, give each one a manual fallback.

Automation is still worth pursuing, because well-built workflows recover faster and lose fewer leads than manual ones. Our overview of small business automation covers where it delivers the most value, and our piece on AI governance in business explains how to keep accuracy and accountability in step with speed.

A Practical Resilience Checklist for Business Owners

  • List every tool your website, checkout and lead flow depend on, and note where each is hosted.
  • Confirm backups are stored outside a single region and have been restored in a test.
  • Set up outside-in uptime monitoring with alerts to more than one person.
  • Add retry queues to forms, CRM syncs and payment callbacks.
  • Cache static pages and serve them through a content delivery network.
  • Keep plugins, themes and certificates current, and know how to pause ads quickly.

If you would rather have an expert audit your setup, book a call and walk through your dependencies with our team, or use the contact page to describe your situation. Our digital consulting and process automation service is built around exactly this kind of planning.

Frequently Asked Questions About the Amazon AWS Outage

What caused the October 2025 AWS outage?

AWS said a latent race condition in the DynamoDB DNS management system produced an incorrect empty DNS record for the regional endpoint in US-EAST-1, and the automation failed to repair it. That triggered a cascade affecting EC2 launches and Network Load Balancer health checks.

How long did the AWS outage last?

The October 2025 disruption lasted well over half a day for many customers, with one monitoring analysis reporting more than 15 hours. The May 2026 thermal event in one US-EAST-1 zone lasted about 28 hours.

Was the AWS outage a cyberattack?

No. Published analysis pointed to an internal software race condition in October 2025 and a cooling failure in May 2026. Early speculation about attacks was not supported by the official reports.

Will AWS compensate customers for downtime?

AWS service level agreements generally provide service credits, commonly a percentage of the affected service’s monthly bill. Credits do not cover lost revenue or reputational damage.

How do I protect my website from an AWS outage?

Store backups outside a single region, cache static content, remove hidden third-party dependencies, monitor from outside your network and keep a written incident plan.

Are finelo reviews and complaints related to AWS outages?

Not directly. Finelo is an investing education app, and its reviews focus mainly on subscription billing and refunds. The link is behavioral, because users who hit login or payment errors during any outage often search for reviews and complaints before realizing the cause is infrastructure.

Final Thoughts: Plan for the Outage You Cannot Predict

The October 2025 outage, the May 2026 thermal event and the Middle East disruptions all point to the same conclusion. Cloud infrastructure is remarkably reliable, yet it is not immune to software bugs, physical faults or events beyond anyone’s control. Reliability comes from design, not from trust in a single provider.

You do not need an enterprise budget to improve your position. Start with backups outside one region, outside-in monitoring and a written plan, then add fallbacks where the cost of failure is highest. Businesses that treat resilience as an ongoing habit are the ones whose customers barely notice when the next outage makes the news.\