What If Every Server on Earth Crashed at Once?
Technology

What If Every Server on Earth Crashed at Once?

• 7 min read

Happy System Administrator Appreciation Day. The sysadmins warned you. They sent emails about redundancy. They requested budget for backups. They put sticky notes on the server room door that said "DO NOT TURN OFF" and someone turned it off anyway. Now, at 00:00 UTC on the last Friday of July, every server on Earth crashes simultaneously.

Every rack in every data centre. Every cloud instance on AWS, Azure, Google Cloud, Alibaba Cloud. Every quietly humming box under a desk in a small business. Every Raspberry Pi running a home media server. Every virtual machine. Dead.

Not destroyed. Not wiped. Just crashed. Blue screen, kernel panic, power cycle required. Every single one, all at once.

The first second

The internet disappears. Not slowly, not in patches. Completely. Every website goes offline. Every API returns nothing. Every DNS server stops resolving. You type google.com and your browser stares at you like you've asked it something unreasonable.

Email stops. Not "delayed." Gone. Every email server on the planet has crashed. Your outbox contains unsent messages that may as well be written on paper for all the good they'll do.

Dark server room with all rack lights off

Cloud storage is inaccessible. Every Google Doc, every Dropbox file, every iCloud photo library. If it's not on a device in your hand, you can't reach it. The "cloud" turns out to have been other people's computers all along, and those computers are currently refusing to participate.

The first minute

Financial systems go dark. Stock exchanges halt. Not a trading pause. A total cessation. The New York Stock Exchange, the London Stock Exchange, the Tokyo Stock Exchange, NASDAQ: all run on servers. All down. Billions of dollars in trades that were mid-execution freeze in whatever state they were in when the servers died.

Payment processing stops. Visa processes approximately 65,000 transactions per second on a normal day. All of those transactions now fail. Every chip-and-PIN terminal, every contactless payment, every online checkout: declined. Not because your card is wrong but because the system that verifies it doesn't exist at the moment.

Cash becomes the only working payment method. ATMs, which are networked to bank servers, don't work either. Whatever cash you have in your wallet is your entire purchasing power until the servers come back.

The first hour

Hospitals run on servers. Patient records, imaging systems, lab results, drug interaction databases, appointment scheduling. Most modern hospitals have backup generators for power, but those keep the lights on and the ventilators running. They don't restart crashed server software. The IT department is already on the phone (phone networks are partly working, since mobile base stations have local processors, but anything that routes through a central server is broken).

Air traffic control relies on networked systems for radar processing and flight tracking. Planes already in the air can still fly (the aircraft's own computers are separate from ground servers) but ground-based traffic management is blind. Every airport in the world issues a ground stop. No takeoffs until the system is restored. Pilots in the air are directed to land at the nearest available airport using radio communication and manual procedures that most controllers haven't practised since training.

Emergency services are partly functional. The 999 and 911 call systems route through local switches that may or may not involve servers. Dispatchers lose their computer-aided dispatch systems. They're back to paper maps and radio. Ambulances, fire engines and police cars still work (they're vehicles, not servers) but coordination degrades badly.

The supply chain freezes

Modern logistics runs on servers. Warehouse management systems, fleet tracking, inventory databases, automated ordering. A Tesco supermarket doesn't know what's on its shelves without its stock management system. The lorries delivering tomorrow's bread are routed by server-based logistics software that is currently a blank screen.

Empty supermarket shelves with a few scattered products

Amazon's entire operation stops. Same for every e-commerce platform, every delivery service, every just-in-time manufacturing process. Modern supply chains carry almost no buffer stock because servers manage the flow in real time. Remove the servers and the flow stops. Within 48 hours, grocery shelves start to thin. Within a week, they're empty in major cities.

Fuel distribution is server-managed. Refineries, pipeline controls and petrol station point-of-sale systems all depend on networked computers. Pumps might still physically work (they're mechanical at the base level) but payment systems don't, and the logistics of getting fuel to stations depends on the same crashed software as everything else.

How long until recovery?

This depends on what "crashed" means. If every server needs a simple reboot, the recovery time is hours to days. Data centres have staff. Servers can be power-cycled. The systems come back, check their databases, reconcile their logs and resume.

But "all at once" creates a problem that individual crashes don't. When one server crashes, it can resync from others. When all of them crash simultaneously, there's nothing to resync from. Every database needs to recover from its own logs independently. Every distributed system needs to negotiate a consistent state with peers that were all in an unknown state at the moment of failure.

Distributed databases like Google's Spanner or Amazon's DynamoDB are designed to handle individual node failures, not total simultaneous failure of every node. The recovery procedures for "literally everything failed at once" are, in most cases, untested. Because nobody tests for that. Because it's not supposed to happen.

Realistic recovery time: major cloud providers back online in 6-24 hours. Smaller operations, days. Some systems with corrupted databases or incomplete transaction logs, weeks. A few, where the crash caused data corruption that can't be automatically resolved, permanently damaged.

What the sysadmins would say

They'd say "I told you so." And they'd be right. Every sysadmin has given a presentation about single points of failure that nobody listened to. Every sysadmin has requested disaster recovery budget that got cut. Every sysadmin has watched a manager dismiss redundancy planning as unnecessary expense because "the servers never go down."

The servers are down now.

System Administrator Appreciation Day exists because the people who keep the infrastructure running are invisible until the infrastructure stops. They're the IT equivalent of sewage workers: nobody thinks about them until the system backs up, and then suddenly everyone has urgent questions.

After recovery, sysadmins would briefly become the most valued employees in every company. Disaster recovery budgets would triple. Redundancy planning would become a board-level priority. This would last approximately eighteen months before everyone forgot, cut the budget again and went back to assuming the servers would never go down.

They would, of course, go down again. They always do. The only variable is whether anyone bothered to prepare, and the answer, reliably, is no.