All posts

3GW data center outage exposes fragile grid balance

Manaal KhanJuly 26, 2026 at 3:32 PM7 min read
3GW data center outage exposes fragile grid balance

Key Takeaways

A Drive Through 'Data Center Alley'

3GW data center outage exposes fragile grid balance
Source: TechCrunch
  • A downed power line caused 3.1GW of data centers to disconnect in 30 seconds, spiking voltage across PJM's grid from Northern Virginia to Chicago
  • The incident was twice as large as a similar 2024 event, showing the problem is accelerating as AI data centers proliferate
  • Solutions exist: ON.Energy is installing 3GW of campus-wide UPS systems that hide data centers behind battery banks, allowing them to absorb grid fluctuations

A single power line fell outside Washington, DC this week. What should have been a routine grid hiccup instead triggered 3.1 gigawatts of data centers to simultaneously switch to backup power, yanking their load from the grid in about 30 seconds. The result: voltage spikes from Northern Virginia to Chicago that took 11 minutes to stabilize and caused lights to flicker across the region.

The incident didn't cause a blackout. But it demonstrated something grid operators and data center builders can no longer ignore: the AI boom has concentrated so much computing power in so few locations that a minor fault can cascade into a major grid disruption. Northern Virginia hosts the world's highest concentration of data centers, and they just proved they can act like a single, massive, unpredictable load.

Advertisements

What actually happened to the PJM grid?

PJM Interconnection manages grids from New Jersey to Illinois, serving 67 million customers. It's the largest grid operator in the United States. When that power line went down, data centers sensed the voltage dip and did exactly what they're designed to do: switch to backup power to protect their servers.

The problem is they all did it at once. According to PJM data, 3.1 gigawatts of load vanished within 30 seconds. The grid partially recovered, then more loads dropped off. At peak, PJM had an extra 3.49 gigawatts of electricity with nowhere to go.

3.49 GW
Peak excess electricity on PJM's grid during the event, roughly 3% of total demand

Three percent of total demand sounds manageable until you remember that electrical grids operate in near-perfect balance. Supply and demand must match closely, or voltages sag or spike. Small fluctuations are tolerable. Large ones trigger failsafes. And when dozens of data centers trip their failsafes simultaneously, they amplify the very problem they're trying to escape.

"It's the canary in the coal mine," Ricardo de Azevedo, CTO at ON.Energy, told TechCrunch. These sorts of events involving large loads like data centers are "happening more and more."

Data collected by Ting Labs, a startup running an IoT sensor network through residential electrical sockets, confirmed that voltage fluctuations rippled across the entire PJM territory. Lights flickered in homes hundreds of miles from the original fault.

Why did data centers all disconnect at once?

Data centers make power decisions in milliseconds. When voltage dips, their transfer switches flip to backup generators or batteries almost instantly. That's by design. You don't want servers crashing because a transformer blew three blocks away.

But Ali Zain Banatwala, senior market models specialist at the Independent Electricity System Operator, identified the core issue: data centers located near each other all sensed the same voltage dip at the same time and all decided to disconnect within seconds of each other.

"We need to figure a way for these loads that are located next to each other to sequentially either disconnect or reconnect," Banatwala told TechCrunch. A staggered approach would give grid operators time to compensate. The current approach, where everyone runs for the exits simultaneously, makes things worse.

Also Read
OpenAI admits its AI autonomously hacked Hugging Face

Another example of AI infrastructure creating unexpected cascading effects

This problem is getting worse, not better

The July 2026 event was twice as large as a similar incident in 2024, when 60 data centers disconnected simultaneously and pulled 1.5 gigawatts from the grid. In two years, the potential disruption doubled. Given the pace of AI infrastructure buildout, there's no reason to think the trend will reverse.

AI training runs are particularly demanding. A single large language model training job can require sustained power draws measured in hundreds of megawatts. When multiple hyperscalers cluster their AI facilities in the same region for proximity to fiber infrastructure and power substations, they create concentrated loads that grids weren't designed to handle as single points of failure.

Northern Virginia became the world's data center capital because of its fiber connectivity, favorable power rates, and proximity to government customers. Those same advantages now make it a single point of vulnerability for the entire PJM grid.

What solutions exist for data center grid stability?

ON.Energy has developed what amounts to an uninterruptible power supply for entire data center campuses, covering servers, chillers, and all other equipment. The concept: hide the data center behind a bank of batteries connected to sophisticated power conversion equipment. From the grid's perspective, the facility appears as one consistent, predictable load rather than a volatile collection of individual systems.

The system can absorb grid fluctuations rather than reacting to them. When voltage spikes, extra power charges the batteries. When voltage dips, the batteries dispatch power to servers. The whole system tracks the grid's state within milliseconds, preventing the kind of sudden load drops that caused this week's problems.

ON.Energy is currently installing 3 gigawatts worth of these systems across four data center campuses, according to de Azevedo. That's roughly the same amount of load that disconnected during the incident, which suggests the company sized its initial deployments to address exactly this scale of problem.

ℹ️

Logicity's Take

This incident should reshape how CTOs evaluate colocation providers. Ask whether the facility has campus-wide power buffering or just rack-level UPS. Providers like Equinix, Digital Realty, and QTS all tout uptime SLAs, but those metrics focus on your servers staying online, not on whether the facility is a good grid citizen. The next version of due diligence needs to include grid stability architecture. Expect ERCOT's new 'ride through' requirements to spread to other grid operators, potentially adding compliance costs for data center operators who haven't invested in buffering technology.

Advertisements

Regulators are starting to act

Grid managers have noticed the pattern. ERCOT, which manages the Texas grid, will require large loads like data centers to "ride through" disruptions rather than disconnecting, de Azevedo said. The specifics of what counts as a successful ride-through aren't yet public, but the direction is clear: data centers will need to absorb short-term grid fluctuations rather than making them worse.

PJM hasn't announced similar requirements, but after two incidents in two years, with the second being twice as severe, regulatory pressure seems inevitable. The 67 million customers who depend on PJM for electricity shouldn't have their lights flicker because a hyperscaler's transfer switch is too twitchy.

Also Read
Claude Opus 5 beats Fable 5 on most benchmarks at lower cost

AI model efficiency gains could reduce per-query power consumption at data centers

Approaches to data center grid stability

ApproachHow it worksGrid impactAdoption status
Traditional UPSRack-level battery backup for brief outagesStill disconnects from grid during eventsUniversal
Campus-wide battery bufferingEntire facility behind battery bank with power conversionAbsorbs fluctuations, appears as steady load3GW being installed (ON.Energy)
Staggered disconnection protocolsCoordinate neighboring facilities to disconnect sequentiallySpreads load drop over timeProposed, not implemented
Ride-through mandatesRegulatory requirement to absorb short fluctuationsPrevents mass disconnection eventsERCOT adopting; other grids considering

What this means for companies running AI workloads

If you're running AI training or inference at scale, your facility's power architecture is now a strategic concern. The July 2026 incident didn't cause downtime for the data centers involved; their backup systems worked. But it did demonstrate that facilities optimized purely for uptime can externalize costs onto the grid and, ultimately, onto regulators who will respond with mandates.

Companies building or leasing new AI compute capacity should ask providers specifically about grid interaction architecture. Is the facility designed to ride through voltage fluctuations? Does it use campus-wide buffering? How does it coordinate with neighboring facilities during grid events?

The answers will increasingly affect both operational reliability and regulatory compliance. ERCOT's requirements are likely the beginning, not the end, of a regulatory trend that will spread to other grid operators as AI infrastructure continues to concentrate.

Also Read
Google Maps defaults to fuel-efficient, not fastest routes

Another example of infrastructure efficiency becoming a default rather than an option

Frequently Asked Questions

How much power did the data centers remove from the grid?

Approximately 3.1 gigawatts disconnected within 30 seconds, with total excess power peaking at 3.49 gigawatts. This represented about 3% of PJM's total demand at the time.

Why can't the grid handle a 3% load change?

Electrical grids must maintain near-perfect balance between supply and demand. The problem wasn't the 3% figure itself but the speed: 3.1GW vanishing in 30 seconds gives operators no time to compensate. The same load change spread over 10 minutes would be manageable.

Will this cause data center construction to slow?

Unlikely. Demand for AI compute continues to grow. But facilities will face new design requirements and regulatory mandates. ERCOT's ride-through requirements signal where other grid operators are headed.

How does ON.Energy's solution differ from standard UPS systems?

Traditional UPS protects individual racks for brief outages but still disconnects from the grid during events. ON.Energy's campus-wide system puts the entire facility behind a battery bank with power conversion, so the grid sees a steady load and the facility absorbs fluctuations rather than reacting to them.

How long did the grid take to stabilize?

Approximately 11 minutes from the initial disconnection cascade. The grid partially recovered after the first wave of disconnections, then experienced additional load drops before finally stabilizing.

ℹ️

Need Help Implementing This?

If you're evaluating data center providers or designing power architecture for AI workloads, Logicity's consulting team can help assess grid stability requirements and vendor options. Contact us at consulting@logicity.in.

Source: TechCrunch / Tim De Chant

M

Manaal Khan

Tech & Innovation Writer

Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.

Related Articles