SpaceXAI has proven it can stand up AI infrastructure at a pace that’s hard to match: the first Colossus cluster, with roughly 100,000 NVIDIA H100 GPUs and about 130 MW of compute power, came online in 122 days. Colossus II was even faster. But a report from The Information points to a different problem: the earliest facilities reportedly suffered frequent outages, and SpaceX is said to be reworking part of their design to boost reliability. The question is no longer just how many GPUs it can install, but how much redundancy it needs to turn that capacity into a cloud service it can sell reliably to third parties.
Colossus and the reliability challenge in 30 seconds
- SpaceXAI stood up Colossus’s first 100,000 H100s in 122 days, and Colossus II’s first block in 91.
- Its two facilities together total roughly 1 GW of compute power, according to SpaceX documentation.
- The Information reports the earliest facilities suffered frequent outages and that parts of them are now being redesigned.
- SpaceX is already selling capacity to third parties, including Google, raising the bar for availability.
- Power, batteries, cooling, networking and redundancy could determine how far this model can scale.
The shift matters because Colossus can no longer be analyzed simply as the supercomputer used to train Grok. Following xAI’s integration into SpaceX, the company is now selling part of its massive GPU infrastructure to outside customers.
A filing with the US Securities and Exchange Commission (SEC) shows just how far that shift goes. SpaceX signed a contract with Google on June 5, 2026 to provide capacity based on approximately 110,000 NVIDIA GPUs, plus CPUs, memory and other components. Google agreed to pay $920 million a month between October 2026 and June 2029, though the contract includes termination conditions and adjustments if SpaceX fails to deliver the committed capacity.
With contracts of that size, building fast is still an advantage. Keeping the infrastructure running is now just as important.
From 100,000 H100s in 122 days to more than 1 GW of compute
The Colossus story begins with a decision that explains much of its speed.
Instead of building a conventional data center from scratch, xAI repurposed a former Electrolux factory in Memphis. The company brought in roughly 100,000 NVIDIA H100 GPUs and about 130 MW of compute power in 122 days.
SpaceX now uses that timeline as a selling point. In its SEC filing, it compares the project to a reference timeline of roughly two years to bring a 100 MW greenfield data center online. That comparison comes from the company itself and doesn’t mean the two kinds of projects are directly equivalent.
Colossus kept growing. SpaceXAI says it later doubled the system to about 200,000 GPUs over another 92 days. Its public materials cite more than 0.5 exabytes of storage and up to 2.8 Tb/s of per-server connectivity for distributed workloads.
Colossus II pushed the model even further.
According to SpaceX’s documentation for investors, the first cluster added roughly 110,000 NVIDIA GB200 GPUs and 210 MW of compute in 91 days.
The second block took just 64 days to add another 110,000 GB300 processors and around 220 MW.
| Deployment | Hardware reported | Compute power | Time |
|---|---|---|---|
| Colossus, phase one | ~100,000 H100 | ~130 MW | 122 days |
| Colossus, expansion | up to ~200,000 GPUs | — | +92 days |
| Colossus II, first cluster | ~110,000 GB200 | ~210 MW | 91 days |
| Colossus II, second cluster | ~110,000 GB300 | ~220 MW | 64 days |
| Colossus + Colossus II | multiple generations | ~1 GW | combined capacity |
These figures come mainly from SpaceX and should be understood as company-reported data. Independent research group Epoch AI uses satellite imagery, permits and cooling models to estimate the capacity of major AI facilities, and currently places Colossus II above a million H100-equivalent GPUs, though that’s an estimate, not an official inventory of physical GPUs.
The distinction matters. An “H100-equivalent” tries to normalize compute capacity across different generations and doesn’t mean that exact number of physical H100s is actually installed.
Building fast and designing for failure aren’t exactly the same problem
Extreme speed has a potential trade-off.
The Information reported in September that the first data centers xAI stood up suffered frequent outages. According to sources cited by the outlet, the earliest phase reportedly prioritized getting compute up and running quickly over deploying all the infrastructure meant to boost reliability.
The report also says SpaceX reorganized its data center team over the summer and brought in engineers from its rocket-launch organization. Part of the existing infrastructure is reportedly being redesigned.
SpaceX hasn’t published detailed metrics that would let outsiders quantify these alleged outages, such as annual availability, average duration, affected GPUs or specific causes. So it isn’t possible to say publicly which components failed or to attribute the incidents to a specific redundancy gap.
The report does raise a relevant technical question, though.
An internal training cluster and an infrastructure service sold to third parties can tolerate risk very differently.
To train a model, distributed software can save checkpoints and resume work after certain failures. Losing a node also doesn’t necessarily mean stopping the whole system if the architecture is designed to isolate it.
But the bigger the cluster, the larger the absolute number of components that can fail.
Hundreds of thousands of GPUs also mean servers, power supplies, switches, optical links, cooling systems, storage, pumps, transformers and miles of cabling. The availability of the whole system doesn’t depend solely on the GPUs themselves being reliable.
Traditional data center redundancy is specifically designed to keep a single critical component’s failure from taking down too large a share of the infrastructure.
That can mean power delivered via different paths, UPS systems, batteries, generators, extra cooling capacity, redundant pumps, networks with alternate routes, or architectures that can isolate racks and servers without stopping the rest.
But every added layer costs money, takes up space and can slow down deployment.
That’s Colossus’s dilemma: part of its edge comes precisely from cutting or simplifying the processes that slow down traditional construction.
Power shows how the design is evolving
The electrical infrastructure around Memphis offers a window into that evolution.
SpaceX explained in regulatory filings that Colossus and Colossus II were initially brought online using largely on-site generation. The company considers its ability to build that energy infrastructure quickly to be part of its competitive advantage.
That strategy has also drawn controversy.
Reuters reported in July on the use of dozens of gas turbines in Tennessee and Mississippi and on the regulatory and legal dispute over their permits and emissions. SpaceXAI maintains these installations are temporary while it moves toward permanent solutions.
The company later said it will retire the temporary turbines as it moves toward a permanent 1.2 GW gas plant, though the timeline extends to July 2027.
Generation is being paired with electrical storage at an unusual scale.
Canary Media identified 720 Tesla Megapacks at Colossus II via satellite imagery in July. Depending on the specific version installed, the outlet calculates the array could approach 2.8 GWh and deliver between roughly 720 MW and 1,400 MW of power. SpaceXAI hasn’t yet published a full technical spec sheet confirming those figures.
A battery array of that size can serve several functions: stabilizing power delivery, absorbing transients, managing demand peaks or sustaining loads during certain grid events. Its existence alone doesn’t reveal the facility’s full redundancy architecture, but it does show the power system is becoming more complex.
SpaceXAI is also seeking more supply from the Tennessee Valley Authority (TVA) while developing its own generation.
The result is moving further away from the image of a building full of GPUs hooked up to some temporary turbines. Colossus is becoming an energy-and-compute campus with multiple sources, storage and grid connections.
Selling compute changes the rules
Reliability carries even more weight because SpaceXAI is opening Colossus up to other companies.
In May, it announced an agreement to provide Colossus 1 capacity to Anthropic. According to SpaceXAI, that facility has more than 220,000 NVIDIA GPUs, across H100, H200 and GB200. Anthropic said it would use the additional capacity for its Claude Pro and Claude Max services.
The Google deal offers a more detailed look at the commercial consequences of failing to deliver capacity on time.
The SEC filing states that if SpaceX doesn’t provide the committed number of GPUs by September 30, 2026, Google can, after a one-month grace period, terminate the agreement or accept a lower amount with a proportional reduction in payments.
That turns availability into money.
When Colossus was mainly xAI’s internal infrastructure, an outage could translate into training delays and operating costs. When the same infrastructure is sold to third parties for hundreds of millions of dollars a month, delivered and available capacity becomes a contractual obligation.
And that shift helps explain why Colossus’s next phase may depend less on breaking another construction record and more on finding the balance between speed, cost and redundancy.
SpaceX already knows this problem well in another domain.
Its space business depends on designing systems that can tolerate failures, detect anomalies and keep operating with redundant components. The fact that engineers from its launch organization are now working on data center infrastructure, as The Information reports, fits with a push toward more reliability-focused engineering, though there still isn’t enough public information to know exactly what changes are being applied.
Scale doesn’t seem to be slowing down either. SpaceX reports roughly 1 GW of combined compute power between Colossus and Colossus II, while maintaining additional energy capacity to keep growing. Its stated public goal for Memphis has reached as high as a million GPUs.
The first Colossus proved a company could turn a factory into one of the world’s largest AI clusters in four months. Colossus II proved that process could be accelerated even further.
The next test is different. SpaceXAI has to prove that this way of building can also deliver the availability that outside customers expect when they’re contracting tens of thousands of GPUs.
In AI infrastructure, getting 100,000 accelerators up and running quickly is a competitive edge; keeping them available when something breaks is what turns them into cloud.

