Server CPUs Now Take 25-30 Weeks to Ship as AI Agents Enter the Picture

Server processors are joining the list of components running short because of artificial intelligence. Consultancy TrendForce now puts server CPU lead times at 25 to 30 weeks, compared with the 16 to 20 weeks it considers typical of a balanced market. After GPUs, HBM memory, and enterprise SSDs, the pressure has reached a component that in recent years had taken a back seat to accelerators.

The server CPU shortage in 30 seconds

  • TrendForce raises server CPU lead times to 25-30 weeks, up from 16-20 weeks under normal conditions.
  • AI agents like Meta Muse use one virtual machine per user, each with its own CPU, memory, and storage.
  • A scenario with 100 million daily users would require about 12.5 million physical cores, not 50 million.
  • There’s no proof that Muse is the main cause: demand for AI servers and preventive buying also play a role.

The figure comes from TrendForce’s weekly bulletin, which links the shift to the rise of AI agents, and has been picked up by outlets such as Wccftech. It’s worth reading with some nuance. Longer lead times are a market signal, but on their own they don’t say how much new demand there is or exactly where it’s coming from.

Why agents need CPUs, not just GPUs

A traditional chatbot receives a question and returns an answer generated on a GPU. An agent does more: it runs multi-step tasks, saves its state, opens files, runs commands, and calls external tools. All of that happens in a conventional compute environment, with a processor, memory, and disk, while the GPU only steps in when the model needs to think through its next action.

Meta Muse, the personal agent Meta launched in September, is the most visible example. Each user gets their own computer in the cloud — a Linux virtual machine configured with 2 vCPUs, 8 GB of RAM, and 100 GB of SSD storage. According to researchers who questioned the agent itself, cited by Tom’s Hardware, those machines run on servers with AMD EPYC Turin processors. Muse passed half a million daily active users in its first week and, according to several reports, has already shown capacity problems.

This publication already covered Meta’s push in this area with Muse Code, its programming agent. What matters now is that a consumer product with hundreds of millions of potential users turns every single person into a running virtual machine — and that translates directly into servers.

A vCPU isn’t the same as a physical core

The risk in this debate is overestimating demand. Assigning 2 vCPUs to a virtual machine doesn’t mean continuously using two physical cores. Agents spend a good part of their time waiting: for the model’s response, for the user’s, for a tool’s, or for the network. That’s why cloud providers can oversubscribe — that is, hand out more vCPUs than physically exist.

DeepSeek’s case illustrates this. Its DSec platform, described in a technical paper published this month, maintains more than 380,000 simultaneous isolated environments (sandboxes) for training agents on a production unit of about 160 nodes. According to the paper itself, around 90% of those environments use no more than 5% of the processor capacity they request. The gap between what’s allocated and what’s actually consumed is enormous.

Using that logic, analyst Freda Duan has outlined a scenario for Muse. If it reached 100 million daily users with two hours of average activity, and you apply a peak-to-average factor of 2.5x plus a 20% margin, that would require about 25 million virtual machines active at once. At 2 vCPUs each, that’s 50 million vCPUs — but with an estimated 0.5 physical cores per machine, the real figure would be around 12.5 million cores.

Scenario assumptionValue
Daily active users100 million (hypothetical)
Average activity per user2 hours a day
Peak-to-average factor and margin2.5x and 20%
Simultaneous active virtual machines25 million
Allocated vCPUs (2 per VM)50 million
Physical cores (0.5 per VM)12.5 million

It’s an estimate built on debatable assumptions, not a figure from Meta. Muse is far from 100 million users today, and actual usage per machine could be higher or lower. But it’s useful for gauging the order of magnitude — and for understanding why comparing prices or capacity by vCPU can be misleading, something also seen in enterprise cloud pricing.

Muse doesn’t explain all the pressure

Pinning the longer lead times solely on Muse would be premature. The product has only been on the market for a few weeks, and processor orders are planned months in advance. The pressure predates it: back in April, TrendForce had already lowered its forecast for general server shipments in 2026 due to component shortages, and warned that suppliers were prioritizing capacity for the more profitable AI servers.

Other factors matter too. Every AI server carries high-end CPUs that orchestrate the GPUs, and demand for that equipment remains strong. When lead times stretch out, many buyers also place preventive orders to secure supply, which tightens the market even further. And the shortage of DRAM memory and enterprise SSDs is already making full server configurations more expensive and slower to deliver.

If the demand holds up, the winners won’t be limited to Intel, AMD, and Arm-based designs. Server makers, cloud providers, memory and storage vendors, networking companies, and data center operators stand to benefit too. The open question is what will weigh more over the coming quarters: the rapid adoption of agents, or cloud providers’ ability to pack more virtual machines onto each physical core.

Frequently Asked Questions

How long does it currently take to get server CPUs delivered?

According to TrendForce, between 25 and 30 weeks, compared with the 16 to 20 weeks of a balanced market.

Why do AI agents need CPUs?

Because they run multi-step tasks, maintain their state, and use tools inside virtual machines with a processor, memory, and disk, in addition to the GPU that generates the responses.

What virtual machine does Meta Muse assign to each user?

A machine with 2 vCPUs, 8 GB of RAM, and 100 GB of SSD storage, hosted on servers with AMD EPYC Turin processors, according to published reports.

Why can’t demand be calculated just by adding up vCPUs?

Because agents spend a lot of time waiting, and providers oversubscribe physical cores. An allocated vCPU isn’t the same as a core used continuously.

Sources:

Scroll to Top