OpenAI is developing Jalapeño as an inference accelerator built primarily to cover its own compute needs. Richard Ho, the company’s head of hardware, told Tom’s Hardware that the chip could be used by other customers, though the priority for the next few years will be meeting OpenAI’s own growing internal demand.
Jalapeño in 20 seconds
- Jalapeño is OpenAI’s first in-house inference accelerator, developed together with Broadcom.
- The company plans to deploy it within its compute infrastructure by the end of 2026.
- OpenAI says it also works with outside models such as DeepSeek R1 and Kimi K2.5.
- Ho leaves the door open to future availability for third parties, but the current priority is internal consumption.
- The design is part of a multi-generation platform, with Gen 2 already in development.
OpenAI’s position is more pragmatic than commercial. The company isn’t presenting Jalapeño as a product that will immediately compete with NVIDIA’s accelerators on the market, but as a piece of its own infrastructure. Ho acknowledges the hardware could be used outside OpenAI, but he believes internal capacity demand will be enough to absorb most of the production for quite a while.
A Chip Built Around Inference
OpenAI introduced Jalapeño in June 2026 together with Broadcom. The company describes it as its first Intelligence Processor and as the first generation of a compute platform that will span several generations. The chip was designed specifically for large language model (LLM) inference workloads, rather than adapting an existing accelerator to OpenAI’s needs.
The project went from initial design to manufacturing in nine months, according to OpenAI. Broadcom contributes the silicon implementation and networking technologies, while Celestica handles the manufacturing of boards, rack systems, and other elements needed to bring the accelerator to large-scale facilities. OpenAI, for its part, designed the architecture around its own models, kernels, inference software, and real serving patterns.
The company has explained that one of its main goals was improving efficiency. In a data center dedicated to artificial intelligence, increasing the useful work each watt can deliver has a direct effect on how much capacity can be deployed with a given power infrastructure.
The first results OpenAI published in August point in that direction. In its tests with GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, Jalapeño delivered between 1.5 and 1.9 times more AI work per watt at peak performance compared with the systems used for reference. The company also measured end-to-end latency 1.7 to 3.6 times lower and, on highly interactive workloads, 2.1 to 4.1 times more throughput. These are figures published by OpenAI, so they describe its own measurements and methodology.
The Door to Third Parties Stays Open
The possibility that Jalapeño ends up outside OpenAI hasn’t been ruled out. Ho told Tom’s Hardware that the accelerator is programmable and general-purpose within its scope, and that the team was able to get models running that OpenAI hadn’t used during initial development.
That point matters because one of the open questions around an ASIC — an integrated circuit designed for a specific application — is how tightly it ends up tied to a single model or workload. OpenAI says Jalapeño isn’t hard-coded exclusively for its own models.
The team demonstrated this with outside models. According to OpenAI, Jalapeño has run GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. Ho also said engineers managed to get DeepSeek and Kimi running roughly two months after receiving an A0 sample of the chip, ahead of presenting the results at Hot Chips.
That doesn’t mean OpenAI is about to sell Jalapeño as a commercial product. Ho’s answer makes the current situation clear: the company has internal compute demand that keeps growing with its user base, new models, Codex, and other features, and it wants to cover those needs first. The possibility of using it elsewhere stays open, but without a timeline or a commercial announcement.
There’s also a manufacturing capacity issue. Deploying in-house accelerators at scale requires not just building the chip, but also having memory, networking, boards, complete systems, and data center capacity available. The strategy announced with Broadcom envisions gigawatt-scale deployments across several generations, while OpenAI continues using NVIDIA accelerators and other suppliers for both training and inference.
Jalapeño Is Part of a Bigger Strategy
OpenAI isn’t treating Jalapeño as a one-off project, either. The company is already talking about a second generation in development and a third that’s starting to take shape. The goal is for each generation to build on the experience gained from the previous ones to improve speed and efficiency.
The strategy fits the growing interest among major model developers in controlling more of the infrastructure stack. Co-designing models, inference software, memory, networking, and silicon makes it possible to tailor hardware to real production workloads, something an outside vendor has to serve across many different customers instead.
For OpenAI, that integration has a direct payoff: running its own models with lower power consumption and lower latency can increase available capacity without relying solely on buying more general-purpose accelerators. At the same time, though, the company is keeping a multi-vendor strategy, so Jalapeño doesn’t mean an immediate replacement for NVIDIA or the rest of the hardware used in its data centers.
The commercial question, then, takes a back seat. OpenAI has an accelerator that, according to its own tests, can run both its own and outside models with good results, but its first job will be serving the company’s own demand. If enough manufacturing capacity exists down the road and the architecture keeps its edge over other alternatives, the same design could eventually be used outside OpenAI. For now, that possibility remains open and isn’t an announced product for external customers.
Frequently Asked Questions
What Is OpenAI’s Jalapeño?
Jalapeño is the first inference accelerator designed by OpenAI. It’s built to run large language models with a particular focus on efficiency and low latency.
Will Jalapeño Be Sold to Other Companies?
OpenAI hasn’t announced a commercial launch for third parties. Richard Ho has said the chip could be used outside the company, but the current priority is covering internal compute needs.
What Models Can Jalapeño Run?
OpenAI has published results with GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, in addition to its own workloads.
When Will OpenAI Start Using Jalapeño?
OpenAI plans to start deploying the chip within its compute infrastructure by the end of 2026. The company is already working on a second generation.

