IBM and Together AI have signed a multi-year agreement valued at $240 million to deploy a large NVIDIA infrastructure cluster on IBM Cloud, specifically designed for AI inference. The platform will utilize NVIDIA HGX B300 systems with Blackwell Ultra GPUs and Spectrum-X Ethernet fabrics, with availability expected in the first quarter of 2027.
The key points of the IBM and Together AI deal in 30 seconds
- The multi-year contract between IBM and Together AI reaches $240 million.
- IBM will deploy an NVIDIA HGX B300 system cluster on IBM Cloud, scheduled for Q1 2027.
- Together AI will mainly use the infrastructure for serving open models in production.
- The company currently processes approximately 400 billion tokens per month.
- The deal highlights the growing focus on inference over training within AI infrastructure.
The operation combines three different market segments. IBM provides cloud infrastructure and its ability to operate enterprise environments; NVIDIA supplies the accelerated platform and networking connecting the GPUs; and Together AI overlays its inference platform, training, fine-tuning, and agent-based workloads.
According to IBM, the result will be the first large-scale dedicated inference cluster built on IBM Cloud with HGX B300 systems.
The number of servers or GPUs comprising the cluster has not been disclosed, so the $240 million contract does not directly translate into a specific number of accelerators or computing power.
NVIDIA B300 for inference that increasingly demands more infrastructure
Choosing the NVIDIA HGX B300 enables understanding the direction of significant AI investment trends.
Each HGX B300 system integrates eight NVIDIA B300 Blackwell Ultra GPUs, connected via fifth-generation NVLink and NVSwitch. NVIDIA’s technical documentation specifies 288 GB of HBM3e memory per GPU, raising total memory to approximately 2.3 TB per node. The internal interconnect offers 14.4 TB/s of aggregate bandwidth.
Networking is also a key component of the architecture. NVIDIA pairs HGX B300 with ConnectX-8 and Spectrum-X Ethernet to build clusters where GPUs can exchange large amounts of data with low latency.
NVIDIA’s reference architecture contemplates configurations of 32, 64, and 128 nodes, corresponding to 256, 512, and 1,024 B300 GPUs respectively. Each GPU can support Ethernet connections up to 800 Gb/s via ConnectX-8.
This does not mean the IBM and Together AI cluster will necessarily use one of these specific configurations. The companies have not disclosed the size yet. However, it helps contextualize the type of infrastructure they are preparing.
Large-scale inference no longer involves just loading a model onto a few GPUs and responding to requests. Reasoning models, extended contexts, agent-based applications, and high numbers of concurrent users are increasing memory, communication, and computing requirements.
NVIDIA states that its current HGX architectures are designed precisely to combine training with inference for medium and large models.
Together AI already processes 400 billion tokens monthly
For Together AI, the new cluster means increased capacity at a time of strong growth in their inference service.
The company reports processing 400 trillion tokens per month, or about 400 billion tokens, using the numeric scale in Spanish.
Founded in 2022, the firm specializes in infrastructure and services for open models. Its platform includes inference, training, fine-tuning, and agent-based application tools.
Together AI recently closed a Series C funding round of $800 million, with a valuation of $8.3 billion, according to IBM’s announcement of the deal.
Having dedicated infrastructure also offers greater control over one of the main cost factors in inference: the expense of generating each token.
When millions of requests rely on the same GPUs, small improvements in utilization, memory, batching, networking, or software can lead to significant differences in overall costs.
Together AI already offers various generations of NVIDIA GPUs. As of August 2026, options include H100, H200, and B200, while HGX B300 configurations are available upon request.
Infrastructure is back at the center of AI competition
The deal also demonstrates how investment patterns around artificial intelligence are evolving.
During the early years of generative AI, much attention centered on building and training ever larger models. Now, massive deployment of these models is shifting focus increasingly toward inference.
Every query in an assistant, enterprise application, or agent consumes computing capacity. When a service reaches hundreds of billions of tokens, GPU performance and token costs become critical factors in determining its economic viability.
Together AI emphasizes that partnering with IBM and NVIDIA is driven by the need for GPU capacity aligned with their growth and to reduce token costs.
IBM also benefits by handling significant workloads that bolster its position as an AI infrastructure provider. The company has long advocated for a hybrid strategy allowing clients to combine on-premises data centers, cloud, and enterprise software.
The agreement extends a previous collaboration between IBM and NVIDIA. In March 2026, both companies announced new joint initiatives around GPU-accelerated analytics, unstructured data processing, and AI infrastructure in both private data centers and cloud environments.
Together AI now adds a dedicated layer focused on open model execution.
A key difference from just renting GPU resources on demand is the provision of dedicated capacity at scale over multiple years, giving Together AI greater predictability to support their growth.
Production deployment is expected, if schedules are met, in the first quarter of 2027. Until then, some key details remain unknown: number of HGX B300 systems, total GPUs, cluster location, associated power requirements, and final inference capacity.
FAQs
How much is the IBM and Together AI deal worth?
The multi-year agreement announced by both companies is valued at $240 million. Its purpose is to deploy AI infrastructure on IBM Cloud for Together AI’s services.
What GPUs will the new cluster use?
The infrastructure will be based on NVIDIA HGX B300 systems, which feature eight B300 Blackwell Ultra GPUs per system. IBM and Together AI have not disclosed the number of nodes or GPUs planned for deployment.
What will Together AI use this infrastructure for?
Primarily for running inference on open models at large scale. Together AI also offers training, fine-tuning, and agent-based applications.
When will the cluster be available?
IBM anticipates availability in the first quarter of 2027. The company notes that their statements about future plans are targets and may change.
via: newsroom.ibm

