IBM and Together AI have signed a multi-year contract worth $240 million. IBM is deploying a large cluster of Nvidia HGX B300 systems on IBM Cloud, which Together AI will use for inference of open-source models. The capacity is expected to be available in the first quarter of 2027.
According to IBM, this is the first dedicated, large-scale cluster on IBM Cloud built specifically for inference. In addition to the HGX B300 systems, Nvidia Spectrum-X Ethernet will be used for the network layer. Nvidia states that this combination delivers 30 times more AI factory output than the previous generation.
With this cluster, Together AI is targeting workloads from companies that want to run open models without the cost of proprietary alternatives. Together AI chose IBM and Nvidia for their product roadmaps and the speed at which GPU capacity can be delivered at the lowest possible token cost. The company’s platform covers inference, training, fine-tuning, and agentic workflows. For its inference product, Together AI now reports 400 trillion tokens per month.
“Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale,” says Vipul Ved Prakash, CEO of Together AI. Partnering with IBM and Nvidia provides that foundation, he notes.
Broader IBM-Nvidia collaboration
Alan Peacock, General Manager of IBM Cloud, discusses scalable, cost-effective, enterprise-level AI infrastructure. Dion Harris of Nvidia compares AI factories to electricity and telecommunications as essential business infrastructure.
The cluster is the latest step in a broader collaboration. IBM and Nvidia recently reported progress on GPU-native data analysis, unstructured data extraction, on-premises and cloud infrastructure, and consulting services. IBM had previously expanded its hybrid cloud options for AI as well.