
IBM and Together AI have signed a $240 million multi-year agreement to build a large AI inference cluster on IBM Cloud using Nvidia’s latest Blackwell Ultra infrastructure, giving IBM a concrete role in the spending wave moving from AI model training toward large-scale production use. IBM said the planned deployment will use Nvidia HGX B300 systems and Spectrum-X Ethernet networking, with availability expected in the first quarter of 2027.
Together AI will use the capacity to provide inference for open-source models, meaning the computing work that happens after a model has been trained and is generating responses for users and applications. For investors, the deal is less about a single contract changing IBM’s financial profile and more about where cloud infrastructure demand is moving.
Together AI says its inference service is already handling about 400 trillion tokens a month, and its recent financing plans point to a much larger compute footprint. The agreement also places IBM inside an Nvidia-centered ecosystem that is attracting large amounts of capital from cloud providers, AI startups and financial institutions.
IBM’s cluster is built for production inference
The deployment will be IBM Cloud’s first dedicated large-scale inference cluster using Nvidia HGX B300 systems and Spectrum-X Ethernet, according to IBM. The company did not disclose the number of systems or GPUs involved, the payment schedule, expected margins or the amount IBM will spend to build the capacity.
Those omissions matter when interpreting the headline value. The $240 million figure is the value of a multi-year agreement between IBM and Together AI. It should not be treated as $240 million of immediate revenue for IBM, and the announcement does not specify how the contract will be recognized in IBM’s financial statements over time.
The hardware choice is still notable. Nvidia introduced the HGX B300 NVL16 as part of its Blackwell Ultra platform and says the system can deliver 11 times faster large-language-model inference, seven times more compute and four times more memory than its Hopper generation. Spectrum-X is designed to connect accelerated computing systems over Ethernet at data-center scale, where network performance can become a bottleneck as many GPUs work on the same workloads.
That combination fits Together AI’s business model. The company provides infrastructure and software for running open and custom AI models, including inference, training and post-training workloads. Rather than selling a closed model of its own as the only option, Together AI competes partly on giving customers access to a broader open-model ecosystem and on lowering the cost of serving those models at scale.
The economics are increasingly important as companies move AI applications from experiments into production. Training a model can be an enormous one-time or periodic compute job, but inference becomes a recurring operating workload as users send prompts, agents take actions and applications generate responses. At high usage levels, small differences in hardware utilization, networking efficiency and cost per token can become financially meaningful.
Together AI is committing to a much larger compute footprint
The IBM agreement follows a major capital raise at Together AI. On July 1, the company announced an $800 million Series C funding round that included Nvidia among its investors. Together AI also said it had secured commitments for more than 500 megawatts of compute capacity that would be capitalized independently by new investors to support its expected growth.
That 500-megawatt figure is a capacity commitment, not a disclosed spending total, and Together AI did not say the IBM deployment accounts for all or a specified portion of it. Still, the two announcements point in the same direction: the company is preparing for materially higher demand for inference and other production AI workloads.
IBM’s August 11 announcement adds a current usage measure to that picture. The company said Together AI’s inference product is serving about 400 trillion tokens per month. IBM also cited Together AI’s $8.3 billion valuation following the July financing. These are company-reported figures, but they help explain why Together AI is signing for dedicated infrastructure rather than relying only on short-term, on-demand capacity.
For Together AI, securing a dedicated cluster can provide more control over capacity and economics as usage expands. For IBM, the agreement offers a way to participate in AI growth through cloud infrastructure even when the underlying models are developed outside IBM. The customer relationship is therefore different from a straightforward sale of IBM software: the attraction is the combination of IBM Cloud operations, Nvidia accelerated computing and Together AI’s model-serving platform.
The deal gives IBM a place in Nvidia’s AI infrastructure buildout
IBM is pursuing the contract at a time when its own infrastructure results are mixed. In the second quarter of 2026, IBM reported $17.2 billion of total revenue. Infrastructure revenue was $3.8 billion, down 7% from a year earlier, although distributed infrastructure grew 37%. IBM also said its Power and Storage businesses had built an order backlog of nearly $500 million.
The Together AI agreement does not map neatly onto one reported segment from the information disclosed, so it would be premature to estimate its earnings impact. What it does show is IBM using its cloud platform to win a workload tied directly to the newest generation of Nvidia systems, rather than leaving AI infrastructure demand entirely to hyperscale cloud providers and specialist GPU clouds.
Nvidia’s own actions underline how much capital the broader buildout may require. On August 10, Nvidia announced memorandums of understanding with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR for independent compute-financing platforms intended to mobilize more than $500 billion of third-party capital for AI infrastructure over time. Nvidia said those partnerships are meant to create financing pools for customers seeking large amounts of compute, although the arrangements remain subject to final agreements.
The IBM-Together AI contract is much smaller than that proposed financing effort, but it is a concrete example of the downstream demand such capital is meant to serve. An AI developer needs models and software, but production inference also requires GPUs, networking, data-center capacity, power, financing and cloud operations. Spending can therefore spread well beyond the companies building frontier models themselves.
There is also a timing point for investors. IBM expects the new Together AI capacity to be available in the first quarter of 2027, so the operational milestone is still several months away. Until the cluster is deployed and customer workloads begin running on it, the key facts are the signed multi-year agreement, the planned Nvidia hardware stack and Together AI’s commitment to scale inference on IBM Cloud. The next concrete test will be whether IBM and Together AI bring that capacity online on the stated schedule and whether either company provides more detail on deployment size or financial contribution as the launch approaches.
Latest News
View all news- Senate Pushes Crypto CLARITY Act Cloture Vote to September 15
- Treasury Details $2,500 Employer Contributions to Trump Accounts
- Local Opposition Becomes a Credit Risk for AI Data-Center Lenders
- Treasury’s $58 Billion 3-Year Auction Tests Demand as 30-Year Yield Nears 19-Year High
- Intel Upsizes Stock Offering to $20 Billion After 2026 Rally