IBM and Together AI have agreed on a $240 million deal to build an AI inference cluster powered by Nvidia hardware. Reuters reports the investment, citing Nvidia as a source. The commitment underscores that specialized compute infrastructure for productive AI use has become a strategic bottleneck – not just model training, but inference operations at scale require enormous capacity.
The essentials
- $240 million investment for joint inference cluster
- Partnership between IBM and Together AI using Nvidia hardware
- Focus on production AI workloads rather than model training alone
- Infrastructure designed to enable enterprise-scale deployment
Who does what?
According to the Reuters report, IBM and Together AI are collaborating to create a cluster that allows companies to run large language models and other AI systems in production – without requiring massive individual hardware investments. Nvidia supplies the graphics processors essential for such operations.
The emphasis on inference rather than training is noteworthy: while model training is a one-time, compute-intensive operation, inference runs continuously – every user query to an AI system consumes compute resources. To serve many users, proportionally scaled hardware is required.
Market implications
This announcement illustrates a broader trend: the infrastructure bottleneck is shifting from training to production deployment. Large models are trained; the challenge now is operating them economically. Companies wanting to handle this themselves face either massive capital expenditure or must partner with specialized providers.
For German mid-market companies, this is relevant: anyone wanting to deploy AI systems productively must decide whether building proprietary infrastructure makes sense or whether partnerships with established players like IBM or specialized infrastructure vendors offer better economics. The $240 million investment signals the scale at which such projects operate.
What this means for German enterprises
This investment sends a clear signal: specialized AI inference capacity is becoming scarce. Companies working with large models should clarify early how they'll operate inference workloads – whether via cloud providers, specialized infrastructure partners, or in-house development. The costs and technical complexity are substantial. For many organizations, a model like IBM and Together AI's – centralized infrastructure serving multiple customers – will likely prove more cost-effective than isolated solutions.
Sources
Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.




