NewsAI HardwareInferenceAccelerator

Cerebras CS-4: Double Performance on Same Chip

AI accelerator maker Cerebras has unveiled its new CS-4, claiming it's the fastest inference system on the market. The trick: same 5nm chip, but significantly higher clock speed.

4,400 tokens per second

Cerebras CS-4: Double Performance on Same Chip

Cerebras has unveiled its new AI accelerator CS-4. According to CEO Andrew Feldman, it's the fastest system in the industry. The system doubles the performance of its predecessor CS-3, even though the underlying chip remains identical – Cerebras increases the clock rate through more power and better cooling.

The essentials

  • The CS-4 is a rack-scale solution: A complete server cabinet with compute units, power, and cooling technology for data centers
  • Performance doubled: Up to 4,400 tokens per second per user – according to Cerebras, up to 30 times faster than Nvidia GPU solutions
  • Three instead of two wafers per rack; the 5nm chip WSE-3 remains unchanged
  • Memory constant: 44 GB per wafer, same as the predecessor

Modular design for faster deployment

Cerebras is introducing a new "backpack" design with the CS-4 to accelerate setup. At the same time, the company is pursuing disaggregated inference – collaborating with other vendors like AMD and AWS Trainium. This means: Not everything needs to run on Cerebras hardware; components can be modularly combined with other systems.

Memory capacity per wafer remains at 44 GB – no improvement here. This is a point that SemiAnalysis analysts view critically: progress on the networking side appears rather limited.

Who's already using it?

Cerebras hardware is already in use. OpenAI, for instance, uses the systems for Codex Spark. This shows: The hardware isn't just an announcement but is already being deployed productively by established AI providers.

Further technical details on the CS-4 are expected at the Hot Chips conference.

What this means for you

For infrastructure decision-makers in German mid-market companies, the CS-4 is interesting if you need high-performance AI inference – that is, fast processing of requests in production systems. The promised 30-fold speed advantage over GPU solutions could be relevant at high throughput. However, keep in mind: Cerebras specializes in inference, not training. And the market remains dominated by Nvidia – whether an alternative solution is worth the effort depends on your specific requirements. But the announcement shows: there's movement in the market beyond the big GPU players.

Sources

Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.

Share
← All articles

All analyses are based on i6eal's own measurements or on clearly labelled sources. Figures are snapshots and may change; corrections are disclosed transparently.