Product

AWS taps Cerebras chips to offer faster, premium AI inference in the cloud

What's the deal? Amazon Web ServicesDealroom has a profile for this one. Try Dealroom → has struck a multiyear partnership with AI chip startup Cerebras Systems to offer a new inference computing service. AWS will combine Cerebras's Wafer-Scale Engine processors with its own Trainium 3 chips in its data centers, with the service launching in the second half of 2026. Financial terms were not disclosed.

Why now? The AI industry is rapidly shifting from model training toward inference — the process that lets AI models respond to user queries. GPUs, long the workhorses of AI, are increasingly seen as too slow and costly for inference at scale. Cerebras claims its chips can handle inference decode tasks up to 25 times faster than Nvidia's GPUs.

Cerebras has been on a roll. In January 2026, OpenAI signed a deal worth more than $10 billion to use its chips. In February, the startup raised $1 billion in a new funding round, lifting its total fundraising to $2.6 billion and its valuation to roughly $23 billion.

What could go wrong? The combined Cerebras-Trainium service will be priced as a premium offering — meaning it will only appeal to customers for whom speed justifies the cost. AWS will continue offering cheaper Trainium-only services for less time-sensitive workloads. Whether enough customers will pay a premium for faster inference remains to be seen.

Cerebras has also had a bumpy road to the public markets. It filed for an IPO in September 2024 but withdrew about a year later. A fresh IPO attempt is reportedly in the works, and the AWS deal will help its case — but execution risk remains.

The signal: AWS is the first hyperscaler to commit to using Cerebras chips, a significant validation for the startup. The deal reflects a broader industry push to diversify beyond Nvidia, which faces mounting pressure from custom chip designers. Nvidia itself signaled the shift in December 2025, signing a $20 billion licensing deal with rival inference chip startup Groq.

The race is no longer just about who can train the biggest models — it's about who can run them fastest and cheapest. Cerebras, with its unusual wafer-scale design, is betting it can set that standard.

Sources:
Cerebras Systems
Business Wire
The Wall Street Journal
Reuters
Bloomberg

Image credit:
Cerebras Systems

J.V.

More top stories