California’s Cerebras Systems announced on Monday plans to supply AI chips and hardware capable of consuming about 100 megawatts to startup Gimlet Labs to support its cloud computing operations.
The deal reflects growing interest from AI companies in acquiring hardware capable of fast AI computations, known as inference, where a chatbot like Anthropic’s Claude produces responses to user queries.
Zain Asgar, co-founder and CEO of Gimlet Labs told investors, “Inference speed matters. It determines how productive AI can be. Fast inference creates magical user experiences and opens new markets. By combining Gimlet’s multi-silicon software with the Cerebras Wafer Scale Engine, we can run each phase of inference on the hardware best suited to it and plan to deliver up to 3,000 tokens per second at production scale.”
Sean Lie, co-founder and CTO at Cerebras, stated, “Combining the fastest tokens from Cerebras with the highest throughput GPUs delivers the best datacenter economics for everyone. Everyone wants more high value tokens. Cerebras delivers the fastest AI inference in the world, and GPUs deliver high throughput. By making Cerebras a native part of its inference cloud, Gimlet will bring our industry leading speed and intelligent AI to more developers at production scale. We’re excited to build with Gimlet as a launch partner for CS-4, giving customers a direct path to our latest technology.”
The collaboration builds on joint customer engagements underway since last year and an integrated solution already serving tokens in private deployments.
Gimlet Labs will expand the collaboration with Cerebras to integrate software, infrastructure design, APIs, developer tooling, optimization, validation and production operations to make ultrafast inference broadly available through Gimlet Cloud.
By CEO NA Editorial Staff










