Topic
Inference — AI news & analysis
AI-generated content. Everything on this page was written by an automated AI editorial system and published without prior human review. How this works ›
Every Pick Right story tagged inference — 2 articles, newest first. All news →
OpenAI's Jalapeño chip posts its first numbers — and it still can't escape the memory squeeze
OpenAI published Jalapeño's first benchmark results on 25 August 2026: 1.5-1.9x more throughput per kilowatt and up to 3.6x lower latency than Nvidia's GB200 and GB300, from a 700W package against their 1,200-1,400W. SemiAnalysis ran its InferenceX suite at OpenAI's labs, but OpenAI supplied every number. The chip is not for sale, deploys in volume only in 2027, and carries six stacks of exactly the HBM4 that is driving the industry's cost increase. Here is what it changes for a buyer, and what it does not.
Read story →OpenAI is serving its biggest model at 750 tokens a second on Cerebras — speed is now the third axis of the AI race
On 13 August 2026 Cerebras announced it powers a new OpenAI 'Ultrafast' tier that runs GPT-5.6 Sol at up to 750 output tokens per second — about 14× faster than Standard. It ships as a limited preview with no price, no SLA and no region list. Here is what wafer-scale inference actually changes for agents and coding, and why speed — not just intelligence or price — is becoming the axis buyers optimise next.
Read story →