As of August 27, 2026, OpenAI said its Broadcom-built Jalapeño chip delivered 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower end-to-end latency than the best recorded Nvidia GB200 or GB300 results on three large language models.
The Verge reported the figures on August 25, 2026, from an OpenAI blog post and a briefing by hardware vice president Richard Ho. OpenAI used the InferenceX benchmark platform and compared Jalapeño with the best Nvidia GB200 or GB300 scores then on record, across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. Patentlyze said Ho presented the same ranges on a Hot Chips slide.
Those numbers are OpenAI’s tests against previously recorded Nvidia results, not a third-party bake-off of the same systems in one lab. Wccftech, citing SemiAnalysis, repeated the InferenceX ranges and noted that Jalapeño uses HBM4 while the GB200 and GB300 systems in the comparison use HBM3E, which it said slightly favors the OpenAI chip. Wccftech also quoted a SemiAnalysis post saying Jalapeño beat Vera Rubin July results on output throughput per megawatt. Patentlyze cited SemiAnalysis’s breakdown of the Hot Chips talk as describing Jalapeño’s main compute engine as a weight-stationary systolic array, in the same broad family as Google’s TPU.
OpenAI publicly introduced Jalapeño on June 24, 2026 as an inference ASIC built with Broadcom to run large language models rather than train them. The Verge and a TechCrunch report carried by Yahoo Tech described it as OpenAI’s first custom processor. SiliconANGLE reported the same unveiling of a custom Jalapeño chip for LLM inference. Forbes called it OpenAI’s first proprietary chip; TechSpot called it the company’s first custom AI chip.
At that June announcement, OpenAI said it was still measuring final performance and that early testing showed substantially better performance per watt than then-current state of the art, according to The Verge. The Verge also reported that Broadcom chief executive Hock Tan told Reuters the chip matched Nvidia Blackwell and Google TPU performance. TechSpot wrote the next day that engineering samples were running production-class workloads after a roughly nine-month design-to-tape-out cycle, and that neither OpenAI nor Broadcom had released full public specifications or independent benchmarks at the time.
Ho said OpenAI planned to deploy Jalapeño in small volumes by the end of 2026 and ramp in 2027, without giving unit counts, and did not expect the chip to replace the company’s full compute mix, which still includes Nvidia, The Verge reported.
A TechCrunch report carried by Yahoo Tech said the OpenAI-Broadcom chip partnership was officially announced in October. The Verge said the June 2026 unveil came about nine months after that team-up.