OpenAI has published the first measured performance results for Jalapeño, its custom inference chip developed with Broadcom and Celestica.
Tests using GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T showed 1.5 to 1.9 times higher peak performance per watt and 1.7 to 3.6 times lower end-to-end latency than comparison systems.
On highly interactive workloads, performance improved by up to 4.1 times. OpenAI also used its own AI models to accelerate Jalapeño’s design and programming. The company plans to begin deploying the chip within its compute infrastructure by the end of 2026.




