Models
August 25, 2026

OpenAI announces first test results from its Jalapeño inference chip

OpenAI has shared the first performance results for Jalapeño, its custom AI inference chip, reporting higher throughput, lower latency, and greater performance per watt across several large language models.

OpenAI has published the first measured performance results for Jalapeño, its custom inference chip developed with Broadcom and Celestica.

Tests using GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T showed 1.5 to 1.9 times higher peak performance per watt and 1.7 to 3.6 times lower end-to-end latency than comparison systems.

On highly interactive workloads, performance improved by up to 4.1 times. OpenAI also used its own AI models to accelerate Jalapeño’s design and programming. The company plans to begin deploying the chip within its compute infrastructure by the end of 2026.

#
OpenAI

Read Our Content

See All Blogs
Gen AI

AI agent evaluation for self-improving agent loops

Sarankumar S

August 21, 2026
Read more
LLM Models

Claude watermarking and what it means for enterprise AI teams

Sarankumar S

August 20, 2026
Read more