Models
August 25, 2026

OpenAI announces first test results from its Jalapeño inference chip

OpenAI has shared the first performance results for Jalapeño, its custom AI inference chip, reporting higher throughput, lower latency, and greater performance per watt across several large language models.

OpenAI has published the first measured performance results for Jalapeño, its custom inference chip developed with Broadcom and Celestica.

Tests using GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T showed 1.5 to 1.9 times higher peak performance per watt and 1.7 to 3.6 times lower end-to-end latency than comparison systems.

On highly interactive workloads, performance improved by up to 4.1 times. OpenAI also used its own AI models to accelerate Jalapeño’s design and programming. The company plans to begin deploying the chip within its compute infrastructure by the end of 2026.

#
OpenAI

Read Our Content

See All Blogs
LLM Models

LLM testing of Muse Spark 1.3

Sarankumar S

September 11, 2026
Read more
LLM Models

Harness engineering for AI agents: The missing layer for production deployment

Sarankumar S

September 11, 2026
Read more