OpenAI publishes the first measured results for Jalapeño, its own inference chip
1.5-1.9x more work per watt and 1.7-3.6x lower end-to-end latency, as measured by OpenAI itself.
On August 25, 2026, OpenAI published measured results for Jalapeño, its first custom inference chip. Testing on InferenceX, a public benchmark from SemiAnalysis, across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, OpenAI reports 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads it reports 2.1 to 4.1 times higher performance. On Kimi K2.5 1T, the largest public model tested, it reports roughly 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency. The chip is rated at 700 watts, with measured sustained power at or below 550 watts on the workloads tested. All figures are OpenAI's own measurements. The comparison systems are described only as "leading commercially available AI systems" and are not named. Results were normalised using each accelerator's published chip power rating.
Downward pressure on inference cost is now arriving from a route that does not run through GPU procurement. But this is vendor-run measurement rather than third-party verification, so treat it as a signal of intent until it shows up in pricing.
