OpenAI Details Jalapeno Inference Accelerator

OpenAI presented details of its Jalapeno AI accelerator at the Hot Chips conference on August 25, according to The Register. The publication reports that a 128-chip configuration delivers 1.7 exaFLOPS and 27 TB of HBM, with OpenAI projecting higher inference throughput and lower latency than contemporary GPU systems. Volume production is reported for 2027.
OpenAI presented its Jalapeno custom AI accelerator at the Hot Chips semiconductor conference on August 25, offering its most detailed public account yet of silicon developed with Broadcom, according to The Register. The publication reports that a 128-chip Jalapeno system provides 1.7 exaFLOPS of compute and 27 TB of HBM memory.
The Register describes Jalapeno as OpenAI's first chip in a broader custom-silicon effort and reports that the accelerator is designed principally for inference rather than general-purpose training. According to the publication, OpenAI expects deployments to begin later in 2026, with volume production reported for 2027.
Throughput and latency claims
The Register reports that OpenAI shared early results using SemiAnalysis' InferenceX benchmark before the Hot Chips presentation. Across GPT-OSS-120B, DeepSeek R1, and Kimi K2.5, the reported results showed 1.5x to 1.9x higher peak-throughput AI work and 1.7x to 3.6x lower end-to-end latency than competing systems.
The publication characterizes InferenceX as presumably an unofficial test, so the results are not equivalent to independently standardized performance validation. Workload mix, batch size, precision format, model-serving software, and latency targets can materially affect accelerator comparisons.
A specialized inference design
The Register contrasts Jalapeno with AMD's MI455X and Nvidia's Rubin GPUs, which it describes as designs intended for both training and inference. It reports that OpenAI continues to require training compute and is likely to use AMD and Nvidia hardware before moving applicable inference workloads to in-house silicon.
For ML infrastructure teams, the disclosed specifications underline a familiar systems constraint: large-model inference is often limited by memory capacity and bandwidth as much as raw arithmetic throughput. Custom ASICs aimed at a narrower serving workload can improve token economics and latency, but their practical value depends on compiler maturity, model compatibility, networking, reliability, and availability at scale. Independent benchmarks and production deployment data remain open questions for Jalapeno.
Key Points
- 1OpenAI disclosed a 128-chip Jalapeno configuration with 1.7 exaFLOPS and 27 TB of HBM for LLM inference workloads.
- 2The Register reports early InferenceX results showing higher throughput and lower latency, though independent standardized validation has not been established.
- 3Specialized inference ASICs commonly trade broad programmability for efficiency, making software compatibility and serving-system integration central deployment considerations.
Scoring Rationale
OpenAI's disclosed custom accelerator specifications are notable for teams tracking the future cost, latency, and supply dynamics of large-model inference. The reported performance figures are early and not independently standardized, but the hardware effort is material given OpenAI's scale and its relationship to GPU-centric infrastructure.
Sources
Public references used for this report.
Practice with real Ad Tech data
90 SQL & Python problems · 15 industry datasets
250 free problems · No credit card
See all Ad Tech problems