Overview
Jalapeño is OpenAI's **first self-developed AI inference chip**, jointly announced with Broadcom on June 24, 2026. Manufactured using TSMC's 3nm process, it went from design to tape-out in just **9 months**—setting a new industry record. This marks the first time OpenAI has turned the rumor of a self-developed chip into reality among major tech companies, signaling a partial shift away from absolute reliance on NVIDIA GPUs.
Core positioning: **Specialized for LLM inference** (not training), targeting the "cost black hole" of large-scale API calls like ChatGPT. Broadcom CEO Hock Tan revealed at the launch: **In early lab tests, Jalapeño's inference cost is about 50% lower than mainstream GPUs, with performance comparable to NVIDIA Blackwell**—if this holds true in real-world deployment, it could halve the cost per ChatGPT call for OpenAI.
**Timeline and ecosystem**: Led by a former Google TPU veteran, co-designed with Broadcom, manufactured with assistance from Celestica, and taped out at TSMC. **Initial deployment is planned for late 2026**, and it is explicitly described as "**the first step in a multi-generation chip development plan**." Starting in 2026, OpenAI will collaborate with partners like Microsoft to drive **gigawatt-scale data center deployments**. OpenAI is still evaluating whether to sell the chip externally or keep it for internal use only.
Significance: **The emergence of Jalapeño marks the official transition of major AI companies from the 'buying chips' era to the 'making chips' era**—following Google TPU, Amazon Trainium, and Meta MTIA, OpenAI becomes another player joining the self-development camp, creating the first significant crack in NVIDIA's "shovel-selling" business.
Key Features
- 9-month tape-out: From design to tape-out in just 9 months, setting a new industry record; co-designed with Broadcom, manufactured on TSMC 3nm process
- 50% inference cost reduction: Early lab tests show inference cost about 50% lower than mainstream GPUs (disclosed by Broadcom CEO Hock Tan)
- Performance comparable to Blackwell: Comparable to NVIDIA Blackwell on inference workloads; competes with custom ASICs like Google TPU and AWS Trainium
- Specialized for LLM inference: Optimized for large-scale API calls like ChatGPT, not a training chip—focuses on cost reduction rather than peak performance
- Starting point for multi-generation chip plan: Jalapeño is the first step in OpenAI's multi-generation chip development plan, with continuous iterations to follow
- Gigawatt-scale data centers: Starting in 2026, will collaborate with partners like Microsoft to drive gigawatt-scale data center deployments
- Led by TPU veteran: Design led by a core figure from the original Google TPU team, ensuring strong engineering pedigree
Use Cases
- Large-scale inference workloads for OpenAI's own ChatGPT / API / Codex
- Cost reduction for OpenAI model deployment on Microsoft Azure
- Potential future external sales as an alternative to NVIDIA inference cards for enterprise large-scale inference workloads
- Benchmark for domestic competitors (e.g., Huawei Ascend, Cambricon)
- Key variable for AI infrastructure analysts / investors assessing 'NVIDIA's ceiling'
Pros
- Industry-record 9-month tape-out speed
- Inference cost about 50% lower than mainstream GPUs (lab results)
- Performance comparable to NVIDIA Blackwell
- Clear starting point for OpenAI's self-developed chip journey with multi-generation plan
- Strong engineering team combining Broadcom, TSMC, and Google TPU talent
- A substantive crack in NVIDIA's 'absolute monopoly'
Pricing
**Not for external sale yet**: Currently only for OpenAI's internal use and deployment by deep partners like Microsoft. Whether to open sales externally is still under evaluation. **Indirect benefit**: Expected reduction in user costs for ChatGPT / API usage.
Summary
Jalapeño represents a critical step for OpenAI amid 'NVIDIA GPU shortages and soaring inference costs'—**9-month tape-out + 50% inference cost reduction + performance comparable to Blackwell**—moving AI giants from 'buying chips' to 'making chips'. Currently more of a strategic signal than a directly purchasable product; however, for ChatGPT users and OpenAI API developers, inference costs and availability may significantly improve within the next year. For domestic chip makers (Huawei Ascend, Cambricon, etc.), Jalapeño is a must-study benchmark.
Version History
- Jalapeño (first self-developed inference chip) (2026-06-24): Jointly announced with Broadcom, TSMC 3nm process, 9-month tape-out; inference cost about 50% lower than mainstream GPUs; performance comparable to NVIDIA Blackwell; initial deployment in late 2026