Overview
Jalapeño is OpenAI's first self-developed AI inference chip, jointly released with Broadcom on June 24, 2026. Manufactured using TSMC's 3nm process, it went from design to tape-out in just 9 months—setting a new industry record. This marks the first time OpenAI has turned its 'self-developed chip' from rumor into reality among tech giants, signaling a partial shift away from absolute reliance on NVIDIA GPUs.
Core positioning: **Specializing in LLM inference** (not training), targeting the 'cost black hole' of large-scale API calls like ChatGPT. Broadcom CEO Hock Tan revealed during the launch: **In early lab tests, Jalapeño's inference cost is about 50% lower than mainstream GPUs, with performance comparable to NVIDIA Blackwell**—if these numbers hold in real-world deployment, the cost per ChatGPT call could be halved.
**Rhythm and Ecosystem**: Led by a former Google TPU veteran, co-designed with Broadcom, manufactured with Celestica's assistance, and taped out at TSMC, **initial deployment is planned for the end of 2026**, explicitly described as '**the first step in a multi-generational chip development plan**'. Starting in 2026, OpenAI will collaborate with partners like Microsoft to promote **gigawatt-scale data center deployments**. OpenAI is still evaluating whether to sell the chip externally or keep it for internal use only.
Significance: **Jalapeño's emergence marks the official transition of AI giants from the 'buying cards' era to the 'making chips' era**—following Google TPU, Amazon Trainium, and Meta MTIA, OpenAI becomes the latest player to join the self-developed chip camp, creating the first visible crack in NVIDIA's 'shovel-selling' business.
**Latest Updates**: As of July 2026, OpenAI's official website has released new products such as GPT-5.6 (frontier intelligence model), GPT-Live (real-time interaction product), and Health in ChatGPT (health domain features). These new features, together with the Jalapeño chip, form OpenAI's inference ecosystem, further reducing costs for services like ChatGPT and expanding application scenarios.
Key Features
- 9-Month Tape-Out: From design to tape-out in just 9 months, setting a new industry record; co-designed with Broadcom, manufactured using TSMC's 3nm process
- 50% Inference Cost Reduction: Early lab tests show inference cost about 50% lower than mainstream GPUs (disclosed by Broadcom CEO Hock Tan)
- Performance Comparable to Blackwell: Matches NVIDIA Blackwell in inference workloads; alongside custom ASICs like Google TPU and AWS Trainium
- Focused on LLM Inference: Optimized for large-scale API calls like ChatGPT, not a training chip—aimed at cost reduction rather than peak performance
- Starting Point for Multi-Generation Chip Plan: Jalapeño is the first step in OpenAI's multi-generational chip development plan, with continuous iterations to follow
- Gigawatt-Scale Data Centers: Starting in 2026, will collaborate with partners like Microsoft to promote gigawatt-scale data center deployments
- Led by TPU Veteran: Designed under the leadership of a former core member of Google's TPU team, ensuring solid engineering pedigree
Use Cases
- Large-scale inference workloads for OpenAI's own ChatGPT / API / Codex
- Cost reduction for OpenAI model series on Microsoft Azure
- Potential alternative to NVIDIA inference cards for enterprise large-scale inference workloads if sold externally in the future
- Benchmark for domestic competitors (e.g., Huawei Ascend, Cambricon)
- Key variable for AI infrastructure analysts/investors assessing 'NVIDIA's ceiling'
- Supporting real-time inference needs for new products like GPT-5.6 and GPT-Live, reducing latency and cost
Pros
- Industry record of 9-month tape-out speed
- Approximately 50% inference cost savings compared to mainstream GPUs (lab results)
- Performance comparable to NVIDIA Blackwell
- Clear starting point for OpenAI's self-developed chip journey with a multi-generation plan
- Solid engineering strength from Broadcom + TSMC + Google TPU team talent combination
- A substantive crack in NVIDIA's 'absolute monopoly'
- Synergy with new models like GPT-5.6 and GPT-Live to enhance overall inference efficiency
Pricing
**Not for external sale yet**: Currently only for OpenAI's internal use and deployment with deep partners like Microsoft. Whether to open external sales is still under evaluation. **Indirect benefits**: Users can expect reduced costs for ChatGPT/API usage, especially with new features like GPT-5.6 and GPT-Live further optimizing inference costs.
Summary
Jalapeño represents a critical step for OpenAI amid 'NVIDIA card shortages and soaring inference costs'—**9-month tape-out + 50% inference cost reduction + performance comparable to Blackwell**, moving AI giants from 'buying cards' to 'making chips'. Currently, it is more of a strategic signal than a directly purchasable product for users; however, for ChatGPT users and OpenAI API developers, inference costs and availability may significantly improve within the next year. For domestic chip manufacturers (e.g., Huawei Ascend, Cambricon), Jalapeño is a must-study benchmark. With the release of new products like GPT-5.6, GPT-Live, and Health in ChatGPT, Jalapeño's inference capabilities will support a broader range of real-time application scenarios.