2026-08-25 · ← News
OpenAI shows Jalapeño first: Custom chip radically cheapens model execution
OpenAI doesn't need Nvidia to run agents, only to train them
Jalapeño is an Application-Specific Integrated Circuit (ASIC) developed alongside Broadcom. While Nvidia's GPUs (like the B200) are built for massive parallel computation during training, Jalapeño is solely focused on inference (the phase where a model generates responses for users). And inference makes up the bulk of operational costs for AI platforms.
Early benchmarks on the InferenceX platform highlight a dramatic leap: compared to the best Nvidia-backed systems (GB200/GB300), Jalapeño delivers 50 to 90 % more work per watt. This holds true across a spectrum of models, from GPT-OSS 120B to DeepSeek R1 and Kimi K2.5 1T.
For users, latency (the time before a model starts spitting out a response) is even more crucial. Here, Jalapeño reduced end-to-end latency by 1.7× to 3.6× compared to equivalent Nvidia setups. Hardware VP Richard Ho claims they’ve managed to break the traditional tradeoff between throughput and low latency.
The end of Nvidia's monopoly in artificial intelligence server rooms
For the industry at large, this represents a critical inflection point. Until now, Nvidia held the only keys to the AI server room, dictating both price and availability. Their chips are excellent general-purpose engines, but for a single-purpose task like "generate tokens," they are unnecessarily expensive and power-hungry.
When you build a chip exclusively for inference, you can slash API costs. For developers building complex RAG pipelines or long-running agents on top of OpenAI's models, this should mean one thing: cheaper tokens and smoother responses. For Sam Altman, it’s a move to rein in the largest line item on the P&L statement.
Paper benchmarks against hard production reality
OpenAI is showcasing selected results on specific models. The true test of the hardware will come when it handles diverse production workloads under massive user demand. A secondary constraint is manufacturing capacity: the Broadcom design implies fabrication at TSMC, where OpenAI must wait in line behind Apple and Nvidia.
Richard Ho tempers expectations by noting that the chips will deploy in "small volumes" this year, with a larger ramp-up slated for 2027. Even then, OpenAI isn't cutting Nvidia loose entirely, still calling them "very good partners." Swapping hardware beneath a product as massive as ChatGPT is akin to open-heart surgery while running a marathon.
When will the API price tag drop
The crucial metric to watch is when the touted energy savings and higher throughput translate into lower API prices for developers. As long as OpenAI keeps token prices static, Jalapeño serves merely as a tool to expand their own margins rather than a market disruption.
Lilith's verdict
It is primarily a supply chain optimization. OpenAI doesn't want to pay a tax to Nvidia for every generated word, so they built their own toll booth.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗