> ## Content Index
> Fetch the complete content index at: https://www.tao.media/llms.txt
> Use this file to discover other available public pages before exploring further.

# OpenAI’s First Custom AI Chip Beats Nvidia Blackwell in Early Inference Tests
- URL: https://www.tao.media/openais-first-custom-ai-chip-beats-nvidia-blackwell-in-early-inference-tests/
- Published: 2026-08-26T16:44:58.000Z
- Updated: 2026-08-26T16:44:58.000Z
- Description: OpenAI’s first custom inference chip is already outperforming Nvidia, AMD and Google hardware across several AI workloads, despite being a first-generation design.
- Author: Bart Hillerich
- Tags: OpenAI, News

[OpenAI’s](https://www.tao.media/tag/openai/) first custom AI chip is already outperforming some of Nvidia’s most advanced hardware across key inference benchmarks.

The chip, called Jalapeño, delivered better performance per watt than Nvidia Blackwell systems [across nearly every scenario tested by semiconductor research firm SemiAnalysis](https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia?ref=tao.media), which benchmarked the hardware directly at OpenAI’s labs using its InferenceX suite. SemiAnalysis said Jalapeño also beat every Nvidia, AMD, and Google chip it has tested across multiple major open-weight models.

OpenAI separately reported that Jalapeño delivered 1.5 to 1.9 times more AI work per watt than the Nvidia GB200 and GB300 systems used for comparison across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T. It also achieved between 1.7 and 3.6 times lower end-to-end latency across the three models.

Jalapeño is OpenAI’s first internally designed AI accelerator, putting the company into direct competition with the hardware it currently relies on to serve its models.

[Minos HelixForge Featured in OpenAI Field Report on Agentic AI for Scientific ComputingThe Bittensor genomics project’s GPU-native synthetic genome engine was highlighted as one of the most ambitious case studies in a new report on AI-assisted scientific software.![](https://storage.ghost.io/c/78/0b/780ba906-b1a7-4bf0-873c-bdd5c32e5331/content/images/icon/Group-1321319358-b9f51d93-ba84-4a1e-8102-05f981a088cb.png)Intelligence | Bittensor News, Insights, StoriesBart Hillerich![](https://storage.ghost.io/c/78/0b/780ba906-b1a7-4bf0-873c-bdd5c32e5331/content/images/thumbnail/1f569311-afc9-4445-9dac-b641891348b6-175d76e6-59eb-4d22-a430-240e8f011fa4.png)](https://www.tao.media/minos-co-authors-scientific-computing-paper-with-openai/)

## Jalapeño Is Built Specifically for AI Inference

[OpenAI first unveiled Jalapeño in June](https://openai.com/index/openai-broadcom-jalapeno-inference-chip/?ref=tao.media) as part of a partnership with Broadcom. Unlike Nvidia GPUs that are designed to support both AI training and inference alongside other workloads, Jalapeño is an application-specific integrated circuit, or ASIC, built from scratch around large language model inference.

![](https://storage.ghost.io/c/78/0b/780ba906-b1a7-4bf0-873c-bdd5c32e5331/content/images/2026/08/image-12.png)

OpenAI CEO Sam Altman & Broadcom President and CEO Hock Tan | [OpenAI](https://openai.com/index/openai-broadcom-jalapeno-inference-chip/?ref=tao.media)

Inference is the process of actually running a trained AI model to generate an answer, execute an agent task or serve a user request. As products such as ChatGPT and Codex scale, inference represents an increasingly significant portion of the compute required to operate them.

But Jalapeño does not appear to be narrowly optimized only for OpenAI models.

SemiAnalysis tested the chip on several open-weight models and described it as a generalized LLM inference accelerator capable of performing across different workloads and operating points. OpenAI similarly says the architecture was designed for current and future LLMs across the industry.

On Kimi K2.5, SemiAnalysis found Jalapeño could approach 700 tokens per second per user, while producing more than nine times the throughput of the next-best tested chip at an interactivity level of 100 tokens per second per user. On GPT-OSS, its throughput per megawatt was nearly twice the highest-throughput result SemiAnalysis recorded for Nvidia’s GB200.

![](https://storage.ghost.io/c/78/0b/780ba906-b1a7-4bf0-873c-bdd5c32e5331/content/images/2026/08/image-13.png)

That combination of throughput and latency is important because AI infrastructure typically faces a tradeoff between serving a large number of users efficiently and generating responses quickly for each individual user. 

OpenAI says Jalapeño was designed to perform well on both simultaneously.

## A Potential Challenge to Nvidia’s CUDA Advantage

The results could have implications beyond chip-level performance.

One of Nvidia’s largest competitive advantages is CUDA, the software ecosystem developers use to build and optimize AI workloads for Nvidia hardware. Competing chips have historically struggled not only because of hardware performance, but because Nvidia has spent years building the software required to make its GPUs usable at scale.

SemiAnalysis argues that the speed at which OpenAI has brought Jalapeño’s software stack online could begin to challenge that advantage.

> "The CUDA moat is potentially dead given how fast OpenAI can bring up new models on their silicon."

Jalapeño taped out in late 2025, and SemiAnalysis says OpenAI had only been working with actual silicon for roughly three months before producing the current benchmark results. Performance has also been improving quickly, with throughput at some operating points more than doubling over a period of less than two weeks during testing.

OpenAI has another unusual advantage in developing that software: its own AI models.

The company says earlier OpenAI models helped engineers design and bring up Jalapeño, while newer models are now being used to optimize and program the chip itself. OpenAI previously said AI assistance contributed to an unusually fast development cycle for the processor.

> "They \[OpenAI\] claim that AI assistance in chip design delivered an 8% reduction in SIMD area and a 10% reduction in matrix-engine area during design. While they did not clarify the exact process/voltage/temperature (PVT) conditions, they also mentioned the AI-assisted blocks improved timing and power over the initial blocks."

## Nvidia Isn’t Being Replaced Yet

The results do not mean Jalapeño has comprehensively surpassed Nvidia’s hardware.

SemiAnalysis notes that its current benchmarks use 8,000-token input and 1,000-token output workloads, which are significantly easier to optimize than long-context, multi-turn agent workloads. Jalapeño has not yet been tested on SemiAnalysis’ more demanding AgentX benchmark.

The comparison with [Nvidia’s upcoming Vera Rubin platform](https://www.nvidia.com/en-us/data-center/technologies/rubin/?ref=tao.media) is also more nuanced. SemiAnalysis found Jalapeño and Rubin roughly comparable in output tokens per dollar in current testing, while noting that Rubin benefited from speculative decoding and Jalapeño did not. The firm expects Jalapeño’s economics to improve further as similar software optimizations are added.

OpenAI also has no plans to stop buying Nvidia hardware. Hardware chief Richard Ho said the company’s broader compute strategy will continue to involve partners such as Nvidia, while OpenAI develops second- and third-generation versions of its own silicon.

Jalapeño is expected to begin deployment in small volumes by the end of 2026, with production ramping further in 2027\. OpenAI ultimately plans to deploy its custom compute platform at gigawatt scale across multiple generations.