Skip to main content
  1. Home
  2. Computing
  3. News

OpenAI says its Jalapeno AI chip delivers faster responses than rivals like Nvidia

OpenAI's first AI chip focuses on speed, efficiency, and scaling AI services

Add as a preferred source on Google
openai-jalapeno-ai-chip
OpenAI

Every time you ask ChatGPT a question, a chip somewhere is racing to send back an answer. OpenAI just shared fresh benchmark results claiming its custom Jalapeño chip, co-developed with Broadcom, does that job significantly faster and more efficiently than rival chips in the market right now.

We’ve designed and built our first AI chip: Jalapeño.

Designed from the ground up by OpenAI and brought to production with @Broadcom, Jalapeño is purpose-built for the LLM workloads powering ChatGPT, Codex, the API, and future agentic products.

Chips are foundational to the AI… pic.twitter.com/mHU7DaMMTi

— OpenAI (@OpenAI) June 24, 2026

How much faster is Jalapeno, exactly?

OpenAI tested Jalapeno using InferenceX, an independent benchmark from SemiAnalysis that measures how well a chip handles inference, which is the process of running a trained AI model to generate responses.

Recommended Videos

Across three different AI models, Jalapeno delivered 1.5 to 1.9 times more AI work per watt of power, along with 1.7 to 3.6 times lower latency, meaning noticeably faster response times. For highly interactive tasks like AI agents, that advantage climbed even higher, reaching up to 4.1 times better performance.

Most chips force a tradeoff between speed and efficiency, but OpenAI says Jalapeno manages both simultaneously by minimizing how much data needs to move between different parts of the system during processing.

Why did OpenAI decide to build its own chip?

OpenAI has relied heavily on Nvidia GPUs since its inception, and Jalapeno marks its first real step toward designing hardware in-house. Interestingly, OpenAI’s own AI models helped design portions of the chip itself, and that same approach let a small team bring three additional AI models up to full performance within just two months.

Despite the strong results, OpenAI isn’t abandoning Nvidia anytime soon, since Jalapeno only handles inference, not the separate, compute-heavy process of training new models. The company plans to deploy the Jalapeno chip in small volumes by the end of 2026, scaling up meaningfully in 2027, with a second and third generation chip already taking shape behind the scenes.

OpenAI joins a growing list of AI companies building their own chips. Google has its TPU line, Amazon is pushing into custom silicon, and Anthropic recently confirmed similar plans of its own. Interestingly, Samsung has reportedly used Anthropic’s Claude to speed up its own chip design work, suggesting AI assisting hardware development is quickly becoming an industry norm, not just an OpenAI experiment.

Manisha Priyadarshini
Manisha Priyadarshini is a tech and entertainment writer with over nine years of editorial experience.
Windows 11 is about to turn on a setting that could hurt gaming performance
Microsoft will start enabling Memory Integrity on more Windows 11 PCs in October
A gaming PC with RGB synced lights running Apex Legends.

Microsoft is preparing to enable Memory Integrity on more Windows 11 PCs starting in October. The security feature is meant to protect systems from malicious code, but there is one reason gamers may want to keep an eye on it.

Microsoft has previously acknowledged that Memory Integrity can affect gaming performance on some Windows 11 systems. Back in 2022, the company even published instructions explaining how gamers could temporarily disable Memory Integrity and Virtual Machine Platform if they were causing problems.

Read more
AI chatbots will agree and misinform if you just put a little pressure, warns research
ChatGPT 3.5 proved the most vulnerable when false claims were repeated over and over
Claude AI on an iPhone.

AI chatbots have a well-documented habit of hallucinating information and sometimes agreeing with users even when they are wrong. A new study suggests that simply refusing to take no for an answer can make the problem worse.

Researchers from the University of Arizona tested seven AI models, including GPT-3.5, GPT-4o, GPT-4o-mini, Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama 3 70B, and DeepSeek-R1. Instead of judging them from a single response, researchers kept conversations going while repeatedly feeding the models information they knew was false.

Read more
Google brings its best AI music model Lyria 3.5 to the Gemini app
Developers get full API access through Google AI Studio.
Google-gemini-lyria-3.5

Google has added Lyria 3.5, its most advanced music generation model yet, to the Gemini app. Previously available through Google's AI filmmaking tool Flow, the model is now rolling out to all Gemini users, making it easier to generate polished songs, instrumentals, and soundtracks from simple text prompts or even photos.

Lyria 3.5 makes AI-generated music sound more natural

Read more