Tech and AI
The first published numbers for a model maker's own chip are stated per kilowatt rather than per chip, and the choice of unit is the argument
By Staff Writer | 26 August 2026

OpenAI published benchmark results for Jalapeno on 25 August, claiming 1.5 to 1.9 times more work per watt and 1.7 to 3.6 times lower latency than the systems it was measured against. The part is rated at 700 watts against comparators rated at 1,200 and 1,400.
The company set out the first measured results for its own inference chip on 25 August, alongside a presentation at a processor conference the same day. The part was tested on a public third party benchmark across three openly available models, and the company reports that it delivered between 1.5 and 1.9 times more work per watt at peak throughput and between 1.7 and 3.6 times lower end to end latency than the systems it was compared against. For highly interactive workloads the claimed advantage widens to between 2.1 and 4.1 times.
The bottom line is that the results show a very, very significant performance advance over state of the art.
Richard Ho, head of hardware, OpenAI
The unit of measurement is doing a lot of the work
The company is open about this and states its reasoning plainly: performance is sometimes reported per chip, and it believes performance per unit of power is the more useful standard. It normalised the results using each accelerator's published chip power rating. Its own part is rated at 700 watts, although measured sustained draw stayed at or below 550 watts on the workloads tested, while the comparison systems are rated at 1,200 and 1,400 watts.
A ratio of roughly one to two on rated power is embedded in every number quoted before any silicon is compared. That is a legitimate way to measure a data centre, where the constraint is the electrical supply. It is not the same claim as a faster chip.
The published tables carry the underlying figures rather than only the multiples, which is more than most vendor benchmarking offers. On the largest public model tested, the company reports approximately 1.5 times higher peak performance per watt and 3.4 times lower end to end latency than the comparison system.
Design time, and a claim about who wrote the code
The company says its own models were used in the design, and that this took the part from initial design to tapeout in nine months by shortening the design, measurement and verification loops. It also reports that for selected attention and mixture of experts blocks on one of the tested models, machine generated implementations ran 1.5 to 1.8 times faster than the existing implementations written by human experts. It states that this applies to the selected blocks and not to the full model.
Anyone who has watched a design and build programme compress its own review cycles will recognise both halves of that. A nine month run to tapeout is the sort of number that gets quoted for years afterwards, and the qualification attached to the code figure is the sort that gets dropped.
What has not been demonstrated
Deployment is stated as beginning inside the company's own compute estate by the end of the year, in small volume, with wider deployment in 2027. Production qualification is still running, the software is described as maturing, and validation across more models is continuing. The company also says it will keep buying accelerators from its existing suppliers for both training and inference.
The comparison, in other words, is between a part that is not yet deployed and hardware that is on sale now, and by the time the first is in service the second will have been succeeded. Read the published tables and the power ratings together, treat the multiples as a statement about electrical efficiency rather than raw speed, and note that the benchmark was chosen by the party publishing the results.