Tech and AI
The chip bought for 20 billion dollars last year is now in full production, and the first rack goes to a cloud most people have never heard of
By Staff Writer | 25 August 2026

Nvidia said on 24 August that Groq 3 LPX has entered full production, with a benchmark of 3,400 output tokens a second on a 100,000 token context. The first customer is a specialist artificial intelligence cloud rather than one of the three large platforms.
Nvidia announced at the Hot Chips conference on 24 August that Groq 3 LPX, its inference accelerator, is now in full production. The part extends the company's Vera Rubin rack platform and is aimed at one job only: generating output quickly enough that an agent can complete hundreds of steps while somebody waits.
The company reports a record of 3,400 output tokens a second in independent benchmarking, running an open source model of 31 billion parameters with a 100,000 token context, which it describes as the fastest performance recorded for that model and four times the responsiveness of the nearest alternative platform. The purchase of the underlying processor assets, made in December, is reported at 20 billion dollars, and on that account it is the largest acquisition the company has made.
Generation is the phase of inference that determines how responsive an AI system actually is, and that's exactly what NVIDIA Groq 3 LPX is built to accelerate. As the first AI cloud bringing it to production via Nebius Token Factory, we're making sure every step of an agent's loop feels instant, through the same API developers are already using, with no migration to a new stack.
Danila Shtan, Chief Technology Officer, Nebius
The first adopter is the part that matters
Nebius, a specialist artificial intelligence cloud, is named as the first to bring the part into production, ahead of the large platform operators. That order of adoption tells you something about where the demand for very fast output is: not in general purpose cloud, but in the narrow set of workloads where a human being is waiting for an agent to finish. Nvidia says a second specialist cloud plans to be among the earliest adopters after that.
Separately on the same day the company said a space and artificial intelligence group has adopted its Vera central processors to carry the orchestration, tool use, code execution and data processing that sit behind agent work. Those are the parts of an agent's day that are not model inference at all, and they have been the quiet constraint on the whole approach.
The competitive axis has moved. It is no longer how large a model can be trained, it is how fast the answer comes back, because an agent that takes four hours to do an hour of work is an agent that gets switched off.
Why a construction business should care about token speed
Because it decides the shape of the buildings this industry is being asked to put up. A rack designed for very fast generation is a rack with a different power draw, a different cooling arrangement and a different ratio of compute to network than the training halls that have driven data centre design for three years. Operators specifying a scheme now are specifying it for a workload mix that changed this week.
There is a second reason, closer to the desk. The agent tools being sold into contract administration, programme analysis and document review are all limited by exactly the number this announcement moves. A tool that reviews a hundred thousand words of correspondence in ninety seconds gets used on every claim. The same tool at forty minutes gets used on none of them.
The figures here are the maker's own, taken from its own release, and the benchmark is one model on one configuration. What is not in doubt is the direction of the spend, or that the racks are said to be online before the end of the year.