CoreWeave Ships Nvidia’s Newest Rack, and a Coding Agent Tries It First

  • AI
  • October 1, 2026
  • 0 Comments

The first real customer for Nvidia’s newest supercomputer is a piece of software that writes code. When CoreWeave began delivering the Vera Rubin NVL72 system to customers this week, the initial workloads on the hardware were not demonstration benchmarks. They were the everyday jobs of Devin, the autonomous software engineer built by the AI lab Cognition, running on the machine as a paying production tenant.

CoreWeave said on September 30 that it had begun production-scale delivery of the Nvidia Vera Rubin NVL72 system, with Cognition as the first customer anywhere running real workloads on the platform. The announcement came during Fully Connected, the cloud provider’s AI conference in San Francisco, where the company framed the launch as more than a hardware handoff: the same operating model and tooling customers already use on their existing GB200 and GB300 NVL72 fleets carries over to the new racks.

The numbers Cognition reported are the kind that sell servers. In benchmarks run on CoreWeave’s cloud, the company measured a 4.8-times increase in total token throughput for its SWE-2 inference workloads on Vera Rubin NVL72 compared with a GB200 NVL72 baseline, and a 3.8-times lift in output token throughput for reinforcement-learning jobs. For Devin, an agent that reads a codebase, reasons and executes in a continuous loop, the gain compounds because every step waits on the last one.

“Agentic coding is an unforgiving workload that requires long contexts, high concurrency and rapid reasoning,” said Silas Alberti, senior vice president of research and a member of Cognition’s founding team. The result, he said, is more concurrent Devin sessions per GPU, faster research loops and a lower cost per session, with no loss in generation speed.

Cognition did not stumble into the front of the line. The lab scaled from bridge capacity to thousands of GPUs for training and inference in less than nine months, all on CoreWeave. It stood up its Vera Rubin NVL72 cluster in early September, working alongside CoreWeave’s engineers, and ran what the companies describe as the first customer-executed inference benchmark on the system, tested against a GB200 cluster as the control.

CoreWeave is also widening the door behind that first customer. The company said it had begun offering bare-metal instances running on Nvidia Vera CPUs, and that it was expanding its partner network to take on inference and storage demand that the new systems are expected to pull in. The message is that Vera Rubin is available, not merely announced, a distinction that matters in a market where next-generation hardware has a habit of living on roadmaps longer than it lives in racks.

The systems have already moved past the pilot stage in physical terms. CoreWeave said hundreds of Nvidia Rubin GPUs are deployed across multiple regions, with customer workloads onboarding in production rather than in test. The rack handover was co-designed with Nvidia, an arrangement the companies say let Cognition stand up a working cluster in days instead of the months a first-generation deployment once took.

The backdrop is a scramble over scarce compute. AI labs are signing long-term contracts for the fastest chips before they leave the factory, and cloud providers have spent the past two years fighting over who can stand up the newest generation first. CoreWeave, which went public as a Nasdaq-listed company, has built its pitch on being the specialty cloud that runs AI workloads closer to the metal than the big hyperscalers.

Nvidia’s role in the announcement is the quiet part. The Vera Rubin platform, with its Spectrum-X networking, is the company’s bet on the agentic-AI workloads that now dominate spending, software that reasons for minutes instead of predicting the next token. Getting the first production customer on the platform gives Nvidia a reference point it can show every other lab deciding whether to commit to the new generation.

The counterweight is that one early customer is still one customer. The 4.8-times figure is a vendor- and customer-reported benchmark run against a single baseline, not an independent audit, and early performance numbers on new silicon have a way of settling as workloads diversify. Cognition’s confidence is a signal, not yet a market.

What the deal demonstrates more than speed is a new kind of flagship buyer. The first customer for a machine this expensive is no longer a research institution or a national lab. It is an applied-AI company whose product is an agent that writes code, spending on compute because its own economics demand it. The customers deciding the future of Nvidia’s hardware are, increasingly, software that is itself deciding what to build next.

CoreWeave did not say how many systems it had delivered or which customers would come after Cognition. The answer to that question will determine whether the launch is remembered as the first delivery of a new generation, or as the moment the queue for it started to form.

Related Posts

  • October 1, 2026
  • 16 views
A Bad Prompt Exposes 95,000 Customer Emails at Bee Cheng Hiang

Bee Cheng Hiang, the Singapore company that has sold bak kwa, or barbecued pork jerky, for close to a century, decided in April to try something new. An employee asked…

  • October 1, 2026
  • 17 views
Google’s New Gemini Model Opens to Cyber Defenders First

Google introduced a new flagship artificial-intelligence model on Tuesday, but most people cannot use it yet. Gemini 4 Argon, the company’s first new frontier model since Gemini 3 last November,…