Cognition Runs First Production Workloads on Vera Rubin
Cognition has become the first disclosed customer running production workloads on Nvidia Vera Rubin NVL72, using systems supplied through CoreWeave rather than waiting for the platform’s broader rollout. CoreWeave says it has deployed hundreds of Rubin GPUs across multiple regions under limited availability and brought Cognition onto the new infrastructure within days of rack handover. In Cognition’s SWE-2 workload, the company measured 4.8 times the total token throughput of a GB200 NVL72 configuration; for reinforcement-learning workloads it measured 3.8 times the output-token throughput per GPU at matched interactivity. Those are customer-and-provider measurements rather than universal benchmarks, but they move Rubin from a future product selection into an operating commercial environment.
Cognition develops Devin, an AI software-engineering agent whose workload combines long-context reasoning, repeated inference and tool interaction. Such applications can stress a system differently from a single prompt-response service because an agent may make hundreds of model calls during one task. Faster decoding and stronger interconnect performance can therefore reduce elapsed task time even when model quality is unchanged. Cognition has expanded its use of CoreWeave to thousands of GPUs in less than nine months, according to the companies, providing an established workload against which the new system can be compared. CoreWeave supplies the operating environment and Nvidia supplies the rack-scale architecture, networking and software underneath it.
The deployment is commercially different from an enterprise announcing an intention to adopt Rubin in the future. Hardware is installed, customer workloads are running and the customer can compare the new system with its existing Nvidia fleet. That creates a reference for other model companies considering whether the performance improvement is large enough to justify migration, capacity reservations or premium pricing. CoreWeave also reduces the migration cost because customers can use similar operational tooling across GB200, GB300 and Vera Rubin infrastructure. A cloud provider able to make generational upgrades look incremental at the software layer helps Nvidia turn each hardware cycle into usable capacity more quickly.
Limited availability still matters. Hundreds of GPUs across several regions are significant for early production but far below the volumes needed for hyperscale adoption. Early users may also receive unusually intensive engineering support, so their deployment experience may not predict the ease of a mass rollout. The most useful evidence will be whether multiple customers can reproduce the performance and operational results as CoreWeave expands capacity. Nevertheless, Cognition provides something that roadmap announcements cannot: a paying production workload generating measurable results on the architecture.
Analysis
Rubin’s first production customer shortens the interval between benchmark promise and commercial utilization. Cognition’s reported 4.8x SWE-2 throughput gives CoreWeave a basis to charge for higher-productivity capacity, while Nvidia gains an upgrade path from existing GB200 and GB300 estates. The economic test is attachment and utilization: if customers migrate workloads rapidly enough to keep expensive Rubin racks busy, generational performance converts directly into cloud revenue and new Nvidia orders. A smooth operational transition can be as valuable as the raw benchmark because it lowers the customer’s switching cost within Nvidia’s own platform.