CoreWeave maintains industry-leading AI cloud performance across multiple compute generations with its software platform
CoreWeave Inc. (Nasdaq: CRWV), The Essential Cloud for AI™, today announced the availability of NVIDIA Vera Rubin NVL72 on CoreWeave, with Cognition as the first customer anywhere running production workloads on the system. Customers, like Cognition, run the system under the same operating model and tooling as their existing NVIDIA GB200 NVL72 and GB300 NVL72 fleets, with performance engineering from CoreWeave’s team. The news was shared during Fully Connected, CoreWeave’s AI cloud conference, which brings together more than 4,500 customers, partners, developers and AI leaders to share how they are building and running AI in production.
This press release features multimedia. View the full release here: https://www.businesswire.com/news/home/20260930041932/en/

NVIDIA Vera Rubin NVL72 systems deployed on CoreWeave Cloud for production AI workloads.
“Bringing up NVIDIA Vera Rubin NVL72 so quickly, and having a customer already seeing performance gains within days, is the payoff from years of engineering our platform across GPU generations,” said Chen Goldberg, executive vice president of product & engineering at CoreWeave. “With customers like Cognition, that investment shows up in the ability to get production workloads running within days. When it comes to agentic tasks, long contexts, repeated model calls and thousands of concurrent tasks put pressure on the entire platform. Our job is to make compute, networking and software work as a single system, so customers can build increasingly complex agents without taking on the infrastructure complexity themselves.”
Cognition is the first customer in production with Vera Rubin NVL72
Cognition, the applied AI lab behind Devin, runs training, reinforcement learning and production inference for Devin on CoreWeave. The company scaled from bridge capacity to thousands of GPUs for training and inference in less than nine months. Cognition worked with CoreWeave to stand up a Vera Rubin NVL72 cluster in early September, and Cognition’s own engineers ran the first customer-executed Vera Rubin inference benchmark, measured against a GB200 NVL72 cluster baseline.
“Agentic coding is an unforgiving workload that requires long contexts, high concurrency and rapid reasoning,” said Silas Alberti, SVP research & founding team at Cognition. “By deploying the NVIDIA Vera Rubin NVL72 on CoreWeave, our engineers are seeing up to a 4.8 times increase in total token throughput for SWE-2 inference workloads. For an agentic workload where every step waits on the last one, that compounds into real work Devin gets done. CoreWeave continues to deliver the bleeding-edge rack-scale acceleration we need to push the boundaries of AI.”
Cognition’s engineers benchmarked Vera Rubin
In independent benchmarks run on CoreWeave Cloud, Cognition measured a 4.8 times increase in total token throughput for its SWE-2 inference workloads on NVIDIA Vera Rubin NVL72 compared to a GB200 NVL72 baseline. Additionally, the team recorded a 3.8 times boost in output token throughput for reinforcement learning workloads. For Cognition, that translates to more concurrent Devin sessions per GPU, drastically accelerated research loops and lower cost per session, with no loss in generation speed.
Proven across every NVIDIA generation
The relationship between CoreWeave and NVIDIA dates to 2017, beginning with the NVIDIA Volta generation, which is still in commercial service today on CoreWeave Cloud, demonstrating the long useful life of NVIDIA compute and the value CoreWeave’s platform can continue to draw from it. CoreWeave’s full-stack software platform—including CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control®, CoreWeave Sandboxes and serverless inference—gives customers a consistent way to deploy and manage workloads across GPU generations. Customers can match each workload to appropriate capacity, keep using existing infrastructure as needs evolve and adopt new NVIDIA architectures through a familiar operating environment.
“CoreWeave has consistently demonstrated the infrastructure expertise required to bring each new generation of NVIDIA accelerated computing into production,” said Ian Buck, vice president of Hyperscale and High-Performance Computing at NVIDIA. “Its full-stack expertise across multiple generations of NVIDIA infrastructure is helping AI innovators like Cognition quickly put NVIDIA Vera Rubin NVL72 to work on demanding production workloads.”
CoreWeave, in collaboration with Dell Technologies, was among the first cloud providers to deploy Dell PowerRack systems featuring NVIDIA GB200 and GB300 NVL72, and is one of the first to deploy NVIDIA Vera Rubin. CoreWeave has also published the industry’s first measured silicon performance numbers on the platform, showing 10 times the token throughput per megawatt over NVIDIA GB200 NVL72 on the DeepSeek R1 reasoning model at matched interactivity.
CoreWeave consistently delivers industry-leading performance, demonstrated by record-breaking MLPerf benchmark results in inference and training and its position as the only AI cloud to earn the top Platinum ranking in SemiAnalysis ClusterMAX™ three times in a row.
About CoreWeave
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a fully integrated platform of technology and teams that enables innovators to move at the pace of innovation, building and scaling AI with confidence. Trusted by 9 of the 10 leading foundation model providers, CoreWeave serves as a force multiplier by combining superior infrastructure performance with deep technical expertise to accelerate human discovery. Established in 2017, CoreWeave completed its public listing on Nasdaq (CRWV) in March 2025.
View source version on businesswire.com: https://www.businesswire.com/news/home/20260930041932/en/
Contacts
Media
press@coreweave.com
If you believe this article contains misleading, harmful, or spam content, please let us know.
Report this article