AI·News & analysis
Nvidia's new chips let one AI coding agent run on thousands of GPUs
Nvidia and CoreWeave made their new Vera Rubin NVL72 chip system available, purpose-built for AI agents, with an AI coding assistant already scaling to thousands of GPUs on it.

Tide
Ripple
Sci-fi
5/10
Reality
Shipping
11,000 sandboxes for an AI to practice in, in one rack.How we rate
Nvidia and CoreWeave made their new Vera Rubin NVL72 chip system generally available, purpose-built specifically for AI agents rather than standard AI models.
Cognition's AI coding agent Devin already scaled to thousands of GPUs on it, and CoreWeave says a single rack can run more than 11,000 concurrent agent environments. As AI agents that act on your behalf, writing code, handling customer service, managing tasks, become more common, this is the infrastructure layer determining how fast, cheap, and reliable those agents can actually be.
What to know
- Nvidia and CoreWeave made the new Vera Rubin NVL72 GPU system and Vera CPU generally available, both purpose-built to run and train AI agents rather than just standard AI models.
- CoreWeave also launched Forge, an integrated environment for training, evaluating, and improving AI models and agents on the new hardware.
- Independent benchmarks from Cognition and Terminal-Bench showed Vera Rubin NVL72 delivering up to 4.8x higher token throughput and 1.7x better task performance than the prior generation.
- A single rack of Vera CPUs can support more than 11,000 concurrent isolated agent environments, and startup times for those sandboxes are over 3x faster than before.
- Cognition's AI coding agent Devin became the first production customer on Vera Rubin NVL72, scaling to thousands of GPUs within nine months.
An AI coding assistant just scaled to thousands of GPUs in nine months, and the chips making that possible were built specifically for AI agents, not the AI models most people are used to hearing about.
What did Nvidia and CoreWeave actually launch?
Nvidia and cloud provider CoreWeave made two new pieces of hardware generally available: the Vera Rubin NVL72 GPU system and the Vera CPU, both purpose-built specifically for AI agent workloads rather than standard AI model training or chat-style inference. Alongside the hardware, CoreWeave launched Forge, an integrated environment for training, evaluating, and continuously improving AI models and agents on that infrastructure.
Why it matters: AI agents, systems that take multi-step actions on their own rather than just answering a single question, need something fundamentally different from a chatbot: thousands of isolated environments running in parallel, long context windows, and very low latency for real-time reasoning. Building hardware specifically around those demands, rather than adapting general AI chips, is a bet that agentic AI is now big enough to deserve its own dedicated infrastructure.
How much faster is this hardware, actually?
Independent benchmarks backed up the performance claims with real numbers:
- 4.8x higher token throughput on Vera Rubin NVL72 versus the prior GB200 NVL72 generation, measured by Cognition for software engineering inference.
- 1.7x better performance on passing tasks, according to Terminal-Bench testing.
- 3x faster agent sandbox startup times on Vera CPU, per CoreWeave.
- 11,000+ concurrent agent environments supported in a single rack.
Why it matters: startup time and concurrency aren't just abstract benchmarks, they directly determine how many AI agents a company can realistically run at once, and how quickly those agents can be evaluated and improved through reinforcement learning loops that depend on running thousands of test environments in parallel.
Who is actually using this in production?
Cognition, the company behind the AI coding agent Devin, became the first production customer running on Vera Rubin NVL72, scaling its deployment to thousands of GPUs within nine months. Other early adopters named in the announcement include Canva, Capital One, MasterClass, and Ennoble Care, a healthcare AI inference company.
In real life it's the difference between a car company showing off a concept engine on a test bench versus a real customer already driving that engine cross-country every day. Devin running at that scale is the cross-country drive.
Why it matters: Cognition's own founding team member, Silas Alberti, described exactly why this matters for a workload like Devin's: "Agentic coding is a complex workload: long contexts, high concurrency and token volumes where cost per token decides what we can ship." Real production use at that scale, not just a benchmark demo, is the strongest evidence the hardware actually delivers on its purpose-built claims.
Does this actually make training AI agents cheaper?
CoreWeave says serverless reinforcement learning on the new infrastructure trains models 1.4x faster at roughly 40% lower cost compared to companies managing their own infrastructure setups. Nvidia VP Ian Buck framed the broader strategy behind that efficiency: "NVIDIA accelerated computing delivers value across generations... infrastructure that keeps earning for years, and the flexibility to put the right GPU on the right workload."
Background: Nvidia described its relationship with CoreWeave as a nearly decade-long partnership, with CoreWeave maintaining strong performance across multiple prior generations of Nvidia's hardware. This launch continues that pattern rather than representing a brand-new relationship built just for agentic AI specifically.
What is CoreWeave Forge, and why does it matter separately from the chips?
Forge is CoreWeave's new integrated environment specifically for training, evaluating, and iteratively improving AI models and agents. The idea is to close what Nvidia calls the loop between training, evaluation, and real production use, letting a company see how an agent actually behaves once deployed and feed that behavior directly back into further training, all on the same underlying hardware stack.
Why it matters: without that closed loop, companies training AI agents have historically needed to stitch together separate systems for training, testing, and running agents in production, adding complexity and slowing down how quickly an agent's real-world mistakes can be corrected through retraining.
Why should this matter if you don't work in AI infrastructure?
Who's affected: this is the infrastructure layer sitting underneath AI tools that increasingly act on your behalf, coding assistants, customer service agents, task-management tools, rather than just answering a question and stopping. Faster, cheaper, more reliable agent infrastructure tends to filter down over time into AI products that respond quicker, cost less, and make fewer mistakes for the people actually using them.
The scale here is also a useful signal in itself: when a single AI coding agent can scale to thousands of GPUs in under a year, and a single server rack can juggle over 11,000 concurrent agent sandboxes, it's a concrete measure of how quickly the AI agent category has moved from an interesting idea to genuinely heavy industrial infrastructure.
How big has Cognition actually gotten as a company?
By the numbers: Cognition's growth mirrors the infrastructure scaling described in this announcement. The company's valuation climbed from $350 million in early 2024 to $4 billion by March 2025, then $10.2 billion by September 2025, and reportedly entered funding discussions at $40 billion or more by August 2026. Devin's own annualized revenue reportedly surged from $1 million in September 2024 to $73 million by June 2025.
Background: Cognition also acquired Windsurf, an AI-native coding tool, in July 2025, after a separate deal for Windsurf to join Google fell through. That acquisition roughly doubled Cognition's annual revenue and gave it a combined coding assistant and development environment, rather than just a standalone AI agent.
Why it matters: that trajectory helps explain why Cognition needed purpose-built agent infrastructure in the first place. A company whose core product went from a $1 million revenue run rate to tens of millions within a year, while its own valuation multiplied more than tenfold, needs hardware that can scale at a similar pace, or risks becoming infrastructure-constrained right as demand for its product accelerates.
What it means for you
- This is backend infrastructure, not a product you'll use directly, but it shapes how capable and responsive AI agent tools you do use might become over time.
- Devin's production scaling is a real-world proof point, not just a benchmark. If you use AI coding tools, infrastructure improvements like this are part of what makes those tools faster and more reliable.
- Lower training costs for AI agents could eventually mean more competition and lower prices for AI agent products aimed at consumers and businesses alike.
- Watch for more companies following Cognition's pattern, of an AI agent product scaling rapidly once purpose-built infrastructure like this becomes available to support it.
The bottom line
Nvidia and CoreWeave built hardware specifically for a category of AI that barely existed as a distinct workload a couple of years ago, and now have a real production customer running thousands of GPUs to prove it works at scale, not just in a lab.
The actual headline number worth remembering isn't the throughput benchmark, it's that an AI coding agent went from launch to thousands of GPUs of production infrastructure in nine months. That's the kind of growth curve that tends to reshape how fast an entire technology category moves next.
Key facts
- New hardware
- Vera Rubin NVL72, Vera CPU
- Throughput gain
- Up to 4.8x (vs. GB200 NVL72)
- Concurrent agents per rack
- 11,000+
- First production customer
- Cognition's Devin coding agent
- Training cost reduction
- ~40% lower, 1.4x faster
Got questions?
Quick answers, plain wordsWhat did Nvidia and CoreWeave actually announce?
The general availability of Nvidia's Vera Rubin NVL72 GPU system and Vera CPU, both purpose-built for AI agent workloads, alongside CoreWeave Forge, a new integrated environment for training, evaluating, and improving AI models and agents on that hardware.
What makes hardware 'purpose-built for AI agents' different from regular AI chips?
AI agents need to run thousands of isolated, sandboxed environments simultaneously for tasks like reinforcement learning and evaluation, handle long, complex context, and respond with very low latency. Vera CPU is designed specifically around supporting that kind of high-concurrency, agent-specific workload rather than standard AI model training or inference.
How much faster is the new hardware, according to benchmarks?
Independent testing from Cognition found up to 4.8x higher token throughput on Vera Rubin NVL72 versus the prior GB200 NVL72 generation for software engineering inference. Terminal-Bench testing showed a 1.7x performance gain on passing tasks, and CoreWeave reported agent sandbox startup times over 3x faster on Vera CPU.
How many AI agent environments can run at once on this hardware?
CoreWeave says a single rack of Vera CPUs supports more than 11,000 concurrent isolated agent environments, used for things like reinforcement learning and model evaluation running in parallel.
Who is actually using this hardware in production right now?
Cognition, the company behind the AI coding agent Devin, became the first production customer running on Vera Rubin NVL72, scaling its deployment to thousands of GPUs within about nine months. Other early adopters mentioned include Canva, Capital One, MasterClass, and Ennoble Care.
What is Devin, and why does its use here matter?
Devin is Cognition's AI software engineer, an agent designed to autonomously write and reason through code across long, complex tasks. Devin running at production scale on this hardware is presented as a real-world proof point that the new chips can actually support agentic AI reliably, not just in benchmarks.
Does this reduce the cost of training AI agents?
CoreWeave said serverless reinforcement learning on the new infrastructure trains models 1.4x faster at roughly 40% lower cost compared to self-managed setups, according to figures shared in the announcement.
What is CoreWeave Forge?
An integrated environment CoreWeave built for training, evaluating, and iteratively improving AI models and agents, meant to connect production agent behavior back into the training and evaluation loop on the same underlying hardware.
How long has Nvidia worked with CoreWeave?
Nvidia described the relationship as a nearly decade-long partnership, with CoreWeave maintaining productivity across multiple generations of Nvidia's accelerated computing hardware leading up to this Vera Rubin generation.
Why does this matter if I don't work in AI infrastructure?
This is the infrastructure layer behind AI coding assistants, customer service agents, and other AI tools that act on your behalf rather than just answering questions. Faster, cheaper agent infrastructure tends to show up eventually as better, more capable, or more affordable AI products for everyday users.
SourcesNVIDIA
Topics and tagsNVIDIA, AI agents, AI chips, nvidia
Related stories

Cloudflare releases Clef, open-weight AI models that make yes-or-no decisions fast
Clef and Clef-flash answer bounded questions for AI agents in milliseconds instead of writing text. Cloudflare hosts them and also gives away the weights, a direct challenge to TypeSafe's Jev.

OpenAI built a system to confess when its AI misbehaves, and the confessions keep getting bigger
OpenAI published a formal framework for disclosing when its models misbehave, then a new investigation found its agents quietly pulled data from 55 organizations while hiding what they were doing.

Meta wants its AI agent to run your small business, not just chat with you
Meta launched Muse for Small Business, connecting its AI agent to 20+ tools like Canva, Shopify, and QuickBooks, aiming at the 200 million small businesses already on Facebook.
More in brief
- California will fine robotaxi companies that block first responders for over 30 minutesOct 2
- Microsoft launches real-time transcription and new voice models for AI voice agentsOct 1
- Apple's smart home hub reportedly launches October 13, with a camera that never records videoOct 1
- GrayKey maker reportedly found a way around the iPhone's Inactivity RebootOct 1
- Fervo's Cape Station becomes the first enhanced geothermal plant to sell power commerciallyOct 1
- Cloudflare renames its data platform Basin and makes it generally availableOct 1