CoreWeave Deploys Next-Generation NVIDIA Vera Rubin Systems for Production AI Agent Workloads

CoreWeave announced the availability of NVIDIA Vera Rubin NVL72 systems equipped with Spectrum-X 102.4T Ethernet networking, with Cognition becoming the first customer to run production workloads on the new hardware. Cognition's benchmarks show Vera Rubin delivering up to 4.8x higher token throughput compared to GB200 NVL72 baseline systems when running real-world software engineering tasks. CoreWeave is also introducing NVIDIA Vera, a CPU purpose-built for AI agents, alongside CoreWeave Forge, a platform designed for training, evaluating and improving models and agents on NVIDIA accelerated infrastructure.
CoreWeave's infrastructure strategy demonstrates sustained value across multiple GPU generations, with systems from nearly a decade ago still handling production workloads alongside newly deployed hardware. The company's integration of NVIDIA's latest accelerators with specialized networking and software tools positions it to serve the emerging category of AI agent applications, which present distinctive computational challenges including sustained high token throughput and real-time responsiveness requirements.
Cognition's benchmarking results on production software engineering tasks provide early evidence of Vera Rubin's performance characteristics in demanding agentic workloads. The rapid deployment timeline—days from receiving hardware to production cluster operation—reflects the depth of engineering coordination between CoreWeave and NVIDIA, suggesting that established partnerships may accelerate adoption cycles for next-generation infrastructure.
The availability of specialized AI accelerator infrastructure could reshape competitive dynamics in AI development by reducing barriers for smaller teams to scale production systems. Organizations with limited capital for infrastructure investment may face widening performance gaps, while those accessing CoreWeave's platform gain access to bleeding-edge hardware with minimal deployment friction. The demonstrated performance improvements may also influence which computational tasks become economically viable to automate, potentially affecting labor markets in software engineering and similar domains requiring intensive reasoning.