NVIDIA Advances Vera Rubin Inference With New LPX and CPX Platforms for Faster AI Performance, Lower Token Costs
NVIDIA is enhancing its Vera Rubin NVL72 system with new LPX and CPX platforms, now in full production, to accelerate AI inference for agentic systems. These advancements aim to deliver faster token generation and lower costs, with key partners like SpaceXAI, CoreWeave, a…
Intelligence analysis by Gemini 2.5 Flash

NVIDIA is shifting its focus to optimizing the entire AI factory for inference, emphasizing an integrated approach rather than single chip breakthroughs. The Vera Rubin system, featuring the Groq 3 LPX, is designed to meet the escalating demands of long-context inference and multi-agent AI, demonstrating significant performance gains in benchmarks.
Imagine you have a super-smart robot helper that needs to think and talk really fast to do its jobs. NVIDIA has built a special super-computer system, like a super-efficient factory, that helps these robot helpers think and talk much, much quicker. It's like giving them a super-fast brain and mouth so they can understand long stories and give quick answers, making them much better at helping us with tricky tasks.
Analysis
NVIDIA's latest announcement signals a strategic evolution in its approach to AI infrastructure, moving beyond individual component optimization to a holistic "AI factory" concept. The company emphasizes that the future of AI inference lies in the seamless integration of compute, networking, and acceleration layers. This integrated strategy is particularly geared towards the burgeoning field of agentic AI, which demands unprecedented levels of throughput, responsiveness, and economic efficiency for processing large context windows and generating tokens rapidly.
Vera Rubin NVL72
The NVIDIA Vera Rubin NVL72 rack-scale system is positioned as the versatile foundation for these advanced AI factories. It is engineered to accelerate inference as AI agents engage in increasingly complex reasoning over extended sequences of data. The platform's design focuses on delivering low latency, extreme throughput, and scalable economics, which are critical for the interactive and real-time demands of agentic applications. By extending the Vera Rubin NVL72 with specialized accelerators, NVIDIA aims to ensure that valuable GPU infrastructure remains fully utilized, helping agents complete tasks faster and more predictably.
Groq 3 LPX
The NVIDIA Groq 3 LPX is a new low-latency inference architecture specifically codesigned to work alongside the Vera Rubin NVL72 platform. Its primary role is to accelerate token generation at ultrafast speeds, directly addressing the decode latency challenge inherent in agentic AI workloads. As AI agents reason, use tools, and interact, they generate responses one token at a time, making even minor delays significant. The Groq 3 LPX, in conjunction with Rubin GPUs handling large-scale context processing, is designed to eliminate the traditional trade-off between speed and throughput, enabling responsive, large-scale inference for next-generation AI applications. Benchmarks show it delivering 3,400 output tokens per second for 100,000-token long-context use cases, a fourfold improvement over alternative platforms.
SpaceXAI
SpaceXAI has announced its plans to adopt NVIDIA Vera CPUs to power its next generation of agentic AI, from data centers on Earth to orbital satellites. This partnership highlights the comprehensive nature of NVIDIA's full-stack AI platform, which includes Vera CPUs, accelerated computing, networking, and software. SpaceXAI intends to deploy Vera CPUs for CPU-intensive tasks such as orchestration, tool use, code execution, data processing, and simulation, which are fundamental to agentic AI. This collaboration underscores the platform's capability to scale AI architectures to unprecedented levels, providing leading per-core performance and predictable performance under load, crucial for complex, distributed AI operations.
Key points
- NVIDIA's Vera Rubin NVL72 system is extended with new LPX and CPX platforms for AI inference, now in full production.
- The NVIDIA Groq 3 LPX delivers 3,400 output tokens per second for 100,000-token long-context use cases, 4x faster than alternatives.
- Key industry partners, including SpaceXAI, CoreWeave, and Nebius, are adopting Vera Rubin platform solutions.
- NVIDIA employs an "extreme codesign" approach, optimizing compute, networking, and inference acceleration as a unified system.
- The goal is to build "token factories" for efficient, low-latency, and scalable agentic AI applications, addressing decode latency challenges.
The advancements could lead to significantly more responsive and powerful agentic AI applications, enabling complex problem-solving and real-time interactions at scale. Lower token costs and increased efficiency could democratize access to advanced AI capabilities for developers and enterprises, fostering innovation across various industries.
While promising, the complexity of integrating these full-stack solutions might pose adoption challenges for some organizations, requiring significant infrastructure overhauls and specialized expertise. The reliance on proprietary NVIDIA hardware and software could also limit flexibility and foster vendor lock-in for AI developers and cloud providers.
Market signals
- NVDA The announcement of new platforms in full production and key partner adoptions signals strong market adoption and technological leadership for NVIDIA in the AI inference space.
AI-generated analysis of potential market relevance. Not financial advice.



