AgentX
Nvidia's Rubin NVL72 provides 67x better performance per dollar for agentic inference. This platform is the first co-designed across six specific products: the Rubin GPU, Vera CPU, NVLink 6 Switch, ConnectX-9, BlueField-4, and an additional component. These hardware advancements support agentic AI workflows that demand 15 times more tokens than standard chat requests. The architecture aims to reduce the cost per 1 million tokens by up to 35 times while increasing AI factory throughput per megawatt by up to 30 times compared to the GB300 NVL72.
What changed
Rubin NVL72 is now identified as a co-designed six-product platform delivering 67x better performance per dollar for agentic inference.
Live updates
-
Rubin NVL72 achieves 67x performance per dollar for agentic inference
Nvidia's Rubin NVL72 provides 67x better performance per dollar for agentic inference. This platform is the first co-designed across six specific products: the Rubin GPU, Vera CPU, NVLink 6 Switch, ConnectX-9, BlueField-4, and an additional component. These hardware advancements support agentic AI workflows that demand 15 times more tokens than standard chat requests. The architecture aims to reduce the cost per 1 million tokens by up to 35 times while increasing AI factory throughput per megawatt by up to 30 times compared to the GB300 NVL72.
Why it matters
Agentic AI requires repetitive context processing and parallel sub-agent calls, creating high resource demands. Nvidia uses the TensorRT-LLM Python API and runtimes to optimize these inference workflows. The Rubin architecture targets the specific efficiency needs of programming assistants and other complex AI agents.
What is confirmed
- The Rubin platform is co-designed across six products including the Rubin GPU, Vera CPU, NVLink 6 Switch, ConnectX-9, and BlueField-4.
- Rubin NVL72 provides 67x better performance per dollar for agentic inference.
What to watch next
- Confirmation of the sixth product in the Rubin co-designed platform
- Deployment data for Rubin NVL72 in active AI factories
confidence 90%Sources used for this update (5)
- ee.ofweek.com — 深度丨OpenAI第一代自研芯片亮线,超预期背后的算力新打法
- markets.businessinsider.com — SOLOWIN HOLDINGS (AXG) Targets 100+ MW AI Computing Capacity By 2028 With 1 GW+ Development Pipeline, Advances Quantum and Regulated Stablecoin Strategy
- www.nola.com — Three men cited for alligator harvesting violations in St. Martin, St. Mary parishes
- github.com — modal-projects
- www.europesays.com — Rubin NVL72 Agentic Inference: 67x better Performance per Dollar
-
Nvidia Vera Rubin architecture targets agentic AI efficiency
Nvidia's Vera Rubin NVL72 architecture increases AI factory throughput per megawatt by up to 30 times compared to the GB300 NVL72. This performance gain targets agentic AI workflows, which require 15 times more tokens than standard chat requests. The architecture reduces the cost per 1 million tokens by up to 35 times. These improvements address the high resource demands of AI agents, such as programming assistants, which require repetitive context processing and parallel sub-agent calls. Nvidia supports these inference optimizations through its TensorRT-LLM Python API and runtimes.
Why it matters
Agentic AI differs from traditional chatbots by executing multi-step workflows and unpredictable activity patterns. This shift necessitates a change in inference infrastructure to handle massive context reuse. The Vera Rubin architecture specifically optimizes for these DeepSeek V4-Pro workloads.
What is confirmed
- The Vera Rubin NVL72 increases AI factory throughput per megawatt by up to 30 times over the GB300 NVL72.
- Multi-step AI workflows consume 15 times more tokens than standard chat requests.
- The Vera Rubin architecture reduces the cost per 1 million tokens by up to 35 times.
- TensorRT-LLM provides a Python API and C++ runtimes to orchestrate inference execution on Nvidia GPUs.
Still unconfirmed
- Programming assistants and agent systems generate unpredictable activity through parallel sub-agent calls and repeated context processing.
What to watch next
- Deployment data for DeepSeek V4-Pro on Vera Rubin NVL72 clusters.
confidence 90%Sources used for this update (4)
- github.com — Pull requests: NVIDIA/TensorRT-LLM
- forbesjapan.com — エヌビディア、AMD比最大5倍のコスト効率を記録──AI推論向け新ベンチマーク
- www.mlbtraderumors.com — Jeff Greenberg Withdraws From Red Wings GM Search
- news.qq.com — llm-d如何最大化利用现有硬件资源
-
Nvidia Vera Rubin achieves 30x power efficiency gain over Blackwell in agentic benchmarks
Nvidia released performance data for its Vera Rubin architecture showing a massive increase in power efficiency for agentic AI. Using the SemiAnalysis AgentX benchmark with DeepSeek V4-Pro workloads, the Vera Rubin NVL72 increased AI factory throughput per megawatt by up to 30 times compared to the GB300 NVL72. This shift targets the rising demand for multi-step AI workflows, which consume 15 times more tokens than standard chat requests. The new architecture also reduces the cost per 1 million tokens by up to 35 times.
Why it matters
AI agents are moving from single-turn interactions to complex workflows involving tool calls and sub-agent coordination. This transition increases the average prompt token count by approximately 4 times. Nvidia is positioning Vera Rubin to handle these high-token demands more efficiently than the current Blackwell system.
What is confirmed
- Vera Rubin NVL72 provides up to 30 times more AI factory throughput per megawatt than GB300 NVL72 in AgentX DeepSeek V4-Pro workloads.
- The cost per 1 million tokens for Vera Rubin is up to 35 times lower than for the GB300.
- Single agent requests consume 15 times more tokens than general chat requests according to OpenRouter data.
- Average prompt tokens per request have increased approximately 4 times based on an analysis of 100 trillion tokens.
What to watch next
- Independent verification of Vera Rubin's 30x efficiency claims via third-party benchmarks
- Comparison of Vera Rubin's throughput per watt against OpenAI's Jalapeño chip
- Official release date and availability of the Vera Rubin NVL72 system
confidence 90%Sources used for this update (4)
- www.econovill.com — 엔비디아, 베라 루빈 실측 공개… "블랙웰보다 전력당 처리량 30배"
- github.com — hermes-skill
- github.com — nexthop-ai
- markets.businessinsider.com — SOLOWIN HOLDINGS (AXG)’s AlloyX HK and EvolveQ Signed Strategic MOU to Integrate AI and Quantum Computing for Next-Generation Financial Infrastructure
-
OpenAI Jalapeño Chip Challenges Nvidia Blackwell and Rubin Performance
OpenAI's Jalapeño inference chip outperforms Nvidia's Blackwell system in power efficiency and latency according to public benchmarks. The chip utilizes TSMC's N3P process and features 15.4TB/s of memory bandwidth. While Nvidia reported strong FY2027 Q2 results with 962 billion USD in revenue and 117% growth in data center earnings, the Jalapeño chip threatens Nvidia's CUDA ecosystem and the upcoming Rubin architecture. OpenAI claims the hardware provides 1.9x more throughput per watt and 3.6x lower latency than the Blackwell flagship.
Why it matters
Nvidia currently leads the agentic AI infrastructure market with hardware like the Vera Rubin NVL72. OpenAI is attempting to reduce its dependence on Nvidia by developing custom silicon. This shift represents a broader industry trend where AI software leaders build proprietary hardware to optimize cost and performance.
What is confirmed
- Nvidia reported FY2027 second quarter revenue of 962 billion USD, a 106% increase year-over-year.
- Nvidia data center revenue grew 117% year-over-year to 890 billion USD.
- OpenAI's Jalapeño chip exceeds Nvidia's Blackwell system in power efficiency and latency in public benchmarks.
Still unconfirmed
- The Jalapeño chip utilizes TSMC's N3P process and offers 15.4TB/s of memory bandwidth.
- Jalapeño is expected to enter mass production in 2027.
- Jalapeño's performance and cost-efficiency threaten the upcoming Rubin architecture.
What to watch next
- Confirmation of Jalapeño mass production dates
- Independent third-party verification of internal OpenAI test data
- Nvidia's official response to Jalapeño benchmark results
confidence 80%Sources used for this update (5)
- wingsoverscotland.com — Wings Over Scotland| The world's most-read Scottish politics website
- www.manilatimes.net — SOLOWIN HOLDINGS (NASDAQ: AXG) Announces AI Infrastructure Expansion with 100 MW High-Performance Computing Target
- stock.10jqka.com.cn — 营收暴涨106%,却先跌3%:英伟达的超预期,为什么不及格?利空
- www.mirrormedia.mg — 黃仁勳惡夢成真?OpenAI晶片實測嗆爆輝達
- zaikei.co.jp — OpenAI独自推論チップ「Jalapeño」、Nvidia「Blackwell」比で電力効率最大1.9倍 測定条件には留意
-
Nvidia Vera Rubin NVL72 boosts agentic AI throughput 30-fold
Nvidia is scaling its agentic AI infrastructure with the Vera Rubin NVL72, which delivers 30 times higher throughput per megawatt than the GB300 NVL72. This hardware supports agents that research, code, and reason through multiple steps. While Nvidia maintains a cost-efficiency lead over AMD in coding-agent benchmarks, OpenAI is challenging this dominance with its custom Jalapeño inference chip. OpenAI claims the Jalapeño chip provides 1.9x more throughput per watt and 3.6x lower latency than Nvidia's Blackwell flagship.
Why it matters
Agentic AI requires iterative reasoning and massive context lengths, creating significant power and cost constraints for data centers. Nvidia is combining Vera CPUs, Rubin GPUs, and Groq 3 LPX accelerators to establish a foundational framework for this era. OpenAI is pursuing vertical integration of its AI stack ahead of a planned IPO.
What is confirmed
- The Vera Rubin NVL72 delivers up to 30 times higher throughput per megawatt for agentic AI workloads compared to GB300 NVL72 systems.
Still unconfirmed
- Nvidia's Vera Rubin NVL72 provides a 35x reduction in cost per token when running DeepSeek V4 Pro models compared to the GB300 NVL72.
What to watch next
- Independent verification of Jalapeño chip benchmarks against Blackwell
- Deployment timelines for Vera Rubin NVL72 on Starmind AI satellites
- Official IPO filing from OpenAI
confidence 80%Sources used for this update (7)
- uk.news.yahoo.com — One Agent Benchmark Puts Nvidia 5x Ahead Of AMD On Cost
- eu.36kr.com — NVIDIA's Stunning First Benchmark of Vera Rubin: DeepSeek Throughput Surges 30-Fold
- officechai.com — OpenAI’s New Jalapeno Chip Beats NVIDIA’s Blackwell On Some Parameters, Company Says
- technosports.co.in — NVIDIA’s New Chip Delivers 30x More AI Work Per Watt
- glitchwire.com — OpenAI's Jalapeño Chip Posts Spicy Benchmark Results That Challenge Nvidia. Here's What It Means for Users.
- datacenters.economictimes.indiatimes.com — Nvidia Vera Rubin Boosts AI Throughput Per Megawatt
- finance.yahoo.com — Solowin Holdings (AXG) Stock Price, News, Quote & History - Yahoo Finance
-
Nvidia Launches Groq 3 LPX and Vera Rubin for Agentic AI
Nvidia has entered full production of Groq 3 LPX inference accelerators to support agentic AI, which requires iterative reasoning and massive context lengths. The company is deploying these alongside Vera CPUs and Rubin GPUs. SpaceXAI is adopting the Vera CPU for agent orchestration and plans to deploy Vera Rubin NVL72 hardware on Starmind AI satellites. Performance data for the Vera Rubin NVL72 shows a 30x increase in throughput per megawatt and a 35x reduction in cost per token when running DeepSeek V4 Pro models compared to the GB300 NVL72.
Why it matters
Agentic AI differs from standard chatbots by executing hundreds of reasoning steps, calling tools, and managing contexts of hundreds of thousands of tokens. Nvidia is shifting its infrastructure focus from raw training power to token generation speed, energy efficiency, and cost. This transition aims to redefine the economic viability of autonomous AI agents.
What is confirmed
- Nvidia has started full production of Groq 3 LPX inference accelerators.
- SpaceXAI is using NVIDIA Vera CPUs to accelerate agentic AI.
- The Vera Rubin NVL72 achieves up to 30 times the throughput per megawatt of the GB300 NVL72 when running DeepSeek V4 Pro models.
- Token costs for DeepSeek V4 Pro can be reduced by up to 35 times on the Vera Rubin platform.
- Groq 3 LPX reached 3,400 output tokens per second in tests using the Gemma 4 31B model with a 100,000 token context.
- Nvidia's Groq racks will be online this year following a 20 billion dollar purchase.
Still unconfirmed
- Nvidia is 5x ahead of AMD on cost according to one agent benchmark.
- SpaceXAI will deploy Vera Rubin NVL72 on first-generation Starmind AI satellites.
- The Vera CPU is 1.8 times faster than x86 for agent AI orchestration.
What to watch next
- Deployment of Vera Rubin hardware in Nebius data centers
- Confirmation of Starmind AI satellite orbital computing performance
- Further cost-comparison benchmarks between Nvidia and AMD for agentic workloads
confidence 95%Sources used for this update (14)
- SemiAnalysis — AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing?
- CNBC — Nvidia says Groq racks will be online this year following $20 billion purchase
- NVIDIA Newsroom — SpaceXAI Adopts NVIDIA Vera CPU to Accelerate Agentic AI at Massive Scale
- Yahoo Finance — NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI
- The Information — Nvidia Announces New Customers For Vera CPU, Groq LPX Racks
- consent.yahoo.com — One Agent Benchmark Puts Nvidia 5x Ahead Of AMD On Cost
- news.qq.com — 英伟达新一代算力平台补齐拼图 结合DeepSeek模型Token成本可降35倍
- tech.ifeng.com — 英伟达宣布Groq 3 LPX全面投产:Vera Rubin推理效率实现数倍提升
- www.stnn.cc — 万亿美元“买债救市”要来了?金价飙升
- news.cnyes.com — 輝達宣布Groq 3 LPX全面投產 Vera Rubin推理效率大提升 SpaceX加碼AI代理與太空AI
- news.pedaily.cn — 英伟达震撼首测Vera Rubin,DeepSeek吞吐暴涨30倍
- www.wikitree.co.kr — 엔비디아 ‘베라’ CPU, 에이전트 AI 오케스트레이션 전담해 x86보다 1.8배 빠르다