Report Ads

Nvidia Vera CPU Launch Fires First Direct Salvo in the Server Processor War

Nvidia
From gaming to AI, Nvidia drives visual computing innovation. [TechGolly]

Key Points:

  • Nvidia released the architectural specifications and benchmarks for its new custom “Vera” data center CPU, code-named the Olympus core.
  • Branded as “the CPU for agents,” Vera features 88 custom cores and 176 threads designed specifically to power complex agentic AI loops.
  • In SPEC CPU 2026 benchmarks, Vera delivered a 10% performance advantage over AMD’s EPYC Turin and a 55% lead over Intel’s Xeon.
  • Early shipments of the custom processor began in June to premier clients, including OpenAI, Anthropic, and SpaceX.

The battle for supremacy in the global data center market has officially expanded into a new, highly competitive arena. The semiconductor giant Nvidia has released the comprehensive architectural specifications and industry-standard benchmarks for its first custom-designed central processing unit, the Vera CPU. Branded specifically as “the CPU for agents,” this high-performance processor represents a direct, historic challenge to the traditional x86 server monopolies held for decades by AMD and Intel. By fusing its dominant graphics hardware with custom-designed host processors, the tech giant is attempting to seize complete control over the physical and economic foundations of modern AI factories.

Unlike previous generations of server chips that relied on off-the-shelf processor designs, the newly unveiled CPU features an entirely in-house microarchitecture code-named the Olympus core. The physical chip packages 88 custom Olympus cores and 176 threads, running on a highly advanced Arm-based architecture. Rather than building the processor from multiple smaller chips wired together—a design known as “chiplet” architecture—the company built the processor on a single, massive piece of silicon. This unified, monolithic design ensures that every single processing core sits at the exact same distance from the memory, eliminating the high data-movement latencies that plague traditional chiplet processors.

This strict focus on uniform memory latency is necessary to support the rapidly growing market for “agentic” artificial intelligence. While a standard chat query barely taxes a host processor, autonomous AI agents—which can write software, execute SQL database queries, and take multi-step actions across various applications—generate highly complex, sequential control loops. An agent session typically runs these loops 100 to 300 times, requiring the host CPU to execute rapid tool-calling and data services between GPU processing cycles. If the CPU lags during these sequential phases, the expensive graphics processors are left sitting idle, crushing the overall cost efficiency of the data center.

To feed these demanding sequential loops, the custom CPU incorporates a highly advanced, ultra-fast memory subsystem. The processor utilizes the latest SOCAMM2 LPDDR5X memory technology to deliver up to 1.2 terabytes per second (TB/s) of direct memory bandwidth, alongside a massive memory capacity of up to 1.5 terabytes per CPU. This memory subsystem delivers over four times the memory bandwidth per core compared to traditional x86 server processors. It also maintains highly consistent, predictable latency even when running hundreds of independent, automated agent sandboxes in parallel.

The company’s newly published SPEC CPU 2026 benchmarks demonstrate the immense, real-world performance advantage of this custom architecture. In the industry-standard integer suite, a dual-socket configuration of the new processor outperformed AMD’s flagship 128-core EPYC 9755 “Turin” CPU by approximately 10% in select single-threaded workloads. The performance lead was even more dramatic against Intel, with the custom chip delivering a massive 55.3% advantage over a single-socket, 128-core Intel Xeon 6980P processor, proving that a highly optimized, custom-designed core can easily outperform older-generation x86 architectures.

The commercial deployment of these advanced processors is already well underway. Early-access shipments of the custom CPU began in June to some of the world’s most prominent artificial intelligence developers and cloud providers. The prestigious roster of early adopters includes OpenAI, Anthropic, and SpaceX (SpaceXAI), which are integrating the host processors directly into their massive internal training clusters. Major hyperscale cloud providers, including ByteDance, CoreWeave, and Oracle Cloud Infrastructure, have also received shipments of the new chips to power their own public, high-throughput AI services.

To support the massive, high-volume commercial rollout of the new technology, major global system manufacturers are rapidly expanding their hardware portfolios. Industry giants like Dell Technologies, Hewlett Packard Enterprise, Lenovo, and Supermicro have completed engineering designs to build and deliver standalone servers powered by the custom CPU. Other prominent manufacturing partners, including ASUS, Gigabyte, Foxconn, and Quanta Cloud Technology, are also adopting the platform, ensuring that enterprises can easily acquire the advanced host processors through their preferred hardware channels.

The primary deployment vehicle for the new host processor will be the company’s massive, next-generation “Vera Rubin” AI platform. Designed specifically to run gigascale AI factories, the Vera Rubin NVL72 platform unifies 72 advanced GPUs with 36 of the new custom CPUs in a single, liquid-cooled server rack. By connecting these processors together using high-speed NVLink interfaces capable of 1.8 TB/s of bandwidth, the system operates as a single, massive supercomputer, allowing developers to optimize their full hardware and software stack to achieve the highest performance per watt and the lowest possible token cost.

This aggressive push into the CPU market represents a long-term, structural threat to traditional semiconductor manufacturers. Historically, companies like Intel and AMD dominated the lucrative data center server market, collecting billions in revenue by selling host processors to sit alongside Nvidia’s GPUs. By developing its own high-performance CPU and tightly integrating it into its unified server systems, the GPU giant is successfully transforming the host processor from an external, third-party purchase into a core, proprietary component, potentially capturing a massive portion of the $200 billion total addressable server market.

Ultimately, the release of the custom CPU specifications and benchmarks marks a major milestone in the evolution of artificial intelligence hardware. By delivering a processor purpose-built to handle the unique, sequential control loops of autonomous AI agents, the company has successfully eliminated a major bottleneck in modern data center architecture. As high-volume deliveries to OpenAI, Microsoft, and Oracle expand in the second half of the year, the success of this custom chip will demonstrate whether a unified, full-stack hardware model can permanently redefine the economics of global computing.

Newsroom
Newsroom
Al Mahmud Al Mamun leads the TechGolly Newsroom team. He served as Editor-in-Chief of a world-leading professional research Magazine. Rasel Hossain is supporting as Managing Editor. Our team is intercorporate with technologists, researchers, and technology writers. We have substantial expertise in Information Technology (IT), Artificial Intelligence (AI), and Embedded Technology.