AMD Brings EPYC Venice to Every Layer of Agentic AI Infrastructure

AMD has released a new evaluation of its sixth-generation EPYC processors aimed at a scenario where artificial intelligence (AI) is no longer limited to training models. The company argues that infrastructure built for agentic AI needs CPUs capable of handling everything from databases and web services to inference, analytics, and high-performance computing, and positions the EPYC 9996 as its highest-capacity offering in this generation.

AMD EPYC Venice in 20 seconds

  • AMD evaluates the EPYC 9006 “Venice” family across enterprise, cloud, AI, and high-performance computing workloads.
  • The EPYC 9996 reaches, per AMD’s internal estimates, up to 2.24 times the platform performance of NVIDIA Vera on SPECrate 2026 Integer.
  • Against the Xeon 6980P, AMD reports gains of 1.8 to 3.13 times across several scientific workloads.
  • The family spans configurations from 8 to 256 cores and also includes rack-scale nodes for AI.
  • “Venice” is already in production, and AMD expects the first cloud provider deployments by late 2026.

AMD’s thesis starts from a shift in how infrastructure gets sized. According to the company, two years ago AI system planning was heavily driven by training needs. Since then, inference has taken on more weight, and now AI agents add another variable: a single request can trigger retrieval queries, tool calls, code execution, and result generation.

That workflow can change both how much compute is needed and where each part of the task runs. For AMD, that makes infrastructure designed around a single workload profile less well suited to the job.

The company’s new evaluation is laid out in a white paper dedicated to the AMD EPYC 9006 “Venice” processors, comparing general-purpose, enterprise, cloud-native, AI, and high-performance computing workloads. AMD is trying to show how the family behaves across a wide range of scenarios, rather than limiting the comparison to a handful of tests.

The EPYC 9996 measured against NVIDIA Vera and Intel Xeon

One of the figures AMD highlights comes from SPECrate 2026 Integer. In the company’s internal estimates, a two-socket system with the EPYC 9996 delivers platform performance of roughly 2.24 times that of a two-socket platform based on NVIDIA Vera. The figure comes from preliminary results and AMD’s own engineering projections, so it isn’t equivalent to an independent or final result.

The same comparison puts the sixth-generation EPYC’s per-core performance at roughly 1.2 times that of NVIDIA Vera. AMD notes that the Vera figures used in this comparison are also based on the configurations and estimates used for the analysis.

In enterprise and cloud-native workloads, AMD reports gains of between 2.4 and 3.7 times in certain tests, including Java application server, OpenSSL, MongoDB, Redis, NGINX, and transaction processing. In its internal tests, for example, the 256-core EPYC 9996 reaches up to 2.9 times the Redis throughput of the Xeon 6980P, and up to 3.7 times for NGINX.

The figures depend on each workload and the configuration used. For MongoDB, AMD reports up to 3.5 times the Xeon 6980P’s performance, while for OpenSSL it puts the gap at up to 2.5 times. For MySQL, using a workload derived from TPC-C and run with HammerDB, the company reports up to 2.6 times the performance of AMD’s processor versus the Xeon 6980P.

These tests were run by AMD using specific memory, software, BIOS, and operating system configurations. The company itself cautions that results can vary and that workloads derived from benchmark standards shouldn’t be confused with official results from those benchmarks when they don’t meet their publication requirements.

From scientific computing to 100 kW racks

The agentic AI scenario AMD describes isn’t limited to traditional server applications. Part of the white paper also covers scientific computing workloads such as GROMACS, NAMD, Quantum ESPRESSO, and WRF.

In GROMACS, for example, AMD records 9.195 nanoseconds per day for a two-socket EPYC 9996 system with MRDIMM memory, compared with 2.937 ns/day for the Xeon 6980P in the configuration being compared. In Quantum ESPRESSO, run time drops from 2,864.897 seconds on the Intel system to 1,594.417 seconds with two EPYC 9996 processors and MRDIMM. These are results from AMD’s internal tests and depend on the specified configurations.

The company also frames the problem from the perspective of a rack limited to 100 kW. In its model, a rack using two-socket EPYC 9996 systems reaches roughly 3.4 times the performance of an NVIDIA Vera-based configuration across a mix of six general-purpose workloads. AMD clarifies that this is a preliminary analysis based on engineering measurements or projections from July 2026.

The difference between per-core performance and core density stands out as one of AMD’s arguments here. A processor with high per-core performance can handle latency-sensitive workloads, while a higher core count allows more simultaneous tasks to be packed in. The company presents its lineup as a way to choose different profiles without having to use the same configuration across all of its infrastructure.

The EPYC 9006 family covers, according to AMD, four processor families built on a common software base. The lineup ranges from 8-core configurations aimed at certain edge deployments up to 256-core processors and CPU nodes designed to accompany rack-scale AI systems.

AMD also says “Venice” is already in production. Major OEMs plan to launch platforms based on these processors, and the company says the first cloud providers will begin deploying them during the second half of 2026.

The strategy’s relevance lies in the role CPUs can play within an increasingly distributed AI architecture. An agent doesn’t necessarily run all of its work on a GPU: it can query a database, access services, process information, execute code, or coordinate other tasks before and after an inference step. For AMD, that diversity justifies having different CPU profiles within the same infrastructure while keeping a common software base.

These performance figures should be read alongside AMD’s specific methodology and configurations. Some are engineering estimates and others come from internal tests, so the multipliers don’t automatically carry over to any given server or application. The “Venice” pitch rests precisely on widening the range of scenarios evaluated, from enterprise services to scientific computing and systems built for AI workloads.

FAQ

Which processor does AMD highlight within EPYC Venice?

The EPYC 9996 is the highest-capacity model cited in the evaluation, with configurations of up to 256 cores per processor. AMD also offers other profiles for different types of deployments.

What performance does AMD claim versus NVIDIA Vera?

In a SPECrate 2026 Integer comparison, AMD estimates roughly 1.2 times more per-core performance and 2.24 times more platform performance for a two-socket configuration versus NVIDIA Vera. These are preliminary internal estimates.

What difference does AMD report versus the Intel Xeon 6980P?

The company reports gains of between 1.8 and 3.13 times across several scientific computing workloads. In specific enterprise applications such as NGINX, MongoDB, or Redis, it also reports differences of up to 3.7, 3.5, and 2.9 times, respectively.

When will servers with EPYC Venice arrive?

AMD says “Venice” is already in production, that major manufacturers are preparing their platforms, and that the first cloud providers will begin deploying these processors during the second half of 2026.

Scroll to Top