7 min Devices

CPUs are finally having their AI moment

CPUs are finally having their AI moment

Long dismissed as traffic cops for the GPUs doing the real work, new CPUs are set to significantly boost throughput for agentic AI tasks. AMD and Intel, wary of their Arm-based rivals, are well positioned. Long-term, however, a major mismatch between CPUs and AI accelerators remains unresolved.

We write this hot on the heels of AMD’s Advancing AI event, dominated by its rackscale Helios systems and strengthened partnerships from Microsoft to OpenAI, Anthropic and rival chipmaker Cerebras. Along the way, AMD gave a first, high-level look at Zen 6. It’s the first time a Zen generation ships in EPYC before any consumer-facing Ryzen chip. That choice is telling in and of itself. But why are CPUs suddenly deemed relevant for AI at all?

The CPU was always an AI accelerator

No AI workload runs without a CPU. It’s fairly well-known that CPUs route traffic to LLMs residing in GPU memory, but the workload is far bigger. CPUs turn inputs into tokens, decode images and video, orchestrate the KV cache (the LLM’s ‘short-term memory’). These are all jobs that balloon as context windows stretch toward a million tokens or more. Agents, a later arrival in the AI boom, then made the CPU indispensable in a new way. Ask an LLM to perform any task, and it does so via the CPU. The deterministic aids that made LLMs competent at code, for example, are conventional tools running on conventional CPU cores.

Why the obscure reputation, then? Well, scarcity in other areas of the IT hardware business. First Nvidia GPUs got hard to come by, then memory of virtually all shapes and sizes. Intel and AMD stocks stayed earthbound while Nvidia romped to 1, then 2, then 5 trillion dollars because CPU supply was never the limitation for AI build-outs, which means CPUs command none of the scarcity pricing that GPUs and HBM do.

How agents-per-watt changes the game

Agents-per-watt sounds like yet another metric in a world already far too full of them. We’ve had time-to-first-token, tokenomics, tokens-per-watt and even some out-there entries like ‘intelligence per Joule’. The metric is CPU-specific and AMD has put real numbers behind it with the new EPYC generation. Against Nvidia’s Vera, the Arm-based CPU that ships by default in every Vera Rubin rack, AMD claims 20 percent higher per-core performance, 2.2 times the throughput per socket and up to 2.8 times more agents per watt. Against its own outgoing Turin chips, Venice is said to deliver 1.8 times the tokens per second and 1.7 times the agents per watt.

Agents-per-watt is actually useful, although we suspect vendors will spin this thing their own way and independent benchmarks will be tough to calibrate to real workloads. AMD certainly has fun with the metric: since no standard benchmark for agentic workloads exists, it counts hardware threads as a proxy for agents, and Vera has not shipped yet, so every Vera figure is an AMD estimate based on Nvidia’s public materials. To say a grain of salt is needed is to undersell it.

Nvidia published its Vera white paper early last week, complete with SPEC results showing its 88-core design narrowly edging out a dual-socket EPYC. AMD, as it happily pointed out, got to respond using Nvidia’s own numbers thanks to this timing. Intel, meanwhile, was barely mentioned. It will become competitive by 2027, when Diamond Rapids arrives with built-in matrix acceleration as its one genuine differentiator. Regardless, AMD is shipping now. Venice SP7 is in production and heading to customers, with servers from the major manufacturers following in the fourth quarter.

The framing is clearly different from before. CPUs are now presented as enablers of AI rather than bottlenecks to it. The strongest evidence is that Nvidia feels the need to play the CPU game at all, positioning Vera as a sandbox CPU for AI agents. That push is proving a little frail next to its gigantic GPU advantage, however. Vera, even on paper, seems under-equipped when put next to the new “Venice” EPYC chips. Besides the metrics, Nvidia is ironically running into an issue all too familiar to AMD. Just as CUDA has been a moat for Nvidia on the GPU side (though AMD vehemently disagrees it is now), the maturity of EPYC as a powerful all-workloads processor is giving AMD a clear advantage over Nvidia.

A CPU renaissance

CPUs are having their AI moment, or will do so shortly. The logic is straightforward: the CPU generations arriving now are the first designed with agentic workflows fully known, and with a far deeper understanding of the anatomy of an AI task than their predecessors had.

Still, expect many more changes to occur in this hardware landscape. Google is hinting at yet another move with Frozen v2, for now a fledgling side project that would make the actual architecture of an LLM physical. The downside of such a design is a complete lack of flexibility. Nevertheless, the potential advantage cannot be denied. Then again, we already have very railroaded designs in Groq’s LPU and the wafer-scale units from Cerebras. These are impressive, and mature enough by now to get paired for serious AI workloads, but they remain just one step beyond being one-trick ponies for AI. Anyone calling them ‘agentic’ is missing the fact CPUs are needed for anything to do with agents. GPUs are far more maneuverable, but still need CPUs. Therefore, the latter chips are unlikely to be cut out whatsoever. You could imagine a future in which any number of XPUs survive, or one with just GPUs, or one full of Frozen-style hardwired designs. A future without the CPU is not on the menu.

The big problem: memory

One big problem remains. GPUs benefit enormously from their perceived importance, with heaps of investment flowing into future architectures and the elimination of their bottlenecks. The net result is that HBM, the GPU-side memory pool, is outpacing CPU-side DRAM on an already escalating speed trajectory. A single Helios rack carries 31 TB of HBM4 delivering 1.7 petabytes per second; the Venice socket next to it moves 1.6 terabytes per second. That gap widens every generation, and it is structural rather than something a firmware update will patch.

This means the balance will shift beyond 2026. AI infrastructure is incredibly lopsided right now: just look at the massive memory footprint compared to actual GPU utilization, with LLMs hogging HBM on GPUs whose compute they don’t need. This will shake out if and when AI development slows down or standardizes somehow. Hardware is just barely keeping up with architectural designs and not remotely keeping up with demand. That is temporary. Long-term, CPUs are bound to have a place, even with the aforementioned challenges.

Conclusion

CPUs are being steered towards AI, with specialized versions that are distinct from HPC or general-purpose computing. This isn’t so much a comeback as it is a moment to reappreciate what the CPU has quietly done for these workloads all along. It might take a while for investors to catch on. You simply don’t need as many CPUs as GPUs (Helios pairs 18 of them with 72 Instinct accelerators), and they are far smaller in die size, which means supply isn’t likely to be constrained the way it is elsewhere in the AI stack. The AI traffic cop, at any rate, is doing a far more complicated job than many have appreciated so far.