3 min Devices

Everpure excels at eliminating AI bottlenecks

Everpure excels at eliminating AI bottlenecks

The MLPerf v3.0 benchmarks show that Everpure FlashBlade//EXA excels in model checkpointing and KV cache utilization. As key benchmarks for AI performance under realistic conditions, this demonstrates that training and inference are being continuously optimized.

The models tested are massive, ranging from 405 billion to 1.25 trillion parameters. With 30 data nodes, FlashBlade//EXA recorded 877.52 GiB/s write bandwidth (17.74 seconds) and 588.28 GiB/s read bandwidth (28.99 seconds). With ten data nodes, the write throughput remained at 327.65 GiB/s, which, according to the company, indicates predictable scalability.

KV cache as a new bottleneck

On the inference side, Pure reports 85,736 tokens per second for Llama 3.1 8B using storage only, 67,642 tokens per second for the same workload using storage plus memory, and 33,403 tokens per second for a 70B storage-only workload. Checkpointing was introduced for the first time in version 2.0 of the benchmark, with save and load tests for models ranging from 8B to 1T parameters.

“These results yet again demonstrate that FlashBlade//EXA delivers the highest performance and scale on the market, staying ahead of the increasing demands of AI workloads,” said Rob Lee, Chief Technology & Growth Officer at Everpure.

It’s worth taking a moment to consider the KV cache, which serves as the short-term memory for AI models. While the weights can already be loaded into memory, the KV cache is flexible and adjustable. For example, model providers can choose to offer AI models with compressed weights, which leads to less accurate results but more efficient use of storage and memory, whereas limiting the KV cache results in a less robust and long-lasting model memory. The balance between these two elements depends on the use case: with large amounts of data, the KV cache becomes more important; in situations where accuracy is the priority, the weights and their compression are more likely to be a limiting factor.

Separate metadata and data traffic

The test setup consisted of 30 FlashBlade//EXA blades with 120 DirectFlash Modules for metadata, in addition to 30 Linux/NVMe data nodes. Together, they provided approximately 866 TB of usable capacity under a single file system. For layout negotiation, the platform uses NFSv4.1/pNFS over TCP; for data transport, it uses NFSv3 over RDMA.

Pure presents this disaggregated architecture as a solution to storage bottlenecks in AI and HPC, claiming 10 TB/s within a single namespace and 3.4 TB/s per rack. EXA is thus the most powerful option above FlashBlade//S and FlashBlade//E. “EXA” stands for exascale, that is, the enormous scale of the AI infrastructures of hyperscalers and the largest AI model developers, OpenAI and Anthropic.

In this segment, Everpure faces several competitors, such as VAST Data and WEKA, as well as traditional players like Dell, HPE, and NetApp. While there are subtle differences in their respective focuses, all parties are primarily striving for efficiency, a goal that is becoming increasingly necessary given the sky-high costs of hardware and, specifically, memory.