100 TB/s to fight GPU hunger: NetApp Novus is “fastest storage ever”

100 TB/s to fight GPU hunger: NetApp Novus is “fastest storage ever”

There are several reasons why AI isn’t succeeding in many organizations. Most of these reasons can be traced back to data. If the wrong data goes into a model, it won’t work, but the storage systems holding the data also need to be able to keep up with the demand from GPUs. Add to that the fact that, not surprisingly, there’s a trust issue surrounding AI, and it’s not easy for organizations to bring AI projects to a successful conclusion. NetApp aims to solve all of these issues. How does it plan to do that? Read on to find out.

NetApp CEO George Kurian, as always, was quite clear in his observations during NetApp INSIGHT, the company’s annual conference. AI and AI agents need access to an organization’s collective memory and context. Without that, AI projects will fail far more often than they succeed. He views the underlying NetApp infrastructure as that memory. In other words, organizations store their data on the storage systems. The context is then built up using the software stack on top of it.

All the announcements NetApp is making at INSIGHT this week revolve around NetApp’s ambition to grant AI that access. From NetApp Novus, the “fastest storage ever,” through NetApp’s agentic AI software stack to the hybrid cloud approach, and of course ONTAP as the foundation for virtually everything related to unified storage.

NetApp Novus

By far the most significant announcement at this INSIGHT conference is NetApp Novus. During conversations with him, Syam Nair, NetApp’s Chief Product Officer, still regularly refers to it as Project Novus, which tells us that this has been an internal project into which a great deal of time and effort has been invested. Novus is therefore unequivocally billed as “the fastest storage ever.” That’s not something Everpure will be happy to hear. Just earlier this month, that company proudly announced that its FlashBlade//EXA had achieved by far the best score in an MLPerf benchmark.

With Novus, NetApp says it is taking a giant leap forward in performance. During tests, it achieved read speeds of 100 TB/s, leaving all other systems on the market far behind. When we asked whether this also constitutes so-called sustained performance, that is, whether Novus can maintain this level of performance over a long period, the answer was affirmative.

Storage must keep pace with compute and networking

Incidentally, speeds like these are no luxury, as crazy as that may sound. Storage lags miles behind compute and network infrastructure when it comes to performance. The largest models, which must be capable of running in AI factories and GPU clouds, among other places, already utilize GPU clusters of 10,000 GPUs. This is expected to grow even further to 50,000 GPUs per cluster for future Nvidia architectures. If we assume 2 GB/s per GPU, that amounts to, coincidentally, 100 TB/s.

In other words, NetApp Novus is designed to ensure that storage is no longer the bottleneck for AI projects. All GPUs must be able to remain operational 100 percent of the time. Currently, that figure is sometimes no more than 5 percent. In other words, the GPUs are idle 95 percent of the time in these massive GPU clusters. That’s a waste of investment.

When it comes to performance, that 100 TB/s isn’t the whole story. NetApp also promises that, in principle, everything can scale endlessly. This is due to the disaggregation it applies to the data (see the next paragraph). It should be possible to scale toward zettascale with NetApp Novus in a single namespace. To put this into perspective, Kurian notes that this represents all the storage sold last year, across tape, HDD, and SSD, in a single environment.

Disaggregation taken to a new level

However, in our view, NetApp Novus’s performance isn’t its only noteworthy achievement. NetApp claims to have achieved this without using proprietary protocols. It’s all based on pNFS, or parallel NFS. This file system is also managed using standard tools.

NetApp Novus achieves its high throughput first and foremost by decoupling metadata from data. In other words, metadata is removed from the data path. The data can flow freely to the GPUs at nearly line rate. The metadata (filename, location, last access time, and so on), from which all kinds of insights and other information can be derived, is sent to metadata servers. It’s no coincidence that PEAK:AIO is set to be acquired by NetApp. According to Kurian, that company has “by far the best pNFS.”

What NetApp has done with Novus can be seen as an even deeper form of disaggregation. Whereas last year NetApp AFX decoupled compute and capacity (i.e., storage) from one another, it is now doing the same with data and metadata. AFX is much more focused on inferencing workloads, by the way, whereas Novus is designed for the massive clusters typically used to train models.

NetApp Novus runs ONTAP

Finally, regarding NetApp Novus, it’s worth noting that its operating system is ONTAP. In other words, it integrates seamlessly into the broader data strategy of companies already using ONTAP.

According to Jurgen Hofkens, CTO EMEA at NetApp, the fact that ONTAP is the operating system for Novus is a detail that shouldn’t be underestimated. It ensures that yet another separate silo doesn’t need to be set up. With Novus, neo-clouds and GPU clouds can add an extremely powerful storage layer beneath their GPUs, but thanks to ONTAP, customers can also easily adopt it. He expects this to significantly boost adoption.

For the sake of completeness, here’s a brief overview of exactly what NetApp Novus is as a product in practice: the NetApp Novus Data Director is the metadata software. It runs on Supermicro servers, while ONTAP data services are delivered by an all-flash AFF A90 array. According to Arindam Banerjee, Chief Platform and Technology Officer at NetApp, a software-defined version is also on the way, which will offer organizations more flexibility in how they actually want to roll it out.

More needs to be done

The AI data story isn’t just about performance. If you have no idea what data you have or which data you need to work with, you still won’t be able to accomplish much. During the previously mentioned AFX launch, NetApp already introduced the AI Data Engine for this purpose. It is essentially a data preparation metadata engine that determines which data is vectorized and thus fed into the models.

The AI Data Engine can now extract metadata from a wide variety of sources. In addition to NetApp ONTAP systems, it also works in combination with NetApp StorageGRID and non-NetApp arrays (provided they use standard protocols like SMB, NFS, and S3, of course). There are also quite a few AI-driven optimizations on the way for the AI Data Engine, designed to make it even better at searching for and analyzing metadata.

When we mention to Kurian during our conversation that this story sounds very familiar in the market (Everpure, among others, told us the same thing) and that things are starting to get a bit messy at that layer of the stack, he agrees that there’s a lot of noise right now. So what makes NetApp’s offering better? “We’re convinced that we can actually deliver on it. We’ve been storing customer data for 30 years. We can show them exactly how their data changes the moment it happens. Furthermore, we use open standards, technologies, and data formats, and we let all customers choose their own AI models.”

A custom data map for every organization

The ultimate goal of what NetApp aims to achieve in this part of the AI data stack is well articulated by Asad Khan, SVP and GM for AI at the company. According to him, unstructured data has still not been fully made transparent to date. A great deal of data has never been processed. Organizations often work with only a small subset. That’s understandable in itself, since data preparation takes time and money. Organizations are also potentially missing out on a great deal in this way.

According to Khan, NetApp’s goal is to provide every organization with its own map of all the data it possesses. On top of that comes MCP, and above that is the AI tooling that you can apply to the data. This can be purchased by organizations as a complete product based on StorageGRID. However, there is also the necessary flexibility.

Essentially, this allows organizations to build a knowledge graph of their own data. When asked how much flexibility NetApp can offer, Khan indicates that this is primarily found in the upper layers. The foundation, how NetApp builds the knowledge graph, is fairly opinionated, even though organizations can also work with other knowledge graphs. We do get the impression that Khan would actually prefer this not to happen. The main flexibility lies in the choice of what organizations use to ultimately query the data. Every organization can bring in the models it wants to use.

NetApp Console

To fully round out NetApp’s story regarding the (AI) data stack, we can’t overlook the NetApp Console. It’s not new, but it is receiving the necessary enhancements.

NetApp Console is the environment that allows you to manage your entire data infrastructure from a central location. Not just ONTAP, by the way, but also StorageGrid and E-Series, two product lines that haven’t received much attention in recent years but are still important to NetApp. The main message regarding Console during the conference is that it will include more autonomous functionality.

Cloud is part of (AI) data stack too

The goal of Console is to enable you to manage your entire hybrid NetApp environment. And when we say hybrid, we’re quickly referring to NetApp’s integrations with the hyperscalers. The main news in this area is that Oracle Cloud Infrastructure will also offer a native ONTAP solution, in addition to the existing integrations with Azure, AWS, and GCP. These are jointly designed storage services that are effectively owned by the hyperscalers and are also sold by them. The new service has been given the somewhat lengthy name Oracle Cloud Infrastructure NetApp Storage Service.

Incidentally, the value of the cloud services is quite significant, as we understand from Pravjit Tiwana’s responses to our questions. According to him, 55 percent of NetApp’s new cloud customers are entirely new to the company. So it’s certainly not just a matter of upselling from the on-premises offering. That makes the cloud offering a healthy part of the business.

The somewhat older Cloud Volumes ONTAP (which allows customers to extend ONTAP from on-premises to the cloud) is also performing well, he says. NetApp is still investing in it, he notes, even though it has received somewhat less attention in recent years. He mentions that there are still thousands of customers. These include customers who use it for disaster recovery, as well as those who prefer to manage everything themselves, for example, from a cost-optimization perspective.

Keystone Sovereign

The last topic we want to discuss in this article is certainly not the least important, at least if you live in Europe (or elsewhere outside the U.S.) and place a high value on sovereignty. Keystone, NetApp’s Storage-as-a-Service offering, is getting a sovereign variant.

Whereas sovereignty used to be primarily about where your data was stored, we’re increasingly seeing other factors come into play as well. Think about who has access to a service, the telemetry that’s exchanged, and, in general, transparency regarding how data flows through networks.

With Keystone Sovereign, NetApp commits to hosting the service’s metadata within a sovereign cloud in Europe. All operational matters, such as support and escalation, are also handled within and via Europe. All access and management processes are under European control, and clear documentation is available.

Especially given the growing popularity of StaaS (at NetApp, the service is growing at double-digit rates every year), the availability of a sovereign version of Keystone for organizations is good news for our region.

Conclusion: The snowball is getting bigger

In just a few years, NetApp has made enormous strides in offering unified storage, managing it, and addressing the specific challenges that AI poses for the storage layer. Whereas a few years ago we mainly wrote about a new line in one of the company’s all-flash appliance series, that changed fundamentally last year with NetApp AFX.

With NetApp Novus, the company is taking things a step further. On the one hand, it’s a logical next step when viewed from the perspective of disaggregation, but it still needed some development. The company has done just that. Memory and context are becoming increasingly prominent in NetApp’s stack.

The main question right now is whether the stack NetApp has developed contains the right components. All the pieces seem to be in place. From data preparation through management and hybrid cloud options to the fastest storage currently available, it’s all there. The path is clear, you might say, provided, of course, that interest in AI remains strong.

However, there’s a chance that all these components still won’t set NetApp apart and that ultimate success will hinge on something else. Kurian spoke of trust. That might well be the key. And in that case, it’s certainly an advantage that NetApp has been around for so long. But virtually all other storage providers have been around just as long. So, delivering on your promises is more important now than ever. That builds trust and, ultimately, leads to success.