Vollständiges Transkript anzeigen (2.781 Wörter)
Dylan Patel: Hello and welcome to SemiAnalysis and Supermicro's AI Lumina at Computex 2026. Thank you so much Vik for joining me.
Vik Malyala: Thank you.
Dylan Patel: So we're just going to have a little bit of a casual chat. You know we've had a couple of drinks already, so words are flowing a little bit. You know can you tell me a bit about what's behind me first of all?
Vik Malyala: So this one is our B300 HGX platform based on Intel Granite Rapids SP. So you can get 32 DIMMs. You know nowadays memory is not cheap. So effectively if you can get more DIMMs that you are able to kind of bring the platform.
Dylan Patel: Can I fill it, can I put only eight in it because it's too expensive?
Vik Malyala: You can put it, but then it's not going to work properly anyway. Yeah.
Dylan Patel: Amazing, amazing. So, you know we're at Computex. Supermicro is got an incredible presence here. I saw in the booth you guys even had a HELIOS rack for AMD MI450x. Can you tell me a bit about that?
Vik Malyala: You know, the the thing is it's like 7,000 pounds. So one good thing about, you know, HELIOS is that we are going to be the launch partner for AMD for that. And obviously the specs and everything, whatever is out there, out there, anything more than that, you know, we shouldn't be talking. But what what we see is that there is a good customer interest in them. And given the fact that it is kind of, you know, fall on to what people already have seen from the competitor, which is NVIDIA here. The GB200 and GB300 people have warmed up and it's been adopted. So people are a lot more open to considering HELIOS. So looking forward to seeing how it will be taken in the market.
Dylan Patel: Yeah, I think just to talk a little bit about the, the, the GPU in the system in ways maybe you're not allowed to. You know AMD's announcing it in a few weeks and it's very exciting. It's 72 GPUs. It's got more memory than Vera Rubin. It's got more, more memory, flops. It's got, it's better than Vera Rubin in most every way. It's a little bit later, I mean 3 to 6 months after. But you know, by being launched partner for HELIOS, you know, that's that's pretty exciting. There's, there's quite a bit of public customer attraction, Meta, Oracle for OpenAI and there's many other customers who are looking to adopt it. I think you know, the other thing that was really interesting to me is that it's twice as wide as a normal rack. I've seen the 19 inch racks, I've seen the OCP racks, you know, Supermicro offers both of those. But can you tell me a bit about, like, you know, this is like twice as wide. What does that, what does that mean? Like, you know, initially I used to worry about that in terms of, you know, how people would take it, given the fact that it's nonstandard for a lot of standards, OR-DW and all. But what I see is whether it's Vera Rubin or HELIOS, people cannot just go fit into the traditional data centers anyway, so they have to go with something that is purpose built or something that is going to be completely retrofitted to use. So the form factor doesn't seem to be a limitation by the customers. Probably it's to do with more like, you know, how we move it during the elevators and all that, but that is a different worry. I'm not I'm not seeing any kind of a pushback from customers because it is wider or bigger.
Vik Malyala: There's a, there's a data center that I know of specifically that is going to be deploying HELIOS, and they are redoing the doors so that they can wheel the HELIOS in. But you know, that's that's trivial compared to these, you know, many, many millions of dollars racks. You know, one of the other things that I thought you guys had at your booth that was really exciting was the Vera Rubin products. Vera Rubin, there's there's also the full scale rack, which I think is very exciting. It's it seems like it'll be a much smoother ramp than GB. You know, GB had a lot of problems because it was brand new, introduced a lot of new things. But Vera Rubin seems like it'll be a smoother ramp. Can you tell me a bit about how you think that's going to go?
Vik Malyala: So if you take a look at the GB200 versus GB300, GB300 is a lot smoother compared to 200, as all of us know. Vera Rubin, to the extent that we know and have seen, it definitely is much cleaner design. So I think it's going to be easier for people to adopt it. I think the limitation is going to be more in terms of data centers being ready to deliver that kind of power per rack than anything else. Cleaner design, for sure. And to answer your question on whether Vera Rubin, HELIOS and all, other thing that I see is that people are looking for alternative, right. And what we feel is that whether it's Vera Rubin or HELIOS, we want to be the first to market and enable the customers with Whatever the best that we got. And in that way, you know, we are quite excited about seeing both of them.
Dylan Patel: Yeah, I think that's something that's interesting in the startup ecosystem. I've seen, you know, there's a lot of AI hardware startups nowadays. They're doing a lot of interesting things. There's a lot of interesting rack architectures. A lot of interesting server architectures. One startup that I know of, Positron, is using Supermicro. There's other startups that I know of. I don't want to say the name because they're still in stealth, but they've got all these, like, weird, you know, cable, not just cable backplane, but cables inside the server, flyovers, routing all over the place to enable their unique network topology. And and Supermicro is the only vendor that's able to, you know, you know, sort of tier one OEM who's able to both have the supply chain deliver anywhere in the world, whether it be US or, you know, Asia or Europe or LatAm, even, sort of all over the world, Supermicro is able to deliver, but also with a very unique server architecture. Right. That's the that's the interesting thing. It's not it's not just cookie cutter. Can you talk a bit about how you guys are able to do that?
Vik Malyala: So a couple of different things. What I have seen is that when you talk about the PCIe accelerators, we have customers. I mean hopefully it's public, gets the likes of D Matrix and Positron as you mentioned. Those are serious rebellions. And we have plenty of others that, you know, are up and coming. These are the ones that are using this accelerators as a part of the system. And I have seen customers looking to have different PCIe topologies, right? Whether it's a single root complex or multi root complex. So we kind of enable that. And initially we started with both Intel and AMD as a across the architecture to support these accelerators. But now we have added the new ARM AGI CPU based platforms as well. So now people can have an option on the compute platform as well as what accelerators that they can support. Yeah as for the connecting these multiple accelerators is concerned, you know if we were to take a look at like a H100 NVL and other thing in the past, or its H200 NVL where 2 or 4 GPUs can be connected, some of them are looking at that architecture to come up with a bigger addressable pull, and many of them are still looking at bringing different amounts of memory and memory architectures into that to support different workloads.
Dylan Patel: I think one of the most interesting things about the Supermicro platform is, you know, we've seen all sorts of new hardware again adopted really rapidly by Supermicro, right. So, you know, we talked about, you know, CPU side, you mentioned ARM, AGI, CPU Phoenix getting adopted by Supermicro and being available. You know maybe first we saw that with AMD you guys are the fastest on AMD out of the tier one OEMs. And nowadays we see this with you know these ASICs. The other thing that's interesting is the connectivity options, right. You know historically Microchip and Broadcom had PCI switches. But now there's new PCI switches, whether it be Marvell or Sterra Labs, 256 lanes through 20 lanes. This enables people to make these really interesting topologies. And I've seen some Supermicro servers based on these newer switches that enable much larger domains of whether it be storage accelerators, NICs, a lot more customization that's available. Can you talk a bit about like the connectivity side of things? Right. Because it's not just about one server now. It's about a whole rack.
Vik Malyala: Yeah, so within a rack obviously a lot more topologies that coming, whether it's NVIDIA based platforms or AMD based platforms. You have the Ethernet which is 400 gig, 800 gig. Now we're talking about 1.46TB based on the Broadcom Tomahawk 6, right? But then to your point, we're working with Astrolabs to kind of look into how does it look at a rack level and multi rack level. And what is that advantage one can get out of it. And especially with the UA link coming up, that's also something that people are quite interested in seeing what it would be. The way I see it is that, you know, we have customers who are trying to bring accelerators integrated with certain NICs coming into different topologies. These are the ones that we are working with them to enable on a platform side. But the adoption point of view, you know, we need to still see because, you know, there are optimizations that can be done. But if it is specific workloads. Now the second thing people need to look at it is how am I going to have something that is going to work for more than one workload? So you cannot make it too specific because then you are stuck with this, right? Otherwise you have to make it more general.
Dylan Patel: Yeah, that's, that's what's amazing about Supermicro's product line. Right. There's, there's the general purpose CPUs. There's storage. Their systems that are sort of bring what you want. You know they have a lot of PCIe slots. So you can bring a bunch of network. You can bring a bunch of storage. You can bring a bunch of accelerators. Um, you know I think that's like what's, what's incredible about it. I guess, like to leave us off, right when I, when I go to Supermicro's booth, when I go to other people's booth, the the thing that I've seen that is like, you know, interesting is everyone wants to talk about AI and show their server off, but there's also a storage. I think that's where Supermicro really hit its stride. You go back two decades, you guys were leading in these two socket storage systems. So can you tell me a bit about, you know, what storage systems you guys are bringing, especially nowadays with storage being, you know, whether it be for KB cash offload or, you know, just incredible rise in cost of SSDs. Storage is something that people used to consider, you know, a back foot. And now it's it's it's the primary thing to focus on given for for a lot of users. So can you tell me a bit about your storage platforms.
Vik Malyala: So storage point of view what we are seeing is that initially it started with the form factors, right. You have U.2s that were popular or still are popular for that matter. Then you have E1.S and E3 drives mainly because of the density that we can bring in a one year or a two year factor, and one good thing that happened is, you know, if you take a look at NVIDIA's new GPUs, CPUs, their PCI Express Gen 6 and same is the case with what is the ARM AGI CPU and upcoming AMD next generation, right when it's platform that all PCI Express, Gen 6. So which actually gives incredible bandwidth that you can get out of these platforms. And we have come up with so-called Petascale solutions, which you can get a very good balance between amount of storage, the storage bandwidth as I/O bandwidth to do that. What we're also doing is, once you get into the PCI Express Gen 6, you have CXL 3.0 devices that can operate in that. That is also another exciting part that can be brought into storage platforms. And KV cache is where, you know, anyone, everyone offers now. That's also another thing that is bringing the value of this high capacity, high performance storage hardware is only one part of it, but what we are doing is working with the likes of Vast, WEKA, DDN, Qumulo, OSNexus and every one of these software vendors. We are able to bring different types of software defined storage into customer environments so that you don't have necessarily get stuck with a traditional storage. So super exciting times for that.
Dylan Patel: So, so Vik, the thing about what you mentioned is that PCIe Gen 6 is very difficult to cool, especially on the storage side, right? We saw this with Gen 5 systems. Initially, a lot of the systems that people released were overheating. The NAND was too hot, and so therefore the lifetime started to degrade or the controller was too hot and so the performance degraded. And so ultimately, this is a big challenge. Now with Gen 6, all these platforms turn, you know, Vera, Phoenix, AGI, CPU and eventually Intel's Diamond Rapids, all these Gen 6 platforms, you're going to have Gen 6 storage. How are you going to cool this, right? It's a, it's an incredible amount of density that you can offer. Can you tell me a bit about cooling? You guys were first on liquid cooling with xAI Colossus, the first large scale 100,000 GPU deployment with liquid cooling. And you've done many more since, and you were the first and everyone else has sort of followed since. So can you tell me about how you're going to deal with cooling these Gen 6 storage platforms?
Vik Malyala: So one thing that we are doing is to bring as many components under the liquid cooling loop. Initially, if we take a look at the compute platforms, we have done primarily the processor only and the VRMs, but then we added the memory as part of it. And some of these platforms were bringing even the E1.S for example. They have liquid cooling loop. We are able to do that as well. But the goal again here is to do that complete thermal analysis and see what are the things that we can do that we can be putting under a liquid cooling loop to cool it down? I mean, various things are happening. I mean, with the likes of Vera Rubin and HELIOS going to hit the market, the power delivery is going to be different, which also means that people are open to nonstandard form factors for the systems, not necessarily 19 inch, but it could be a wider form factor and what not. So that actually gives us some flexibility and how to place the components to make it cooler, better. Even the liquid cooling loop and the power delivery point of view we can do using the bus bar. So this way, a traditional space that is taken for power supplies and other things can be taken out, which also includes the airflow. So we are trying to find different ways to ensure that we have a product that can operate in a different operating environment sufficiently.
Dylan Patel: Thank you so much for explaining your platforms Vik. So SemiAnalysis X Supermicro at AI Lumina at Computex. So, thank you so much for having me. Cheers. You better kill it. Let's go. Gan-bei.