Vollständiges Transkript anzeigen (2.931 Wörter)
All right, welcome to another session of DGX Compare. It's been a while since we did two-node, um, but after 17 hours of configuring, um, I finally got DeepSeek V4 working on two nodes. Um, but a little caveat, it crashes a lot. Um, we're going to walk everybody through it. So, yeah, DeepSeek V4, it's two hill, 284 billion parameters, quantized down FP8, uh, to I think 13, 13 active, uh, parameters mixture of experts, and, uh, it's installed on a cluster of two DGX devices. The primary is the left right there, which is the DGX Spark and the worker node is the, uh, Acer GN100. They're both clustered using, um, QSFP4, uh, connected, let me see, seven, yeah, connected seven, and, uh, yeah, it's, they're connected and, uh, running at 200 gigabytes per second, so it's not that much further off from the 273 megabytes per second, uh, memory bandwidth, uh, for a single node, um, but it's still a little slower, but at least we can get the weights in, right? Again, this is a 284 billion parameter model, uh, clustered on two devices. Um, here we have the, um, the system memory for, uh, the, the Spark, it's running pretty much like 99%, 127 gigabytes allocated entirely to the, the weights and the model itself, um, and, uh, it's actually running off of a Mac, but then I'm doing activity monitor, um, remotely for the GN100, and this GN100, it's around 95%, 120 out of 128 gigabytes total. Um, I don't want to be, I don't want to touch the AC, uh, the Acer, because any kind of, let's say, even, even loading of a browser on the actual, uh, node itself, uh, takes up memory, so I'm trying not to. Um, okay, so let's go ahead and, uh, run the, the actual query that we usually have. Tell me all about Tyrannosaurus Rex. Uh, Tyrannosaurus, misspelled on purpose, and let's see how it runs. So pay attention to the GPU utilization, again, primary node, and this is the secondary node, this is the worker, this will be doing a lot of the work. All right, and we're off. So this is a mixture of experts model, you already saw both GPUs are highly active, once at 95% down here at the worker, and it's, uh, 93%. Um, thinking's done, and it's actually throwing, uh, information out now. Let's see. Look at the, the activity there. I think the first pink one was the initial spike when we queried, and then look at all those cores running. Uh, I see at least three at 100%. It's crazy. Um, but it's running, and the weights are actually throwing on the screen. Um, pretty solid data that's coming out, solid information, I mean. Um, prefill was fast. Uh, I'll have to watch the video again to see how, how long it took, but it's no more than five seconds, um, and, and it's still going. I would say this is probably 15, 15 tokens per second. So, it's usable, not ideal, but usable. I'd run this on the Mac Ultra, but we only have 96 on the Ultra, and this won't even fit. So, to get this model running on a fast device, you'd really need to have at least, at least 256 minimum. Okay, so, it's gone ahead and completed, GPU utilization went down to around zero. Um, so that's solid, hasn't crashed yet. Let's do a follow-up. Uh, what did T-Rex eat, and was this a scavenger or a hunter? Again, watch the GPU usage. So, boom, they both spiked, and, uh, yeah, prefill's just a second, second, and a couple two seconds, tops. And, uh, yeah, it's still going. I would say this is probably 15, 15 tokens per second. So, it's usable. Not ideal, but usable. Try Ceratops. Wow, I didn't know the T-Rex was a Ceratops, and they ate other dinosaurs. Okay. Oh, wow. T-Rexes are cannibals. I, I never realized that. Hmm, okay. So, we're still running stable, hasn't crashed yet. But, based on the forums and based on people who've actually gotten this up and running, the moment the context window gets a little bigger, um, it starts to get suspect, it starts to get problematic, and I get it, right? We're running two devices, this is two CPUs, um, at 128 gigabytes each, and we are running at 99%, right, of system memory, um, so it's really using the entire 256, or let's say 250 or so, 248, because you have some space for the operating system and, and the OpenWeb UI. Um, but you're very minimal on KV cache. In fact, let's talk a little bit about the, the setup here. Um, while that rests a little bit, um, these are all the parameters, I'll show this on the description, um, but we're running tensor parallel, size two. Um, we're using Ray as the coordinator, um, and again here, FP8 for the KV cache type, um, and, uh, we actually, uh, used 80% of memory utilization, um, so that we have probably between 10 and 13% uh allocated to the KV cache, which is not a lot. So the moment you get into a long conversation, that's where it gets, uh, tricky, right? Um, but this is an optimized model for CUDA, um, and again, it took 17 hours to get this working properly because we had to reconfigure over and over and over and over again to get it ideal. We initially started with NVFP4, and then had it quantized down to, sorry, it was quantized down to NVFP4, but then that was just not running properly. Distributed, um, kind of setups like this aren't ideal for, um, even, even the playbooks provided by Nvidia. So you have to do a lot of custom work, and, uh, what we found out was FP8's actually the, the, the, the best, uh, to use, uh, the least is running. Again, see, it's actually running here, right? Um, and another kind of note, we're actually only running, we've only tested this running, uh, single user. Um, the moment we try, we haven't tried double user yet or multi-user, but I would, I would say that if we use multi-user, uh, you would crash. Um, but this is kind of just a proof of concept or, you know, at least loading up the nodes, loading up the model into the nodes. Uh, let me check the heat. Yeah, it's getting, Yeah, it's really hot. Temperature-wise, we're still good. Um, you have five fans actively running, um, but we're only at 87% 80-degree, 87 degrees. So not too bad, but it is getting warmer. So it's still running, yeah. Um, I'd say between 10 and, and 15, um, in terms of token speed. Now, we installed Tools, or Tools is available for this model, so let's give it a try. Um, I'm going to use a new chat, but I'll open up a, I'll do this later where we, um, we actually try to piggyback off of an existing conversation. So let's use our typical file. Um, I have a, a CSV file that's 17k, uh, 17.3k, uh, and, uh, it's got 298 rows and, uh, sorry, 209 rows and then nine columns of car information, historical automobile information, and what I'm going to do is have, DeepSeek analyze the file. So, as you've seen, it spiked up again, let's take a look at the the cores here. Um, there we go, spiked again. Um, it's a thinking model, so it's actually doing a lot of uh work right now in terms of determining and assessing the contents, probably using Python as an algo. Um, it said, let me first look at the full contents of this file to get complete data. Um, and again, this is the worker node, it's doing a lot of the work. Okay, there we go. Um, wow, okay, cool, it's loaded up properly, and it properly identified 398 rows, uh, as well as nine columns, and it's providing us all the data there for statistics. Okay, those look accurate. Um, I'm familiar with this dataset, and it looks good so far. We're looking for displacement. Yep, there we go, 68. The reason why I track this is because sometimes I've noticed when, um, models get busy, they start to spit out incorrect data. Um, I have several other videos where, uh, you know, the dataset that the user, that the model used, was only a subset of the actual dataset because it couldn't fit into memory. Uh, and this time around, DeepSeek's actually able to fit everything into memory correctly and spitting out correct data. Oh, wow. I've actually never seen this. Notable Observations. Okay, cool. Look at that. Notable observations. Okay. Oh, cool. Did that. So, yeah, it's handling RAG pretty well, or at least, uh, tooling. Um, that there goes, it's shut down. I mean, not shutdown, but it's gone ahead and reverted back to, to another CPU, sorry, um, low GPU, um, so that's solid, it hasn't crashed yet. Let's do a follow-up. Uh, can you show me summary statistics, like, uh, mean, median and standard deviation. Okay, let's give it a look. See how like the worker node went first, and then, uh, the primary node we're serving it, and that's because, um, the systems are already kind of allocated the router is actually allocated most of the, the compute, uh, to the worker node. Okay, still okay. It is really hot where I'm at. I have fans running here, running under the table, and running on the side. It's just really hot out today. Um, less because of the, um, the system, but more towards it's just hot in LA today. Um, all right, it's finally loaded up. Based on the contact provided, I can only see three individual car records from the Auto MPG dataset, which is a very small sample. See there we go. I can only see three individual car records from Auto MPG, which is a very small sample. This is exactly what I was talking about earlier, right? So, there are 398 cars, but when it's doing the secondary search for mean, median and standard deviation, um, it can't load the, the space and memory to be able to assess all that, because it's a lot, right? You're actually doing calculations for, for mean and median and standard deviation, which, which could be in, um, intensive. Um, and so it's only using a sample of it. But, at least it's telling me that, um, it's only able to use three. I was using a different model several days ago, and it didn't even self-spell that out, so it started giving statistics like this, but, and, and you're assuming if you don't know the dataset that you're working with the full set, it's only a sample. So, I'm really happy with DeepSeek's done here, um, in that they haven't actually, or they're being transparent about what they're showing. Um, in a few days we'll deploy this to Nebius, um, where we have more, um, I guess memory to work with and, and, and, and CPU to work with, um, but, uh, that's not DGX Compare anymore. I guess that's DGX A300 or A100 Compare. Um, but either way, um, I'm pretty happy with DeepSeek. It's a lot more stable than NemoTron. NemoTron 120 was my favorite, um, but it hasn't provided this kind of accuracy. Now, again I'm not happy with, um, not showing us, uh, the correct information, but, um, what I'm going to do is, we're gonna, we're gonna copy this query, or prompt, we're going to create a new chat, and then we are going to use the same file, let's actually see what happens when we try to upload a file. It should just be working off of cache, because it's actually loaded, pre-loaded in, but let's see what it does. And this time around, I'm going to go straight to querying that, um, that same query that we had earlier that only used three cars. And then, let's see what the results are going to be with a clean context window. All right. That was weird. I did say show me some summary statistics, and it's telling me it's, uh, data provided for the Chevrolet Caprice Classic across three model years. Again, it's hallucinating right now, or not hallucinating, but it's just not getting us the information that we need. Um, it's getting us information that it that the, based on what it's capable of doing, but not what we're actually looking for, right? So this is going to be one of those problems with not enough, uh, KV cache or memory allocated, uh, because you guys know, right, that when you're actually working with, um, uh, context, or, or, or, or, or, retrieval granted log, generation, or generated files to work with, with tooling, um, having a lot of memory available is key. And, you know, this is a 254 billion parameter model and we have two, 256, right, of of RAM, so that's going to be tough. Um, what exactly is DeepSeek, sorry, 284, compressed down, so anything here that talks about, speculative decoding, we have that on, um, we have that all. I mean, more research on, on, on, what the preferred, um, amount of KV cache is. But it's usually 40%, right? So, um, to be comfortable, you should really have 40% allocated, so let's say you run 256, you should at least have, I don't know, um, 96, available, uh, which is not realistic. But again, we're using, 13 parameters active, right, um, and it's fast, but again, you just got to have that kind of, um, available memory in order to do any kind of, uh, history or, or, context. Um, let's do another query, to look at more token speed. Um, I want to go back to our initial chat for the T-Rex, uh, and say how did T-Rex hunt for food? What I'm trying to do now is just generate a lot of, um, context into a single conversation. Um, and let's see if we can get the system to crash. So far, so good. Again, you know, it's been a blue line here for utilization, and system memory, for the integrity. Or not entirely, but for the for most of this, uh, this experiment. And, it's been pretty stable. But look at that, that's crazy. That's just crazy. Oh. That one, or no, sorry, I thought it crashed, it didn't, it just went to screen saver. Phew. And by the way, if, if that, that actually, if that crashed, um, this wouldn't be running anymore. So, so far, so good. Um, 10 to 15, um, tokens per second, pretty solid. Um, let's do a new chat. I wonder if this works dictate? Nope. Um, tell me about elephants. Same thing, pretty fast, pretty fast prefill, that was under a few seconds. And it's straight to decode. You know, other things that I read about in, um, in the forms a lot is that a lot of people are comparing, um, the Mac Ultra, and we have many tests on Mac Ultra as well, comparing the decode speed. Um, again, the DGX is are running at 273 and, and the Mac Ultra, M3 Ultra is running at 819 or 827 or something like that. So you're talking three or four times the speed, and that's exactly true. But the problem is when you're running large models like this, um, it's mostly compute, and the CUDA stack is what's helping these large models, uh, be able to do prefill, um, much faster. So on the Ultra, you'll do a, you'll do a query like that or a prompt like that and it'll, it'll take a couple seconds before any results are spit onto the screen, and they're fast. Um, but if you took it, if you take a look at time to first token, or time to first, um, yeah, time to first, um, yeah, time to first token, um, it's much faster here, it's a second or two on the DGX is, and on the Max, it could be a while. And, I've seen clients where they, they run on an Ultra, and they, they do a query, and they're like, "what's going on?", they think it's crashed, right, and it's not. It's just, it's waiting. So, I'd much rather have slow tokens come out, um, but you have something at least coming out, um, compared to you're waiting and something fast comes out, right? Because maybe at this point, um, it could take sometimes 60 seconds when you're doing crazy queries. Um, and at that point, people are already decided, "Ah, it's crashed," and they don't want to work with it anymore. They've actually gone and done another query. So, all right. Um, look at this, a lot of, a lot of playing and nothing's crashed. And, uh, we've slowed down now, but so far, so good. Uh, we'll end it at that. We'll run this model, uh, for a day, and then, um, we'll give a further report tomorrow. So again, DeepSeek V4, 280, um, 284 billion parameters, quantized down to, um, 13B running FP8, on a clustered DGX, or DGX system, uh, with, uh, DGX Spark as the primary and the worker node is an, uh, an Acer GN100. Thanks a lot.