Hybrides LLM-Inferenz-Setup: DGX Sparks für Prefill und M3 Ultra für Decode

X-PostXaden RyanDiskussion

Xaden Ryan plant ein Experiment zur Aufteilung von LLM-Workloads: Zwei DGX Sparks sollen den rechenintensiven Prefill-Schritt übernehmen, während ein Apple M3 Ultra mit 512 GB Unified Memory die Token-Generierung durchführt. Ziel ist es, lange Kontext-Prompts für GLM-5.3 zu beschleunigen.

Das Wichtigste

  1. Ryan setzt zwei DGX Sparks für die Prefill-Phase ein und plant, den berechneten KV-Cache an einen M3 Ultra (512 GB) für das Decoding zu übertragen.
  2. Im ursprünglichen Beitrag nannte Ryan Deepseek-V4-Flash, stellte im weiteren Verlauf jedoch klar, dass er das Modell GLM-5.3 wegen besserer Reasoning-Fähigkeiten testen möchte.
  3. Das Setup nutzt eine USB-Verbindungsmethode von ashxhart sowie einen Kernel-Patch von exolabs.
  4. Grund für das Experiment ist die unzureichende Prefill-Leistung auf dem M3 Ultra allein, wo ein 200k-Token-Prompt mehrere Stunden in Anspruch nahm.
  5. Offiziell erfordert die Softwarelösung laut Ryan M4-Chips oder neuer; er will die Ausführung auf dem M3 Ultra dennoch testen.

Warum das relevant ist

Bei extrem großen Kontextfenstern stellt der Prefill-Schritt auf Systemen mit Unified Memory einen Flaschenhals dar. Die Aufteilung in eine spezialisierte Prefill-Stufe und eine speicherstarke Decode-Stufe über USB könnte neue Wege für das lokale Hosting anspruchsvoller Modelle eröffnen.

Einordnung

Das Vorhaben verdeutlicht die aktuellen Limitierungen lokaler Hardware beim Betrieb moderner KI-Modelle. Apple Silicon bietet viel Arbeitsspeicher für Gewichte und Cache, verfügt bei langen Eingaben jedoch über begrenzte Rechenleistung für den Prefill. Ryan merkt selbst an, dass der praktische Erfolg des Setups offen ist und künftige Chip-Generationen wie der M5 Ultra derartige Behelfslösungen möglicherweise überflüssig machen.

Original-Post

Xaden Ryan

@XadenRyan · 5. September 2026

2x DGX Sparks secured. Goal is large prefill done on 2x DGX Sparks and then transferring over to M3 Ultra 512 GB to do decode for Deepseek-V4-Flash. Going to try getting this working over the usb connect @ashxhart has been working on and the kernel patch @exolabs got in. t.co/FTftBb6tLI t.co/UX4mrPyazu

84 Likes21 Antworten25 Lesezeichen8838 Aufrufe

Auf X ansehen
Weitere Posts im Thread (9)
  1. @0xkydo @ashxhart @exolabs On the website it says it only supports M4 or newer but I’m gonna test it anyway
  2. @GumbiiDigital @ashxhart @exolabs Sorry I was looking forward to meeting up but I was able to snag some tickets last minute to a van life festival in Oregon that I’ve been wanting to go to for years. Now I’m booking it 3000 miles across the country for this thing next week. I might be back in December tho
  3. @seanhighness @ashxhart @exolabs Maybe, I’m going to try some different experiments. Right now full GLM-5.3 prefill is to slow with a 200k context on the m3 ultra alone (it’s like hours for a single 200k token prompt). So I’m hoping to speed that up with the DGX Sparks. I have no idea if it’s actually gonna work
  4. @en4ble1337 @ashxhart @exolabs Yeah have to send computed kv cache over
  5. @ItsCuthulhu @ashxhart @exolabs I’ve been running it in my 128 GB M4 and the code is definitely cooking. I want something that does reasoning better. So I’m thinking GLM-5.3
  6. @ItsCuthulhu I’m also realizing now that I put DeepSeek on the original post and not GLM-5.3. Whoops.
  7. @ItsCuthulhu Sweet! Thanks!
  8. @spactastical No, It shouldn’t be
  9. @CMichaelGibbs I suspect with the M5 Ultra you won’t need to do this as as it can do much faster prefill
Ausgewählte Antworten (5)
  • @ItsCuthulhu @XadenRyan @ashxhart @exolabs t.co/Kndfupi6zN You could run 3 or more Qwen 3.8 Flash Next on all that. Have them always on. At this point, it's just a better model. Love Deepseek but this is the better model, smaller, cheaper, and you can have many of them running on Hermes. t.co/dWKFXIfb4L
  • @0xkydo @XadenRyan @ashxhart @exolabs does nvidia's pair work here?
  • @GumbiiDigital @XadenRyan @ashxhart @exolabs Sweet! I’ve got extra cables if you need them…when you headed down my way?
  • @ItsCuthulhu @XadenRyan t.co/F9QU2VPtMJ This may be the best bet for now!
  • @CMichaelGibbs @XadenRyan @ashxhart @exolabs Really keen to see how this works. My M5 Ultra should be here in about 4 weeks with 256. I have 4 sparks and am really keen to see if we can have the best of both worlds :D Keep in touch and let me know if you need/want another set of eyes.

Links und Tools aus diesem Beitrag

Zusammenfassung von KI erstellt (Gemini 3.8 Flash, 27. September 2026). Sie kann Fehler enthalten – maßgeblich ist die Originalquelle.

Inhaltlich ähnlich, ermittelt über die KI-Suche.

  • Repository:MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks

    GLM-5.3 Flash EXL3 auf 2x NVIDIA DGX Spark Kits

    Das GitHub-Repository von Mia's AI Lab bietet eine optimierte vLLM-Serving-Konfiguration für GLM-5.3-Flash mit 4-bpw-EXL3-Quantisierung auf zwei NVIDIA GB10 Systemen.

    431SternePython

    KI & AI· Tool

  • Repository:MiaAI-Lab/DeepSeek-v4-Flash-One-DGX-Spark

    DeepSeek-v4-Flash auf einer einzelnen DGX Spark betreiben

    Ein Open-Source-Launcher und Docker-Setup, mit dem DeepSeek V4 Flash 0731 (EXL3) auf einer einzelnen NVIDIA DGX Spark mit 128 GiB Unified Memory ausgeführt werden kann.

    328SternePython

    KI & AI· Tool

  • Video

    Video:Heavy Metal Cloud

    Nvidia DGX Spark im Praxistest: Setup, vLLM-Serving und Vergleich mit dem Mac Studio

    Heavy Metal Cloud testet die DGX-Spark-Varianten von Gigabyte und ASUS mit Grace-Blackwell-Architektur für lokales KI-Inference und Medien-Rendering. Neben einer Schritt-für-Schritt-Anleitung für Docker, vLLM und ComfyUI zeigt das Video Vor- und Nachteile gegenüber Apples Mac Studio.

    KI & AI· Anleitung

  • X-Post:Tech2Wild

    Optimizing Dual GPU Setups for Large Language Models

    This post explores a technical setup for running large language models using dual GPU configurations, specifically highlighting the use of hardware like the Spark and dual 4090s. The author provides a detailed guide on managing context windows and model instances to improve performance and efficiency for local AI workflows.

    52Lesezeichen8671Aufrufe

    KI & AI

  • Artikel:Elite Test Engineering

    DeepSeek V4 Flash auf zwei DGX-Spark-Nodes im Produktivbetrieb

    Ein Erfahrungsbericht dokumentiert das Aufsetzen des Modells DeepSeek-V4-Flash (284B MoE) über zwei DGX-Spark-Nodes mit Grace-Blackwell-Architektur mittels Ray und vLLM. Neben Hardware-Eigenheiten wie Unified Memory und NCCL-Bugs zeigt der Bericht, wie Multi-Token Prediction die Geschwindigkeit um 45 Prozent steigerte.

    KI & AI· Anleitung

Lassen Sie uns über Ihr Projekt sprechen

Standorte

  • Mattersburg
    Johann Nepomuk Bergerstraße 7/2/14
    7210 Mattersburg, Austria
  • Wien
    Ungargasse 64-66/3/404
    1030 Wien, Austria

Dieser Inhalt wurde teilweise mithilfe von KI erstellt.