Artikel:unsloth.ai
Unsloth Benchmarks: Höhere Geschwindigkeit und verringerter VRAM-Bedarf bei QLoRA
Die Dokumentation von Unsloth präsentiert Benchmark-Ergebnisse für das Finetuning von LLMs auf NVIDIA-GPUs. Im Vergleich zur Kombination aus Hugging Face und FlashAttention-2 erreicht Unsloth bei Modellen wie Llama 3.1 und Llama 3.3 eine Verdopplung der Geschwindigkeit, spart über 70 Prozent VRAM ein und ermöglicht deutlich längere Kontextfenster.