Running Qwen3-8B on DGX Spark

2026/09/05

My DGX Spark recently arrived! That’s some 40k HKD I hope well spent :)

I wanted to play with deploying a small LLM locally. I picked Qwen3-8B and tried BF16, FP8, NVFP4, SGLang, vLLM, and DSpark. This post is just a record of the numbers.

DGX Spark unboxing

Summary

All servers use a 32,768-token context limit and CUDA Graph. SGLang uses FlashInfer; vLLM uses its V2 runner, compiled kernels, and full plus piecewise CUDA Graph. Memory figures below are model-weight footprints reported by the server or checkpoint sizes, not total system memory usage.

SGLang BF16 SGLang FP8 SGLang NVFP4 vLLM BF16 vLLM FP8 vLLM FP8 + DSpark vLLM NVFP4 vLLM NVFP4 + DSpark
Checkpoint size 15.26 GiB 8.79 GiB 5.96 GiB 15.26 GiB 8.79 GiB FP8 + 4.42-GiB draft 5.96 GiB NVFP4 + 4.42-GiB draft
Runtime model memory 15.69 GB 9.18 GB 6.61 GB 15.27 GiB 8.88 GiB 11.11 GiB total 5.98 GiB 8.16 GiB total
KV cache dtype BF16 BF16 FP8 BF16 BF16 BF16 FP8 FP8
KV capacity about 463K about 510K about 1,058K 497K 549K 440K 1,135K 919K
Per-request context setting 32K 32K 32K 32K 32K 32K 32K 32K
ARC-C, 300 54.67% 54.00% 56.33% 54.67% 54.00% same FP8 target 56.33% same NVFP4 target
MMLU-Pro, 420 57.86% 59.05% 54.29% 57.86% 59.05% same FP8 target 54.29% same NVFP4 target
C-Eval, 312 80.45% 79.49% 75.96% 80.45% 79.49% same FP8 target 75.96% same NVFP4 target
GSM8K, 300 92.33% 92.67% 86.33% 92.33% 92.67% same FP8 target 86.33% same NVFP4 target
Natural-language decode, B=1 14.9 tok/s 29.2 tok/s 41.5 tok/s about 14 tok/s 23.18 tok/s 52.24 tok/s 39.86 tok/s 72.33 tok/s

DSpark uses exact target verification, so its quality is that of the FP8 or NVFP4 target rather than a separate model.

The 32K context is a deliberately identical server setting, not a datatype limit. Quantization mainly increases total KV capacity and therefore the number of concurrent long requests. It does not automatically change Qwen3-8B’s configured per-request context window.

The earlier draft also mixed checkpoint size with runtime allocation. With the same checkpoint, vLLM and SGLang load slightly different amounts because their quantized-weight packing, scale metadata, allocator, and kernel workspaces are different. The table now reports both quantities separately.

Prefill Throughput

These are aggregate input token/s with one generated token. Each value is BF16 / FP8 / NVFP4. Every batch has an SGLang row and a vLLM row.

B Engine Seq 128 Seq 512 Seq 2,048 Seq 4,096
1 SGLang 1,386 / 2,595 / 3,570 3,854 / 4,769 / 8,671 4,364 / 4,853 / 10,144 4,549 / 4,677 / 9,036
vLLM 1,287 / 2,219 / 2,571 3,914 / 6,343 / 8,696 4,841 / 7,232 / 11,195 4,604 / 6,866 / 10,504
8 SGLang 4,687 / 4,981 / 10,727 4,745 / 4,917 / 10,184 5,331 / 4,880 / 8,021 5,106 / 4,739 / 7,595
vLLM 3,706 / 6,215 / 9,184 4,855 / 7,634 / 12,481 4,920 / 7,483 / 12,114 4,719 / 7,057 / 11,167
64 SGLang 5,368 / 4,918 / 8,137 5,393 / 4,966 / 8,247 5,257 / 4,878 / 8,011 5,037 / 4,698 / 7,584
vLLM 4,823 / 7,715 / 12,112 5,045 / 7,922 / 13,297 4,917 / 7,502 / 12,265 4,707 / 7,025 / 11,229

FP8 prefill is not especially good in this stack: SGLang reports no GB10-specific W8A8 tuning configuration and uses a default kernel. NVFP4 completes FlashInfer SM121 autotuning and is much faster.

Decode Throughput

These are steady aggregate output token/s for 128 generated tokens per request, calculated from TPOT. Each value is BF16 / FP8 / NVFP4.

B Eng. Context 128 Context 512 Context 2,048 Context 4,096
1 SGL 14.92 / 29.46 / 41.69 14.86 / 29.22 / 41.52 14.68 / 28.39 / 40.63 14.43 / 27.40 / 39.59
vLLM 14.13 / 23.29 / 40.07 14.07 / 23.04 / 39.97 13.85 / 22.37 / 39.28 13.55 / 21.65 / 38.39
8 SGL 124.90 / 224.52 / 342.99 121.50 / 213.44 / 327.45 109.71 / 179.16 / 281.61 97.07 / 146.99 / 237.68
vLLM 117.92 / 185.23 / 334.98 112.62 / 174.46 / 318.66 94.40 / 138.44 / 254.57 76.21 / 106.71 / 194.14
64 SGL 808.05 / 1,383.42 / 2,125.81 682.30 / 1,048.48 / 1,664.87 419.48 / 530.54 / 927.21 275.64 / 318.61 / 579.99
vLLM 760.08 / 1,251.41 / 2,129.29 551.11 / 819.85 / 1,487.73 230.04 / 291.60 / 669.73 78.45 / 87.95 / 396.91

At batch one, FP8 is almost exactly 2x BF16 and NVFP4 is about 2.8x. The advantage shrinks with long contexts and large batches as KV-cache traffic and attention account for more work.

DSpark

This test uses vLLM, 20 natural-language prompts, greedy decoding, 256 generated tokens, and concurrency one.

Target Plain DSpark Speedup Plain TPOT DSpark TPOT TTFT change Accept length / rate
FP8 23.18 tok/s 52.24 tok/s 2.25x 43.05 ms 18.73 ms 66.5 → 124.5 ms 3.15 / 30.73%
NVFP4 39.86 tok/s 72.33 tok/s 1.81x 25.01 ms 13.44 ms 45.3 → 112.8 ms 3.06 / 29.49%

DSpark is useful for low-concurrency decode, but not for a saturated server. At batch 64 and context 128, plain FP8 reaches 1,098 tok/s while DSpark reaches 688; plain NVFP4 reaches 1,749 tok/s while DSpark reaches 623.

Compute and Memory Roofs

I measured both an 8192-cubed GEMM and a Qwen-like M=8192, K=4096, N=12288 projection. The best median roof for each datatype is 100.5 TFLOP/s BF16, 198.9 TFLOP/s FP8, and 234.5 TFLOP/s NVFP4. The local 2-GiB CUDA copy reaches 232.4 GB/s; NVIDIA advertises 273 GB/s LPDDR5X.

Phase / engine / datatype Effective model FLOP/s GEMM roof GEMM util Effective weight BW LPDDR util Bound
Prefill B64 L4096 / SGLang / BF16 88.4 TFLOP/s 100.5 TFLOP/s 88.0% 0.315 GB/s* 0.12%* compute
Prefill B64 L4096 / SGLang / FP8 82.5 TFLOP/s 198.9 TFLOP/s 41.5% 0.169 GB/s* 0.06%* compute
Prefill B64 L4096 / SGLang / NVFP4 133.1 TFLOP/s 234.5 TFLOP/s 56.8% 0.185 GB/s* 0.07%* compute
Prefill B64 L4096 / vLLM / BF16 82.6 TFLOP/s 100.5 TFLOP/s 82.2% 0.294 GB/s* 0.11%* compute
Prefill B64 L4096 / vLLM / FP8 123.3 TFLOP/s 198.9 TFLOP/s 62.0% 0.253 GB/s* 0.09%* compute
Prefill B64 L4096 / vLLM / NVFP4 197.1 TFLOP/s 234.5 TFLOP/s 84.1% 0.274 GB/s* 0.10%* compute
Decode B1 L512 / SGLang / BF16 0.229 TFLOP/s 100.5 TFLOP/s 0.23% 243.5 GB/s 89.2% memory
Decode B1 L512 / SGLang / FP8 0.451 TFLOP/s 198.9 TFLOP/s 0.23% 275.8 GB/s 101.0%* memory
Decode B1 L512 / SGLang / NVFP4 0.641 TFLOP/s 234.5 TFLOP/s 0.27% 265.7 GB/s 97.3% memory
Decode B1 L512 / vLLM / BF16 0.217 TFLOP/s 100.5 TFLOP/s 0.22% 230.5 GB/s 84.4% memory
Decode B1 L512 / vLLM / FP8 0.356 TFLOP/s 198.9 TFLOP/s 0.18% 217.5 GB/s 79.7% memory
Decode B1 L512 / vLLM / NVFP4 0.617 TFLOP/s 234.5 TFLOP/s 0.26% 255.8 GB/s 93.7% memory

Effective model FLOP/s counts dense Qwen operations completed per second, not physical low-precision instructions. For prefill, starred bandwidth is only the checkpoint-once lower bound and deliberately excludes activation traffic; compute is the relevant comparison. For decode, effective weight bandwidth is checkpoint size times token rate. The FP8 estimate above 273 GB/s shows the approximation error: checkpoint bytes are not identical to DRAM bytes and caching is ignored. The relevant, bound-setting columns are bold.

Takeaways