My DGX Spark recently arrived! That’s some 40k HKD I hope well spent :)
I wanted to play with deploying a small LLM locally. I picked Qwen3-8B and tried BF16, FP8, NVFP4, SGLang, vLLM, and DSpark. This post is just a record of the numbers.

Summary
All servers use a 32,768-token context limit and CUDA Graph. SGLang uses FlashInfer; vLLM uses its V2 runner, compiled kernels, and full plus piecewise CUDA Graph. Memory figures below are model-weight footprints reported by the server or checkpoint sizes, not total system memory usage.
| SGLang BF16 | SGLang FP8 | SGLang NVFP4 | vLLM BF16 | vLLM FP8 | vLLM FP8 + DSpark | vLLM NVFP4 | vLLM NVFP4 + DSpark | |
|---|---|---|---|---|---|---|---|---|
| Checkpoint size | 15.26 GiB | 8.79 GiB | 5.96 GiB | 15.26 GiB | 8.79 GiB | FP8 + 4.42-GiB draft | 5.96 GiB | NVFP4 + 4.42-GiB draft |
| Runtime model memory | 15.69 GB | 9.18 GB | 6.61 GB | 15.27 GiB | 8.88 GiB | 11.11 GiB total | 5.98 GiB | 8.16 GiB total |
| KV cache dtype | BF16 | BF16 | FP8 | BF16 | BF16 | BF16 | FP8 | FP8 |
| KV capacity | about 463K | about 510K | about 1,058K | 497K | 549K | 440K | 1,135K | 919K |
| Per-request context setting | 32K | 32K | 32K | 32K | 32K | 32K | 32K | 32K |
| ARC-C, 300 | 54.67% | 54.00% | 56.33% | 54.67% | 54.00% | same FP8 target | 56.33% | same NVFP4 target |
| MMLU-Pro, 420 | 57.86% | 59.05% | 54.29% | 57.86% | 59.05% | same FP8 target | 54.29% | same NVFP4 target |
| C-Eval, 312 | 80.45% | 79.49% | 75.96% | 80.45% | 79.49% | same FP8 target | 75.96% | same NVFP4 target |
| GSM8K, 300 | 92.33% | 92.67% | 86.33% | 92.33% | 92.67% | same FP8 target | 86.33% | same NVFP4 target |
| Natural-language decode, B=1 | 14.9 tok/s | 29.2 tok/s | 41.5 tok/s | about 14 tok/s | 23.18 tok/s | 52.24 tok/s | 39.86 tok/s | 72.33 tok/s |
DSpark uses exact target verification, so its quality is that of the FP8 or NVFP4 target rather than a separate model.
The 32K context is a deliberately identical server setting, not a datatype limit. Quantization mainly increases total KV capacity and therefore the number of concurrent long requests. It does not automatically change Qwen3-8B’s configured per-request context window.
The earlier draft also mixed checkpoint size with runtime allocation. With the same checkpoint, vLLM and SGLang load slightly different amounts because their quantized-weight packing, scale metadata, allocator, and kernel workspaces are different. The table now reports both quantities separately.
Prefill Throughput
These are aggregate input token/s with one generated token. Each value is
BF16 / FP8 / NVFP4. Every batch has an SGLang row and a vLLM row.
| B | Engine | Seq 128 | Seq 512 | Seq 2,048 | Seq 4,096 |
|---|---|---|---|---|---|
| 1 | SGLang | 1,386 / 2,595 / 3,570 | 3,854 / 4,769 / 8,671 | 4,364 / 4,853 / 10,144 | 4,549 / 4,677 / 9,036 |
| vLLM | 1,287 / 2,219 / 2,571 | 3,914 / 6,343 / 8,696 | 4,841 / 7,232 / 11,195 | 4,604 / 6,866 / 10,504 | |
| 8 | SGLang | 4,687 / 4,981 / 10,727 | 4,745 / 4,917 / 10,184 | 5,331 / 4,880 / 8,021 | 5,106 / 4,739 / 7,595 |
| vLLM | 3,706 / 6,215 / 9,184 | 4,855 / 7,634 / 12,481 | 4,920 / 7,483 / 12,114 | 4,719 / 7,057 / 11,167 | |
| 64 | SGLang | 5,368 / 4,918 / 8,137 | 5,393 / 4,966 / 8,247 | 5,257 / 4,878 / 8,011 | 5,037 / 4,698 / 7,584 |
| vLLM | 4,823 / 7,715 / 12,112 | 5,045 / 7,922 / 13,297 | 4,917 / 7,502 / 12,265 | 4,707 / 7,025 / 11,229 |
FP8 prefill is not especially good in this stack: SGLang reports no GB10-specific W8A8 tuning configuration and uses a default kernel. NVFP4 completes FlashInfer SM121 autotuning and is much faster.
Decode Throughput
These are steady aggregate output token/s for 128 generated tokens per request,
calculated from TPOT. Each value is BF16 / FP8 / NVFP4.
| B | Eng. | Context 128 | Context 512 | Context 2,048 | Context 4,096 |
|---|---|---|---|---|---|
| 1 | SGL | 14.92 / 29.46 / 41.69 | 14.86 / 29.22 / 41.52 | 14.68 / 28.39 / 40.63 | 14.43 / 27.40 / 39.59 |
| vLLM | 14.13 / 23.29 / 40.07 | 14.07 / 23.04 / 39.97 | 13.85 / 22.37 / 39.28 | 13.55 / 21.65 / 38.39 | |
| 8 | SGL | 124.90 / 224.52 / 342.99 | 121.50 / 213.44 / 327.45 | 109.71 / 179.16 / 281.61 | 97.07 / 146.99 / 237.68 |
| vLLM | 117.92 / 185.23 / 334.98 | 112.62 / 174.46 / 318.66 | 94.40 / 138.44 / 254.57 | 76.21 / 106.71 / 194.14 | |
| 64 | SGL | 808.05 / 1,383.42 / 2,125.81 | 682.30 / 1,048.48 / 1,664.87 | 419.48 / 530.54 / 927.21 | 275.64 / 318.61 / 579.99 |
| vLLM | 760.08 / 1,251.41 / 2,129.29 | 551.11 / 819.85 / 1,487.73 | 230.04 / 291.60 / 669.73 | 78.45 / 87.95 / 396.91 |
At batch one, FP8 is almost exactly 2x BF16 and NVFP4 is about 2.8x. The advantage shrinks with long contexts and large batches as KV-cache traffic and attention account for more work.
DSpark
This test uses vLLM, 20 natural-language prompts, greedy decoding, 256 generated tokens, and concurrency one.
| Target | Plain | DSpark | Speedup | Plain TPOT | DSpark TPOT | TTFT change | Accept length / rate |
|---|---|---|---|---|---|---|---|
| FP8 | 23.18 tok/s | 52.24 tok/s | 2.25x | 43.05 ms | 18.73 ms | 66.5 → 124.5 ms | 3.15 / 30.73% |
| NVFP4 | 39.86 tok/s | 72.33 tok/s | 1.81x | 25.01 ms | 13.44 ms | 45.3 → 112.8 ms | 3.06 / 29.49% |
DSpark is useful for low-concurrency decode, but not for a saturated server. At batch 64 and context 128, plain FP8 reaches 1,098 tok/s while DSpark reaches 688; plain NVFP4 reaches 1,749 tok/s while DSpark reaches 623.
Compute and Memory Roofs
I measured both an 8192-cubed GEMM and a Qwen-like
M=8192, K=4096, N=12288 projection. The best median roof for each datatype is
100.5 TFLOP/s BF16, 198.9 TFLOP/s FP8, and 234.5 TFLOP/s NVFP4. The local
2-GiB CUDA copy reaches 232.4 GB/s; NVIDIA advertises 273 GB/s LPDDR5X.
| Phase / engine / datatype | Effective model FLOP/s | GEMM roof | GEMM util | Effective weight BW | LPDDR util | Bound |
|---|---|---|---|---|---|---|
| Prefill B64 L4096 / SGLang / BF16 | 88.4 TFLOP/s | 100.5 TFLOP/s | 88.0% | 0.315 GB/s* | 0.12%* | compute |
| Prefill B64 L4096 / SGLang / FP8 | 82.5 TFLOP/s | 198.9 TFLOP/s | 41.5% | 0.169 GB/s* | 0.06%* | compute |
| Prefill B64 L4096 / SGLang / NVFP4 | 133.1 TFLOP/s | 234.5 TFLOP/s | 56.8% | 0.185 GB/s* | 0.07%* | compute |
| Prefill B64 L4096 / vLLM / BF16 | 82.6 TFLOP/s | 100.5 TFLOP/s | 82.2% | 0.294 GB/s* | 0.11%* | compute |
| Prefill B64 L4096 / vLLM / FP8 | 123.3 TFLOP/s | 198.9 TFLOP/s | 62.0% | 0.253 GB/s* | 0.09%* | compute |
| Prefill B64 L4096 / vLLM / NVFP4 | 197.1 TFLOP/s | 234.5 TFLOP/s | 84.1% | 0.274 GB/s* | 0.10%* | compute |
| Decode B1 L512 / SGLang / BF16 | 0.229 TFLOP/s | 100.5 TFLOP/s | 0.23% | 243.5 GB/s | 89.2% | memory |
| Decode B1 L512 / SGLang / FP8 | 0.451 TFLOP/s | 198.9 TFLOP/s | 0.23% | 275.8 GB/s | 101.0%* | memory |
| Decode B1 L512 / SGLang / NVFP4 | 0.641 TFLOP/s | 234.5 TFLOP/s | 0.27% | 265.7 GB/s | 97.3% | memory |
| Decode B1 L512 / vLLM / BF16 | 0.217 TFLOP/s | 100.5 TFLOP/s | 0.22% | 230.5 GB/s | 84.4% | memory |
| Decode B1 L512 / vLLM / FP8 | 0.356 TFLOP/s | 198.9 TFLOP/s | 0.18% | 217.5 GB/s | 79.7% | memory |
| Decode B1 L512 / vLLM / NVFP4 | 0.617 TFLOP/s | 234.5 TFLOP/s | 0.26% | 255.8 GB/s | 93.7% | memory |
Effective model FLOP/s counts dense Qwen operations completed per second,
not physical low-precision instructions. For prefill, starred bandwidth is only
the checkpoint-once lower bound and deliberately excludes activation traffic;
compute is the relevant comparison. For decode, effective weight bandwidth is
checkpoint size times token rate. The FP8 estimate above 273 GB/s shows the
approximation error: checkpoint bytes are not identical to DRAM bytes and
caching is ignored. The relevant, bound-setting columns are bold.
Takeaways
- FP8 is the default I would use. It roughly doubles batch-one decode and did not show a statistically significant quality loss here.
- NVFP4 is faster, but the loss is visible. C-Eval drops 4.49 points and GSM8K drops 6.00 points relative to BF16.
- DSpark is for interactive generation. FP8 + DSpark reaches 52 tok/s without the NVFP4 quality penalty, but plain decoding wins at high batch.
- Batch-one decode is already at the memory roof. The simple estimates are around 89–101% of the advertised LPDDR bandwidth.
- SGLang FP8 prefill needs better GB10 tuning. It reaches 41.5% of the measured FP8 GEMM roof in the large prefill case, versus 62.0% for vLLM.