DGX Spark value ladder: 1 box, 2 boxes, 4 boxes vs RTX Pro 6000, H100, and the new DGX Station GB300
As of July 2026. At launch everyone dunked on the DGX Spark's 273 GB/s memory bandwidth. Nine months later, the picture has flipped. The Spark stacks over RoCE (2x, 4x clusters). MoE is the baseline model architecture. Speculative decoding compounds throughput. And the numbers from our own 2-node cluster tell a clear story: a frontier-adjacent model running on a desk, drawing less than half a kilowatt.
This article walks the real-value ladder: from a single GB10 box all the way up to the new DGX Station GB300, with stops along the way at where the Spark wins and where it does not.
| Rung | Build | Price | Best model it runs | Speed | Power |
|---|---|---|---|---|---|
| 1 | Single GB10 (ASUS GX10 1TB) | $3,977 | Qwen3.5-122B-A10B (DFlash) | ~59-81 tok/s | 240W |
| 2 | 2x Spark cluster | $9,500 | DeepSeek V4-Flash 284B (DSpark) | ~61 tok/s | 460W |
| 3 | 4x Spark cluster | $19,500 | Nemotron 3 Ultra 550B at Q4 | fits, throughput unverified | 920W |
| 4 | 1x RTX Pro 6000 Blackwell (96 GB) | $12,200 | Qwen3.5-122B-A10B | ~190 tok/s est. | 600W |
| 5 | H100 80 GB workstation | $32,000 | cannot fit V4-Flash | N/A | 700W |
| 6 | DGX Station GB300 (MSI WS300) | ~$85,000 | up to ~1T params at NVFP4 | 20 PFLOPS FP4 ceiling | 1,600W |
Every figure in this table traces to a dated primary source or our own measured data. Let's walk each rung.
Rung 1: single GB10 box — the 273 GB/s complaint was real, but so is the stack
A single GB10 box like the ASUS Ascent GX10 ($3,977 at Newegg, July 2026) or the DGX Spark Founders Edition ($4,699) has 128 GB of unified LPDDR5X memory at 273 GB/s. That bandwidth number is not wrong. On a dense 70B model, single-stream decode chokes at ~8 tok/s (from our hardware dataset). The reviewers who called it memory-bandwidth-bound were right about that.
But people do not run 70B dense on a Spark. They run MoE.
The Qwen3.5-122B-A10B is a 122B-parameter model where only 10B activate per token. At INT4, the weights take 67 GB. On a single Spark with DFlash block-speculative decoding, it runs at 59 tok/s in general use and up to 81 tok/s on agentic tool-call traffic (NVIDIA DevForum #374328, June 2026). The MTP heads baked into the model cost nothing extra. Spec-decode compounds throughput. The 273 GB/s is still the ceiling, but on a model that only moves 10B active weights per token rather than 70B, the ceiling is not what the early reviews measured.
Where it wins: A single desk box that runs a 122B frontier-class MoE at usable speed. No consumer GPU below $12,000 can load that model at any quant.
Where it loses: Single-stream dense models. Prefill is compute-bound and fast. Decode is bandwidth-bound and slow. A 100K-token prompt takes about 3 minutes on a Spark versus 11 minutes on a Mac (our own "two speeds" measurement).
Rung 2: 2x Spark cluster — where the stacking thesis proves out
Two Sparks over a 200 Gb/s RoCE interconnect cost $9,500 total. That buys you 256 GB of unified memory, enough to load DeepSeek V4-Flash, a 284B-parameter model with 13B active, at NVFP4 weights (~149 GB).
Our own 2-node cluster runs this model at 61.4 tok/s single-stream with DSpark speculative decoding (first-party measurement, June 30, 2026). At concurrency 8 it hits 171 tok/s. At concurrency 16, 261 tok/s. The full stack is NVFP4 weights plus FP8 KV, vLLM at TP=2, DSpark with gamma=5, and a 400K context window. It draws roughly 460W total.
What this means: For under ten thousand dollars, you get a desk-sized cluster running a model from the same family as DeepSeek V4. The H100 system that could load this model costs at least $64,000 (two cards plus host) and fills a server rack. The 2x Spark cluster is 1/6 the price, fits on a desk, and sipping power.
Where it loses: The 273 GB/s bandwidth per node still limits single-stream speed. The H100's 3,350 GB/s would run V4-Flash much faster per token, if you could afford the system, cool it, and power it. The Spark cluster wins on capacity-per-dollar and total cost, not on raw token speed.
Rung 3: 4x Spark cluster — capacity you cannot buy any other way for $19,500
Four Sparks cost $19,500 and give you 512 GB unified memory (~488 GB usable). At Q4, that fits models in the 300-380 GB range: NVIDIA's Nemotron 3 Ultra 550B-A55B (300 GB at Q4, launched June 4, 2026), MiniMax M3 428B-A23B (245 GB), Mistral Large 3 675B-A41B (370 GB), or DeepSeek V3 671B-A37B (380 GB).
The verified benchmark on 4x Spark is from the NVIDIA Developer Forums: V4-Flash 284B at FP8 hits 49 tok/s single-stream and 180 tok/s at concurrency 8 (DevForum #373808, June 2026, TP=4+EP with MTP-2). For models larger than V4-Flash, the 550B-class MoEs that are this rung's real sell, no verified cluster throughput exists yet. The capacity fact is solid: these models load at Q4 with room for KV cache. The throughput number — that part remains unmeasured.
What this buys: A desk-sized 512 GB CUDA cluster that runs the current generation of 550B-675B MoE models at Q4. To get the same capacity from RTX Pro 6000s would cost at least $38,000 (four cards plus a Threadripper Pro host). The 4x Spark is half the price and draws 920W versus 2,200W.
Rung 4: RTX Pro 6000 Blackwell — the comparison that matters
A single RTX Pro 6000 Blackwell with 96 GB GDDR7 ECC and 1,792 GB/s bandwidth costs $12,200 as a whole build (Newegg card at ~$11,100 plus ~$1,100 host, June 2026). Its 1,792 GB/s is 6.5 times the Spark's per-node bandwidth. On models that fit a single card, it will be substantially faster.
How many Pro 6000s does it take to run what each Spark rung runs, and what does that cost?
- Qwen3.5-122B (67 GB at Q4, the single Spark's headliner) fits one Pro 6000 (93 GB usable). So 1 Pro 6000 ($12,200) matches a single Spark ($3,977-4,699) on model capacity. The Pro 6000's 1,792 GB/s bandwidth is 6.5x the Spark's per-node bandwidth, so on models that fit one card it should be faster on decode, but no published 122B-MoE single-stream decode benchmark exists for the Pro 6000 - we compare on capacity, price, and bandwidth ceiling. It also costs 2.6-3x more.
- DeepSeek V4-Flash (~149 GB at NVFP4, the 2x Spark's headliner) needs 2 Pro 6000s (188 GB usable). That is $24,400 versus $9,500, a 2.6x premium. And with no NVLink between Pro 6000s, tensor-parallel throughput is PCIe-bound.
- Nemotron 3 Ultra 550B (300 GB at Q4, the 4x Spark's headliner) needs 4 Pro 6000s (372 GB usable). That bill is $38,000 versus $19,500, or 1.9x the price.
The Pro 6000's advantage is real on models that fit one card: faster decode, faster prefill, higher concurrency. But the Spark cluster runs the same models at 1.9-3x lower cost.
Rung 5: H100 — shortest section in this article
An H100 80 GB workstation costs $32,000. It cannot fit DeepSeek V4-Flash (~149 GB at NVFP4) at any useful quant. To run V4-Flash on H100 hardware, you need at least two H100s: $64,000 plus host for a system that fills a server chassis and draws 1,400W plus cooling.
No published single-stream V4-Flash H100 decode benchmark exists (web-checked July 2026). The H100 belongs in a datacenter where batch serving and training are the job, not in a desk-cluster comparison.
Rung 6: DGX Station GB300 — the whole-ladder inversion
The new GB300-based DGX Station, shipping since June 2026 via ASUS, Dell, HP, MSI, Gigabyte, and Supermicro, changes the question from "can I afford the rung above Spark" to "at what multiple."
The MSI XpertStation WS300 is priced around $85,000 (TechRadar, June 2026). It packs a B300 Blackwell Ultra GPU with 252 GB of HBM3e at 7.1 TB/s, a 72-core Grace CPU with 496 GB of LPDDR5X, and a 1,600W power supply. Total coherent memory: 748 GB. It runs at up to 20 petaFLOPS FP4.
The value inversion is stark:
A 4x Spark cluster ($19,500) fits the same MoE models the Station runs: Nemotron 3 Ultra 550B, Mistral Large 3 675B, DeepSeek V3 671B at Q4. The Station costs 4-6x more. What the Station gives you for that premium: 7.1 TB/s HBM bandwidth (26x a single Spark), the ability to train and fine-tune, 800 Gb/s networking, and a 20 PFLOPS compute ceiling that no desktop cluster matches.
The question is not "is the Station worth it." It is: do you need 7.1 TB/s and 20 PFLOPS on your desk, or does a 2-4 Spark cluster already run the models you need?
The answer to the title question
Did the DGX Spark turn out to be too good? Not for every job. The 273 GB/s bandwidth is real, and it shows on dense decode and long-context prefill. A single RTX Pro 6000 will outrun a single Spark on anything that fits 96 GB, and the DGX Station will outrun everything in this article on anything that fits 748 GB.
But the Spark stacks: two boxes for $9,500 run a 284B MoE at 61 tok/s on a desk drawing 460W, and no other product in its price class can do that. The ladder from 1 to 2 to 4 Sparks scales model capacity at a cost no discrete-GPU build can touch, at a power draw and physical footprint no server can approach.
If you need to run frontier-class MoE models on your desk, the Spark cluster is the cheapest path to them. If you need raw token speed on models that fit one card, buy the Pro 6000. If you need 7.1 TB/s HBM and training capability, save up for the Station. The Spark sits in a gap no other product fills: the gap between "I run small models fast" and "I spend six figures for the frontier."
For the model that best fits your exact machine and memory budget, the LLMRequirements picker matches every open model against your hardware and shows expected speed.
Explore the full DGX Spark specs on the DGX Spark build page, compare options side-by-side on the comparison tool, or browse our verified recipes for the Spark stack.
Sources
- Our hardware dataset: DGX Spark 128 GB ($4,699), 2x DGX Spark cluster ($9,500), 4x DGX Spark cluster ($19,500), ASUS Ascent GX10 1 TB ($3,977), RTX Pro 6000 Blackwell 96 GB ($12,200), 2x RTX Pro 6000 Blackwell ($24,400), 4x RTX Pro 6000 Blackwell ($38,000), H100 80 GB ($32,000)
- Our hardware dataset: DeepSeek V4-Flash 284B MoE (~160 GB at Q4), Qwen 3.5 122B-A10B MoE (~67 GB at Q4), MiniMax M3 428B MoE (~245 GB at Q4), Nemotron 3 Ultra 550B-A55B MoE (~300 GB at Q4), Mistral Large 3 675B MoE (~370 GB at Q4)
- Our own measurement (June 30, 2026): 2x DGX Spark cluster, DeepSeek V4-Flash DSpark, 61.4 tok/s single-stream, 171@C8, 261@C16, NVFP4 KV, vLLM TP=2, 400K ctx
- NVIDIA Developer Forums #374328 (June 2026): 1x Spark Qwen3.5-122B DFlash ~59-81 tok/s, INT4 AutoRound + FP8 experts, vLLM 0.23, 262K ctx
- NVIDIA Developer Forums #373808 (June 2026): 4x Spark V4-Flash FP8 ~49 t/s (c=1), ~180 t/s (c=8), vLLM jasl fork, TP=4+EP, MTP-2, 384K ctx
- NVIDIA official product page (July 2026): DGX Station for Windows, 748 GB coherent memory, 20 PFLOPS FP4, up to 1T parameter models. nvidia.com
- pi3g.com (June 17, 2026): DGX Station available, 748 GB coherent (252 GB HBM3e @ 7.1 TB/s + 496 GB LPDDR5x @ 396 GB/s), 20 PFLOPS FP4, ConnectX-8 dual 400 GbE. pi3g.com
- ServeTheHome (March 2026): DGX Station GB300 revised specs, 7/8 HBM stacks enabled, 252 GB HBM3e, 7.1 TB/s, downgraded from original 288 GB spec. servethehome.com
- TechRadar (June 2026): MSI XpertStation WS300 at $85,000, GB300 Superchip, 748 GB coherent memory, dual 400 GbE. techradar.com
- NVIDIA DevForum #361713 (Feb 23, 2026): DGX Spark FE price increase $3,999 to $4,699
- Newegg (July 2026): ASUS Ascent GX10 1TB at $3,976.99 (N82E16859110044)