AMD Lucebox Beats Nvidia DGX Spark by 3.63x on DeepSeek V4 Flash

GreenGames1 pts0 comments

Lucebox (AMD Radeon AI PRO R9700 + Strix Halo) Beats NVIDIA DGX Spark by 3.63x on DeepSeek V4 Flash Decode Speed | lucebox

By Davide Ciffa<br>Lucebox (AMD Radeon AI PRO R9700 + Strix Halo) Beats NVIDIA DGX Spark by 3.63x on DeepSeek V4 Flash Decode Speed<br>51.1 tok/s (Lucebox) vs 14.09 tok/s (single DGX Spark)<br>DeepSeek V4 Flash decode comparison<br>Measured decode throughput for one request at a time on the full 284B model SystemTested configurationDecode throughput LuceboxROCmFPX compressed model, DSpark draft model, four experts per token51.1 tok/s median 1&times; NVIDIA DGX SparkQ2 compressed model, average of tests at four context lengths14.09 tok/s Measured ratio51.1 &divide; 14.093.63&times; Using the unrounded DGX Spark mean of 14.0875 tok/s gives the 3.63&times; headline ratio. Each platform runs a configuration tuned for its own hardware, so this is a whole-system comparison rather than a GPU-only A/B test. Decode throughput is the rate at which each system generates new tokens.<br>DeepSeek V4 Flash has 284 billion parameters. The 102.3 GB ROCmFPX build is too large for the R9700&rsquo;s 32 GB of memory, so Lucebox gives the two AMD GPUs different jobs. The R9700 handles the dense path and frequently selected experts; Strix Halo holds the remaining experts in its 128 GB memory and runs them in parallel.<br>The headline Lucebox result is the median of three measured requests after two warmups, using an approximately 2k-token prompt and 128 generated tokens. An 11.3 GB DSpark draft model proposes several tokens at once; the full DeepSeek V4 Flash model verifies them before they are accepted. We also ran a separate 53-token prompt to measure the upper end of the implementation.<br>For the single-device baseline and the ROCmFPX compression details, see our companion report on DeepSeek V4 Flash on AMD Ryzen AI MAX+ 395.<br>Complete system price<br>AMD-Powered Lucebox is priced at $6,499 as a complete local inference system, including the custom chassis, 2 TB of storage, power delivery, hardware integration, validation, and warranty. NVIDIA lists one DGX Spark at $4,699, or $9,398 for two.<br>Complete systemCurrent U.S. priceWhat is included 1&times; DGX Spark$4,699Complete system with 4 TB storage AMD-Powered Lucebox$6,499Custom chassis, 2 TB storage, power delivery, integration, validation, and warranty 2&times; DGX Spark$9,398Two complete systems At those list prices, Lucebox costs 38% more than one DGX Spark and delivers about 2.6&times; the decode throughput per dollar in the headline configurations. It costs 31% less than two DGX Sparks.<br>Prices are the U.S. list prices available at publication, before tax and shipping. Sources: Lucebox system configuration and NVIDIA Marketplace. Configurations and prices can change.<br>One DGX Spark, measured directly<br>We ran a Q2 compressed version of the full DeepSeek V4 Flash model on one DGX Spark. We then measured Lucebox at the same 2k, 4k, 8k, and 16k context lengths.<br>Matched single-request decode throughput on the full 284B model ContextLucebox1&times; DGX SparkSpeedup 2k51.0 tok/s14.18 tok/s3.60&times; 4k49.6 tok/s14.24 tok/s3.48&times; 8k47.5 tok/s14.04 tok/s3.38&times; 16k42.9 tok/s13.89 tok/s3.09&times; Mean 47.75 tok/s 14.09 tok/s 3.39&times; The DGX Spark mean is 14.0875 tok/s, which we report as 14.09 tok/s. Lucebox averages 47.75 tok/s across the same four context lengths, a 3.39&times; speedup. The 3.63&times; headline uses the separate 51.1 tok/s Lucebox serving result shown above.<br>DGX Spark ran the Q2 model; Lucebox ran the ROCmFPX model with four experts per token and DSpark speculative decoding. Each system uses the configuration designed for its hardware.<br>For wider context, LocalMaxxing includes two-DGX-Spark results both with the standard target and with a DSpark-specific package. We show both below rather than selecting the more favorable reference.<br>DeepSeek V4 Flash decode, one request at a time · tok/s<br>1&times; DGX Sparkaverage of four tests

14.09

2&times; DGX Sparkwithout a helper model

45.70

AMD-Powered Luceboxserving median

51.10

AMD-Powered Luceboxshort prompt, fastest run

55.00

2&times; DGX Sparkwith DSpark

65.09

The Lucebox values are a three-request serving median and the fastest of three separate runs with a 53-token prompt. The single DGX Spark figure is our mean across 2k, 4k, 8k, and 16k context tests. Public sources: two DGX Sparks without a draft model and two DGX Sparks with DSpark. Prompts, context lengths, software, compression, and routing settings differ, so the public runs provide context rather than a controlled A/B comparison.

The public two-Spark results bracket Lucebox: the standard target is effectively tied with our longer decode run, while the DSpark-specific configuration is faster. They are useful reference points, but not substitutes for a matched benchmark.<br>How the work is split<br>Tensor parallelism usually divides the same calculation evenly between matching GPUs. That does not work well here because the R9700 and Strix Halo...

lucebox times spark model deepseek flash

Related Articles