Command Palette
Search for a command to run...

DeepSeek Launches DSpark to Accelerate Language Model Serving

aiai-infrastructureai-inference-platformsai-modelingai-research-evals 9 posts · 8 accounts

DeepSeek has launched DSpark, a speculative decoding system that speeds up live serving of its DeepSeek V4 language model. The company says the method boosts serving throughput by 51% to 406% under stricter latency targets and delivers 60% to 85% faster per-user text generation at matched throughput. DSpark uses a semi-autoregressive drafter to generate coherent token drafts and a confidence scheduler that verifies only the token prefixes most likely to be accepted, reducing wasted compute during inference.

Benchmarks circulating on social media show the new system outperforming DeepSeek's prior speculative decoding tool, DFlash, as well as EAGLE-3. In published tests, DSpark achieved an average throughput of 127 tokens per second, compared to 111 for DFlash and 81 for EAGLE-3. The open-source AI inference engine vLLM announced it is already working to integrate the DSpark algorithm, which would allow developers to apply the acceleration technique to a wider range of large language models.

From the sources (9 posts)

@askalphaxiv

RT @askalphaxiv: DeepSeek just published DSpark, a speculative decoding system that boosts live DeepSeek V4 serving throughput by 51% to 40…

@teortaxestex

Extending DSpark to GDN-based models will be huge I hope Kimi team in K3 uses it (or some more advanced acceleration in this genre) from the get-go. This also speeds up RL…

@autotrustai

First model to support DSpark speculative decoding.@deepseek_ai DeepSeek-V4-Flash-DSpark-4E: 284B MoE → only 11B activated per token 18% faster inference, +3% MMLU-Pro accuracy 1M context. Fully open (MIT).

@zeningchen42844

How does DeepSeek make AI text generation 60–85% faster with zero quality loss? After reading the DSpark paper, I used Sai to turn the paper into an animated explainer. The visualization made it much easy to me to digest. Strongly recommend

@tech2wild

Got DeepSeek V4 Flash DSpark running on 2x DGX Spark with vLLM TP=2. Single-stream: 66.78 / 62.36 / 58.30 tok/s 62.48 tok/s mean 0.673 draft acceptance 3.36 accepted/draft Built on @rafaelcaricio DSpark vLLM work + @MiaAI_lab 2x Spark ru

@teortaxestex

> the improvement over DFlash: > +20% acceptance length > +14% throughput > avg 127 tok/s vs 111 for DFlash and 81 for EAGLE-3 > for single GPU inference, DSpark's lightweight + early stopping approach is clearly superior …D

@vllm_project

👀 vLLM community is working non-stop to get @deepseek_ai's new DSpark spec decode algorithm for vLLM! Faster inference for everyone!

@thezachmueller

RT @mgoin_: this means GLM 5.2 DSpark on the way btw

@analyticsindiam

NEW: DeepSeek has launched DSpark, a new speculative decoding system that significantly speeds up live serving of DeepSeek V4. Key highlights: • 51%–406% higher serving throughput under stricter latency targets • 60%–85% faster per-user te

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive