Command Palette
Search for a command to run...

DeepSeek V4 Materials Point to Nvidia Training Over Ascend

aiai-modeling 9 posts · 2 accounts

Interpretations of DeepSeek's V4 technical materials suggest the model family was trained primarily on Nvidia hardware, likely Hopper, with Huawei's Ascend 950 used for reinforcement-learning rollouts rather than core backward-pass training. That points to Nvidia GPUs handling the main training workload, while Ascend appears to have had a narrower role in the rollout pipeline.

Benchmark and throughput discussion around the release portray V4-Flash as the more mature variant. Cited figures put Flash at about 90 tokens a second at batch size 256, or roughly 26,000 tokens a second per node, while reported V4-Pro numbers have drawn criticism. The architecture is described as highly complex, and the near-term priority appears to be reinforcement learning rather than additional compute to further refine Pro.

From the sources (9 posts)

@teortaxestex

V4 might be the first time I don't see a ton of dumb censorship "testing" like Tiananmen/Uighurs/Xi Jinpooh questions. I guess these folks believe that it's now distilled from Claude and thus there's no point

@teortaxestex

we have a pretty clear confirmation that DSV4 was trained on Nvidia (and almost certainly Hopper) and Ascend 950 was only used for generating RL rollouts. Basically they never use any FP4 in a backward pass, and FP4 is what they get from As

@teortaxestex

just realized that I'm not seeing V4 on the most interesting benchmark WeirdML from @htihle because of Havard's principle to not expose it to Chinese providers, and native-quality third party Western support … might take time

@teortaxestex

> it's real EQ-Bench is one of the most big-model-smell loaded evals there's no Kimi K2.6 here, though. I think it could contend but overall, yeah, about right. V4 is a big one, though underbaked (V4-Flash is around Sonnet 4.5) https://t

@teortaxestex

This seems absolutely mangled. We "know" now that they can do ≈90 tps on Flash with BS=256 (1.6K/chip, 26K tps/node, <2x of what DS was doing on 8xH800s with V3; meh), or 4722/GPU with 49 tps (≈89% of my H800 estimate! good job Huawei!).

@teortaxestex

bro is pissed at DeepSeek's engineering lol These things serve the purpose of solving problems as they come up, to cram more ability into the model. As you can tell by Flash, it works yes DSV4 is insanely overengineered and has too many mov

@teortaxestex

If DeepSeek has 32T of reasonably unique tokens, they could probably just repeat for 2-3 more epochs to push V4-Pro to the same "cooked" status as Flash, we know it works. This is the Frontier toolkit… But RL takes priority.

@teortaxestex

note that Chong Ruan, Jun Ran and Junlong Li were already departed as of V3.2, so V4 only adds 7 names V4 has been in the works for a while the entire team has significantly grown – V3.2 had 212 names, V4 had 270. So overall they've added n

@poezhao0605

RT @poezhao0605: DeepSeek released V4 this week. The model is impressive. But the technical report is more interesting than the benchmarks.…

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive