DeepSeek V4 Pro Debuts Second in Open-Weights Benchmark
DeepSeek has released V4 Pro and V4 Flash, its first new architecture since V3 and its first two-tier lineup. Artificial Analysis said V4 Pro scored 52 on its Intelligence Index, up from 42 for V3.2 and second only to Kimi K2.6 among open-weight reasoning models, while Flash scored 47. V4 Pro also led open-weight models on the GDPval-AA agentic benchmark with a score of 1554. The models offer a 1 million-token context window, and Artificial Analysis estimated the benchmark would cost $1,071 to run on Pro and $113 on Flash.
Tooling and deployment support appeared quickly. InferenceX said it added day-0 DeepSeek V4 support for Blackwell B300 systems and saw performance five times faster than Hopper, while separate testing reported V4 Flash running on a 4x DGX Spark/GB10 cluster after patches to vLLM, with further kernel optimization still underway. Early developer interest also centered on using Flash in agent software and exploring local deployment.
From the sources (25 posts)
@scaling01RT @j_dekoninck: GPT-5.5 leaps over GPT-5.4 and becomes #1 on MathArena! - Massive jump on BrokenArXiv (+34%) - Solid performance increase…
@tekniumAlso, deepseek v4 is available as well
@victortaelin@BitcoinBananaBY @EffectTS_ what is that T_T please wait for some updates I'll post soon, there were some bugs (specially on OpenAI models that used Codex instead of the API because GPT 5.5 didn't have an API yet)
@artificialanlysXiaomi’s MiMo V2.5 Pro has landed at 54 in the Artificial Analysis Intelligence Index, tied with Moonshot’s Kimi K2.6 - the current top open weights model. MiMo V2.5 Pro’s weights are expected to be released soon, which would make MiMo V2.5
@artificialanlysMiMo V2.5 Pro leads its peer group on agentic tasks. The model scores 1578 on GDPval-AA, and places it in the top tier for real-world work tasks among recent releases
@semianalysis_The Coding Assistant Breakdown: More Tokens Please, Hands On With GPT 5.5, Opus 4.7, DeepSeek V4, Why Benchmarks Are Bad, and Who's Going to Win READ NOW:
@yacinemtbRT @OedoSoldier: @teortaxesTex gorgeous V4 Soviet style poster from CN social media
@brianroemmeleRT @BrianRoemmele: WHALE REPORT! 🐳🐳🐳🐳🐳🐳🐳🐳🐳🐳🐳 15 hours of research. JENSEN OF NVIDA WAS RIGHT! DeepSeek-V4 introduces several architectu…
@victortaelin@theophorus7 I have no idea, I ran all my tests yesterday with codex headless mode and these are the results. Including the manual stuff it did was bad. From the API all is much better. That's all I know. Could be something silly. Thinking
@scaling01LisanBench results for GPT-5.5 - it's good. GPT-5.5 is now the strongest model without Thinking on both metrics! GPT-5.5-medium uses on average ~45.6% less tokens than GPT-5.4-medium while scoring 1.77x higher! (1.14x higher score on the
@victortaelinGPT 5.5 is much smarter than I thought Yesterday, I did one-shots, coding, benchmarks, and was disappointed. Today, I did it all again, except via the API, which is now available. Results changed completely: → one-shot prompts went from ba
@firstadopterOpenAI is back, folks. Told ya! Note: Anything written by @JordanNanos 's team kicks ass. "First we have to highlight GPT-5.5 from OpenAI. In our view, GPT-5.5 is now materially better at some tasks than all other models. We believe that
@scaling01RT @VictorTaelin: GPT 5.5 is much smarter than I thought Yesterday, I did one-shots, coding, benchmarks, and was disappointed. Today, I di…
@scaling01Since GPT-5 the reasoning efficiency of OpenAI models has continually improved! They are reaching higher scores with less and less tokens.
@scaling01@sama it's a good model
@poezhao0605RT @poezhao0605: The AI chip question in China is shifting. From “can Chinese chips match NVIDIA” to “can Chinese model labs reshape worklo…
@yacinemtbRT @zephyr_z9: Massive compute gap At least 40x-80x
@yacinemtbI cannot believe how good the deepseek infrastructure and research team are. I cannot believe how much compute multipliers we still have available for this new technology. None of this is nowhere solved. I can't believe they did what they d
@yacinemtbResearch talent is the spice
@yacinemtbYou know those crazy fuckers at deepseek will open source whatever they train once they actually raise a little bit of money (still 0.1% of amerilard companies)
@semianalysis_RT @SemiAnalysis_: The Coding Assistant Breakdown: More Tokens Please, Hands On With GPT 5.5, Opus 4.7, DeepSeek V4, Why Benchmarks Are Bad…
@simonwAnyone got DeepSeek-V4-Flash running on a Mac yet? 512GB or 256GB or 128GB or smaller?
@teortaxestexThe big question is how well does V4 respond to post-training. The small model seems to be pretty flexible, judging by updates over this 2 month period. If the big one is good too, they can finally get the long context synthetic data machin
@firstadopterThankfully, the media didn’t do the same thing with DeepSeek V4 that they did with DeepSeek R1.
@zephyr_z9Flash at 47, Max at 52 They encountered some serious issues while training V4 Max