Command Palette
Search for a command to run...

Soofi Consortium Launches 30 Billion Parameter Nvidia AI Model, Drawing Benchmark Contamination Claims

aiai-modelingai-open-modelsai-model-releasesai-research-evals 8 posts · 4 accounts

The Soofi consortium released Soofi S, a 30 billion parameter open foundation model for German and English that trains on approximately 27 trillion tokens and runs on Deutsche Telekom’s Industrial AI Cloud. The model utilizes a hybrid Mamba and Transformer mixture-of-experts design with 3 billion active parameters.

Analysis of the published training recipe shows that the model mirrors Nvidia’s open Nemotron 3 Nano architecture and shares roughly 80 percent of its underlying data mixture. Data reviews further indicate benchmark contamination, noting the model was trained on a dataset that rephrases questions from the GPQA Diamond test set for 10 epochs, which reportedly boosts the model’s score on that benchmark by 10 percentage points and accounts for 70 percent of the reported performance gap.

From the sources (8 posts)

@dan_jeffries1

RT @NXT4EU: Germany has launched one of the world's best open-source AI models. Soofi S, made by the Soofi consortium, is a 30B parameter…

@kimmonismus

RT @effi288: 📄 Releasing the Soofi S pretraining tech report: a sovereign, open foundation model for German and English Today we’re publis…

@kimmonismus

Their newly released pretraining report reveals that Soofi S 30B-A3B is based on NVIDIA’s open Nemotron 3 Nano reference architecture, the same hybrid Mamba + Transformer MoE design with ~3B active parameters and nearly identical architectu

@eliebakouch

@kimmonismus there is a huge overlap between nemotron and their mixture, also the recipe is the same and they choose a way to aggregate benchmarks that makes their model the best (you see qwen3 32B dense being worse than apertus 8B lol), st

@eliebakouch

@effi288 @kimmonismus hmm yeah i think my main point is that you created your own "capability index"/aggregation of benchmark. i just took a look and you put gpqa-d as a standalone eval with 20% of the index wieght, this alone account for ~

@eliebakouch

i took a second look since this is going viral, this model is a copy paste of nemotron 3 nano, same arch, ~80% of the mixture in common which oversells very hard the sovereign aspect. but this is not even the issue, they literally train on

@eliebakouch

tl;dr: they trained on 10 epochs of a "very light rephrasing" of EVERY gpqa diamond EVAL data :)

@bdsqlsz

RT @NXT4EU: Germany has launched one of the world's best open-source AI models. Soofi S, made by the Soofi consortium, is a 30B parameter…

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive