StepFun Releases Step 3.7 Flash, 198B Open-Weights Vision-Language Model
StepFun released Step 3.7 Flash, an open-weights vision-language model under Apache 2.0 aimed at agentic, coding, search and multimodal workloads. The company said the sparse mixture-of-experts system has 198 billion parameters with about 11 billion active, a 256,000-token context window, three reasoning levels and throughput of up to 400 tokens a second, and posted benchmark results of 67.1 on ClawEval-1.1, 79.2 on SimpleVQA Search, 56.3 on SWE-PRO and 95.3 on V* Python.
StepFun said the model can interpret interfaces, charts, documents and images before writing code or calling tools. The launch arrived with day-0 support on Nvidia's platform and in vLLM, and users also reported local runs on hardware such as DGX Spark, including one Q4_K_S setup at about 27 tokens a second; initial hands-on comparisons were mixed, with some testers seeing performance near V4-Flash while others judged its vision output weaker than rival models.
From the sources (15 posts)
@nvidiaaiStep 3.7 Flash is here ICYMI: 198B MoE with 11B active params, 256K context, native image + video support. Day 0 support is live on with GPU-accelerated endpoints, deploy with NVIDIA NIM inference microservices, a
@stepfun_ai⚡️ Step 3.7 Flash is here: The new frontier is agent efficiency. #1 ClawEval-1.1 (67.1), #1 SimpleVQA Search (79.2), #2 SWE-PRO (56.3), 95.3 on V* Python. Open weights under Apache 2.0. Built for agentic, coding, search, and multimodal wo
@xianbao_qianStepFun 3.7 Flash is out. This could be a key model as they're marching towards being IPO listed on Hong Kong. - Apache 2.0 - 198B MoE A11B (196B LLM + 1.8B Vision Encoder) - 256k context length - variable reasoning levels - Native FP8,
@teortaxestexI've been waiting for this! They managed to do it before June, and they open sourced it right away! @antirez I've been saying. Look at this model. It's much smaller than V4-Flash, it's multimodal, it's fast. It deserves to be added.
@vllm_project🎉 Congrats to @StepFun_ai on releasing Step-3.7-Flash, with day-0 support in vLLM. - 198B sparse MoE vision-language model, ~11B active params per token, native image + text input - 256K context window for long docs, multi-file repos, and
@teortaxestexFirst impressions: StepFun 3.7 vision is kinda low-res and hallucinatory, behind MiMo 2.5. Kimi >> DS-Vision > MiMo > StepFun. Well, it's their first vision in the Flash series and this is by far the smallest and fastest model.
@adinayakupStep-3.7-Flash 🔥 New VL model from @StepFun_ai ✨ 198B / 11B active - MoE ✨ 256K context ✨ 3 reasoning level ✨ Up to 400 tokens/sec 🤯
@victormustarRT @StepFun_ai: ⚡️ Step 3.7 Flash is here: The new frontier is agent efficiency. #1 ClawEval-1.1 (67.1), #1 SimpleVQA Search (79.2), #2 SW…
@teortaxestexStep 3.7 is generally neck and neck with V4-Flash (which is underappreciated as a powerful agent), I think they targeted it. But it goes to show that vision is a must now. V4.1V can't come soon enough.
@stevibeStep-3.7-Flash Q4_K_S on DGX Spark (GB10, 128GB): > ~27 tok/s generation > 198B sparse MoE, ~11B active > 256K context, native vision > Agentic / tool-calling / reasoning > Apache 2.0 I added a mobile chat screen on the right showing what
@quixiaiI love this chart!
@quixiaiRT @StepFun_ai: ⚡️ Step 3.7 Flash is here: The new frontier is agent efficiency. #1 ClawEval-1.1 (67.1), #1 SimpleVQA Search (79.2), #2 SW…
@theahmadosmanRT @The_Only_Signal: Setup Step 3.7 Flash on two Blackwell RTX PRO 6000 GPUs and got it running and recorded the configs as well as early d…
@e01aiWith launch of Step 3.7 Flash by @StepFun_ai , we explored multiple new use-cases with "Flash Class" models. The balance of speed, cost and quality gives life to new way to interact with the model. This demo shows possibility to "Sample" t
@cyousakuraRT @StepFun_ai: ⚡️ Step 3.7 Flash is here: The new frontier is agent efficiency. #1 ClawEval-1.1 (67.1), #1 SimpleVQA Search (79.2), #2 SW…