Command Palette
Search for a command to run...

StepFun Releases Step 3.7 Flash, 198B Open-Weights Vision-Language Model

aiai-modelingai-model-releasesai-open-modelsai-research-evals 15 posts · 12 accounts

StepFun released Step 3.7 Flash, an open-weights vision-language model under Apache 2.0 aimed at agentic, coding, search and multimodal workloads. The company said the sparse mixture-of-experts system has 198 billion parameters with about 11 billion active, a 256,000-token context window, three reasoning levels and throughput of up to 400 tokens a second, and posted benchmark results of 67.1 on ClawEval-1.1, 79.2 on SimpleVQA Search, 56.3 on SWE-PRO and 95.3 on V* Python.

StepFun said the model can interpret interfaces, charts, documents and images before writing code or calling tools. The launch arrived with day-0 support on Nvidia's platform and in vLLM, and users also reported local runs on hardware such as DGX Spark, including one Q4_K_S setup at about 27 tokens a second; initial hands-on comparisons were mixed, with some testers seeing performance near V4-Flash while others judged its vision output weaker than rival models.

From the sources (15 posts)

@nvidiaai

Step 3.7 Flash is here ICYMI: 198B MoE with 11B active params, 256K context, native image + video support. Day 0 support is live on with GPU-accelerated endpoints, deploy with NVIDIA NIM inference microservices, a

@stepfun_ai

⚡️ Step 3.7 Flash is here: The new frontier is agent efficiency. #1 ClawEval-1.1 (67.1), #1 SimpleVQA Search (79.2), #2 SWE-PRO (56.3), 95.3 on V* Python. Open weights under Apache 2.0. Built for agentic, coding, search, and multimodal wo

@xianbao_qian

StepFun 3.7 Flash is out. This could be a key model as they're marching towards being IPO listed on Hong Kong. - Apache 2.0 - 198B MoE A11B (196B LLM + 1.8B Vision Encoder) - 256k context length - variable reasoning levels - Native FP8,

@teortaxestex

I've been waiting for this! They managed to do it before June, and they open sourced it right away! @antirez I've been saying. Look at this model. It's much smaller than V4-Flash, it's multimodal, it's fast. It deserves to be added.

@vllm_project

🎉 Congrats to @StepFun_ai on releasing Step-3.7-Flash, with day-0 support in vLLM. - 198B sparse MoE vision-language model, ~11B active params per token, native image + text input - 256K context window for long docs, multi-file repos, and

@teortaxestex

First impressions: StepFun 3.7 vision is kinda low-res and hallucinatory, behind MiMo 2.5. Kimi >> DS-Vision > MiMo > StepFun. Well, it's their first vision in the Flash series and this is by far the smallest and fastest model.

@adinayakup

Step-3.7-Flash 🔥 New VL model from @StepFun_ai ✨ 198B / 11B active - MoE ✨ 256K context ✨ 3 reasoning level ✨ Up to 400 tokens/sec 🤯

@victormustar

RT @StepFun_ai: ⚡️ Step 3.7 Flash is here: The new frontier is agent efficiency. #1 ClawEval-1.1 (67.1), #1 SimpleVQA Search (79.2), #2 SW…

@teortaxestex

Step 3.7 is generally neck and neck with V4-Flash (which is underappreciated as a powerful agent), I think they targeted it. But it goes to show that vision is a must now. V4.1V can't come soon enough.

@stevibe

Step-3.7-Flash Q4_K_S on DGX Spark (GB10, 128GB): > ~27 tok/s generation > 198B sparse MoE, ~11B active > 256K context, native vision > Agentic / tool-calling / reasoning > Apache 2.0 I added a mobile chat screen on the right showing what

@quixiai

I love this chart!

@quixiai

RT @StepFun_ai: ⚡️ Step 3.7 Flash is here: The new frontier is agent efficiency. #1 ClawEval-1.1 (67.1), #1 SimpleVQA Search (79.2), #2 SW…

@theahmadosman

RT @The_Only_Signal: Setup Step 3.7 Flash on two Blackwell RTX PRO 6000 GPUs and got it running and recorded the configs as well as early d…

@e01ai

With launch of Step 3.7 Flash by @StepFun_ai , we explored multiple new use-cases with "Flash Class" models. The balance of speed, cost and quality gives life to new way to interact with the model. This demo shows possibility to "Sample" t

@cyousakura

RT @StepFun_ai: ⚡️ Step 3.7 Flash is here: The new frontier is agent efficiency. #1 ClawEval-1.1 (67.1), #1 SimpleVQA Search (79.2), #2 SW…

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive