Command Palette
Search for a command to run...

Thinking Machines Releases Inkling-Small, A 276B-Parameter Model Matching Flagship Benchmarks

aiai-modelingai-model-releasesai-open-modelsai-research-evals 65 posts · 40 accounts

Thinking Machines released Inkling-Small, a Mixture-of-Experts model with 276B total parameters and 12B active parameters, that scores 31.6% on the Humanity’s Last Exam reasoning benchmark and exceeds 80% on SWEBench-Verified, matching its larger 975B-parameter Inkling sibling at a quarter of the size. Built using a refined pre-training data mix and agentic coding reinforcement learning, the open-weights model leverages on-policy distillation to outperform its larger counterpart across active computation budgets.

The model natively processes text, image, and audio with a 1M-token context window and variable thinking effort, making it viable for fine-tuning and local deployment on single Nvidia Blackwell chips. While it edges out the flagship on coding and reasoning tasks, it trails on factual recall, recording 20.9% on SimpleQA Verified versus the larger model’s 43.9%. The weights are licensed under Apache 2.0, with day-zero inference support immediately available on vLLM, SGLang, and Modal.

From the sources (25 posts)

@tomaarsen

Yesss! I've been waiting for this release for some time. Multilingual dense and multi-vector embedding models for retrieval. They look extremely strong, with all data public! - lightonai/mDenseOn - lightonai/mLateOn Great work by @antoi

@thinkymachines

Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available. Fi

@thinkymachines

Efficiency is the point. Across agentic tool use (Terminal-Bench 2.1), reasoning (HLE), and instruction following (IFBench), Inkling-Small delivers more performance per FLOP than Inkling. Variable thinking effort lets you pick your point on

@thinkymachines

Inkling-Small began training after its larger counterpart, so it benefits from everything we learned: an improved pre-training data mix, a refined ML recipe, on-policy distillation with Inkling as the teacher, and two further weeks of agent

@thinkymachines

It matches or exceeds Inkling on reasoning and agentic tasks. 31.6% on HLE, ahead of Inkling’s 29.7%, and the advantage holds at every thinking budget. On SWEBench-Verified it exceeds 80%.

@thinkymachines

Like Inkling, it's natively multimodal. It’s encoder-free, with audio and images processed jointly with text. It nearly matches Inkling across multimodal evals, and it can use Python to crop, zoom, and inspect images while reasoning over do

@thinkymachines

Inkling and Inkling-Small are both available on Tinker with a limited-time discount, and all Tinker models can now be chatted with on Tinker Playground. As always, we're keen to see what you build.

@miles_brundage

RT @thinkymachines: Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its si…

@miramurati

Inkling-Small is comparable to Inkling at a quarter the size. Weights are open, fine-tunable on Tinker today. Look forward to seeing what people make with it.

@_alex_kirillov_

RT @miramurati: Inkling-Small is comparable to Inkling at a quarter the size. Weights are open, fine-tunable on Tinker today. Look forward…

@soumithchintala

Inkling-small. 2 weeks after inkling Nearly as good as Inkling but 4x smaller. We're just getting started...🔥

@rown

RT @thinkymachines: Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its si…

@nvidiaai

Another open-weight release from @thinkymachines 👀 Inkling-Small is here. With native reasoning over audio and images and variable thinking effort, it's a great choice for fine-tuning, with NVIDIA NeMo on NVIDIA DGX Station. NVFP4 checkpo

@_kevinlu

RT @thinkymachines: Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its si…

@thezachmueller

RT @thinkymachines: It matches or exceeds Inkling on reasoning and agentic tasks. 31.6% on HLE, ahead of Inkling’s 29.7%, and the advantage…

@unslothai

@thinkymachines Congrats on the release and thanks for supporting open-source! We made some Inkling-Small GGUFs for you guys to run the model locally 🤗

@techmeme

Thinking Machines releases Inkling-Small, an open-weight model with 276B total and 12B active parameters, saying it "achieves comparable performance" to Inkling (Thinking Machines Lab) (Visit Techmeme dot com for the link and full context!

@huggingface

RT @thinkymachines: Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its si…

@mtslive

SITUATION DETECTED: Thinking Machines has released Inkling-Small, a 276B parameter model with 12B active that it says matches its larger Inkling model at a quarter of the size. The weights are open.

@vllm_project

🎉 Congrats to @thinkymachines on Inkling-Small-- live with Day 0 vLLM support! 276B total parameters with 12B active, native text, image, and audio input, and a 1M-token context window. Open weights, built for agentic and tool-use systems,

@lmsysorg

Inkling-small is out today! With SGLang, you can get 648 tok/s decode with DSpark (simulated acc len=4) and 288 tok/s w/o DSpark, under the same setup (8x @NVIDIAAI B200, TP 8, NVFP4, bs=1). What makes this model different is the size. 276

@woosuk_k

RT @vllm_project: 🎉 Congrats to @thinkymachines on Inkling-Small-- live with Day 0 vLLM support! 276B total parameters with 12B active, na…

@thezachmueller

RT @lmsysorg: Inkling-small is out today! With SGLang, you can get 648 tok/s decode with DSpark (simulated acc len=4) and 288 tok/s w/o DSp…

@arena

Inkling-Small (@thinkymachines) debuts at rank ~#88 (1431 pts, AutoEval) in Text Arena and ~#21 among open. This makes it one of only five US models in the top 25 open models overall. Note: this is an early AutoEval score, in which a Rewa

@mervenoyann

Thinking Machines released Inkling Small (🦖) + NVFP4 12B active 276B total params, the model performs better than larger Inkling on coding 🤯 > check out our blog covering benchmarks, performance and deployment

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive