Command Palette
Search for a command to run...

Liquid's LFM2.5-8B-A1B Beats OpenAI gpt-oss-20b 7/7 to 3/7 in Local Tool-Calling Benchmark

aiai-modelingai-research-evals 2 posts · 2 accounts

A benchmark run by Atomic Chat, a desktop app for running large language models locally, showed Liquid's LFM2.5-8B-A1B beating OpenAI's gpt-oss-20b in a local tool-calling test on a MacBook Pro M5 Max with 64GB of memory. In the trip-planning task, Liquid's model completed all 7/7 tool calls using 4.8 GB of RAM, running at 266 tokens per second and finishing in 6.9 seconds, while gpt-oss-20b completed 3/7 calls, used 11 GB of RAM, ran at 146 tokens per second and took 15.0 seconds.

The workload required three weather checks, two currency conversions, an email and a reminder, making it a test of whether a model can reliably trigger outside tools rather than just answer in chat. While the comparison covered a single local-agent scenario, it highlighted that tool-calling performance can diverge from model size, with Liquid's smaller model outperforming the larger OpenAI model on completion, speed and memory use.

From the sources (2 posts)

@josephjacks_

RT @atomic_chat_hq: Liquid's LFM2.5-8B-A1B smashed OpenAI's gpt-oss-20b on tool calling We ran both locally on a MacBook Pro M5 Max, 64GB,…

@rohanpaul_ai

atomic[.]chat (a desktop app that runs LLMs locally) ran a very revealing comparison for local AI agents, on a MacBook Pro M5 Max, 64GB. Liquid’s much smaller LFM2.5-8B-A1B beat gpt-oss-20b by finishing every required tool call, cutting ru

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive