Liquid's LFM2.5-8B-A1B Beats OpenAI gpt-oss-20b 7/7 to 3/7 in Local Tool-Calling Benchmark
A benchmark run by Atomic Chat, a desktop app for running large language models locally, showed Liquid's LFM2.5-8B-A1B beating OpenAI's gpt-oss-20b in a local tool-calling test on a MacBook Pro M5 Max with 64GB of memory. In the trip-planning task, Liquid's model completed all 7/7 tool calls using 4.8 GB of RAM, running at 266 tokens per second and finishing in 6.9 seconds, while gpt-oss-20b completed 3/7 calls, used 11 GB of RAM, ran at 146 tokens per second and took 15.0 seconds.
The workload required three weather checks, two currency conversions, an email and a reminder, making it a test of whether a model can reliably trigger outside tools rather than just answer in chat. While the comparison covered a single local-agent scenario, it highlighted that tool-calling performance can diverge from model size, with Liquid's smaller model outperforming the larger OpenAI model on completion, speed and memory use.
From the sources (2 posts)
@josephjacks_RT @atomic_chat_hq: Liquid's LFM2.5-8B-A1B smashed OpenAI's gpt-oss-20b on tool calling We ran both locally on a MacBook Pro M5 Max, 64GB,…
@rohanpaul_aiatomic[.]chat (a desktop app that runs LLMs locally) ran a very revealing comparison for local AI agents, on a MacBook Pro M5 Max, 64GB. Liquid’s much smaller LFM2.5-8B-A1B beat gpt-oss-20b by finishing every required tool call, cutting ru