Command Palette
Search for a command to run...

Reachy Mini Adds Local Conversations With Open-Source llama.cpp Realtime API

aiai-productsai-infrastructureai-inference-platforms 3 posts · 3 accounts

Reachy Mini introduced local realtime conversations using an open-source Realtime API built on llama.cpp. The stack combines Parakeet, Gemma 4 E4B and Qwen3TTS, and the announcement said the local setup enables "free chats forever."

The system is designed to run wherever local language models can run, with demonstrations shown on DGX Spark and a 36GB M3 Pro MacBook. The release also said response speeds were high enough that delays had to be hardcoded to avoid interrupting users mid-sentence.

From the sources (3 posts)

@lvwerra

RT @andimarafioti: Introducing local Reachy Mini conversations: free chats forever! So fast that we had to hardcode delays to stop it from…

@thezachmueller

RT @andimarafioti: Introducing local Reachy Mini conversations: free chats forever! So fast that we had to hardcode delays to stop it from…

@andimarafioti

Introducing local Reachy Mini conversations: free chats forever! So fast that we had to hardcode delays to stop it from interrupting you mid-sentence. We built an open-source Realtime API powered by llama.cpp: Parakeet -> Gemma 4 E4B ->

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive