Reachy Mini Adds Local Conversations With Open-Source llama.cpp Realtime API
Reachy Mini introduced local realtime conversations using an open-source Realtime API built on llama.cpp. The stack combines Parakeet, Gemma 4 E4B and Qwen3TTS, and the announcement said the local setup enables "free chats forever."
The system is designed to run wherever local language models can run, with demonstrations shown on DGX Spark and a 36GB M3 Pro MacBook. The release also said response speeds were high enough that delays had to be hardcoded to avoid interrupting users mid-sentence.
From the sources (3 posts)
@lvwerraRT @andimarafioti: Introducing local Reachy Mini conversations: free chats forever! So fast that we had to hardcode delays to stop it from…
@thezachmuellerRT @andimarafioti: Introducing local Reachy Mini conversations: free chats forever! So fast that we had to hardcode delays to stop it from…
@andimarafiotiIntroducing local Reachy Mini conversations: free chats forever! So fast that we had to hardcode delays to stop it from interrupting you mid-sentence. We built an open-source Realtime API powered by llama.cpp: Parakeet -> Gemma 4 E4B ->