Miso One Debuts as Open-Weights 8B Text-to-Speech Model With One-Shot Voice Cloning
Miso One has been released as an open-weights, 8-billion-parameter text-to-speech model aimed at expressive speech generation, offering one-shot voice cloning from a short sample and 110 milliseconds of latency. The model weights are available on GitHub, allowing users to self-host the system rather than rely on an API.
The model is positioned for voiceover work such as shorts, podcasts and educational content. Users can fine-tune it locally and keep audio data on their own machines, while API access is planned for a later release.
From the sources (3 posts)
@kimmonismusMiso One is live: an open-weights voice model built to sound like a real person reading, with actual warmth and pacing where most TTS still goes flat. 8B params, free on GitHub, with one-shot voice cloning from a short sample at 110ms lat
@omarsar0Another banger open-source release. Miso One is an 8B text-to-speech model with real emotional range, so voiceovers carry warmth, hesitation, and excitement instead of sounding flat. It's purpose-built for voiceover work like shorts, pod
@dair_aiRT @omarsar0: Another banger open-source release. Miso One is an 8B text-to-speech model with real emotional range, so voiceovers carry w…