Command Palette
Search for a command to run...

MosiAI Releases 0.9B Open Source Model to Transcribe 90 Minutes of Audio With Speaker Labels in Single Pass

aiai-modelingai-model-releasesai-open-models 3 posts · 3 accounts

MosiAI released MOSS-Transcribe-Diarize-0.9B, a 0.9B parameter open source model that performs speech transcription, speaker diarization, and timestamp generation in a single pass. The system pairs a Whisper style audio encoder with a Qwen3 style causal decoder, allowing it to process recordings up to 90 minutes long without chunking or stitching.

The model operates with a 128K token context window and includes hotword biasing for names and domain terms, achieving approximately 100 tokens per second on an NVIDIA RTX 4090. Day zero inference support is available in vLLM and SGLang, with the architecture released under an Apache 2.0 license on Hugging Face and GitHub.

From the sources (3 posts)

@vllm_project

🎉 Congrats to the @MosiAI_Official team on MOSS-Transcribe-Diarize-0.9B, an open, end-to-end model for multi-speaker long-audio transcription, with day-0 support in vLLM. Most setups chain ASR + diarization + alignment (WhisperX-style). Th

@lmsysorg

🎉 Meet MOSS-Transcribe-Diarize-0.9B from the @MosiAI_Official team, a 0.9B open-source, end-to-end model for long-form multi-speaker ASR. Day-0 support is now live in SGLang! 1️⃣ Audio-in, structured text-out: timestamps + speaker labels,

@mosiai_official

🤗 MOSS-Transcribe-Diarize-0.9B is now open source on @huggingface. Built with an end-to-end audio-to-structured-transcript paradigm: >0.9B open-source ASR model >Apache license 2.0 >128k long-context transcription >Up to ~90-min audio inp

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive