SevenTnewS

Open Source TTS

Audio8's new TTS model clones voices in 11 languages for free

Audio8 releases Audio8-TTS Preview under Apache 2.0, a 0.6B model supporting 11 languages, zero-shot voice cloning, and a bundled 44.1 kHz codec. The release targets edge speech applications that need speed and privacy.

Emmanuel Fabrice Omgbwa Yasse AI-assisted

2026-07-30 · 1 min read

Audio8's new TTS model clones voices in 11 languages for free
Sources : Samuel Zeng thr…

Audio8's new model blurs the line between cloud-based and local speech synthesis. Announced today, the Audio8-TTS preview packs 0.6 billion parameters, zero-shot voice cloning across 11 languages, a bundled 44.1 kHz audio codec, all under Apache 2.0, aligning with the open-weight movement that 41 signatories just defended in a public letter.

At 0.6B parameters, the model runs on consumer hardware and clones a voice from a short sample without fine-tuning. Audio8 calls it a single stack for speech applications that need speed, reach, and privacy. Running locally means no audio leaves the device, a privacy advantage increasingly demanded by enterprise deployments, as covered in a comparison of AI infrastructure strategies.

Small speech models are catching up fast. Voxtral Realtime, an open-source streaming ASR model, matched Whisper's quality with a 480 ms delay earlier this year. Alibaba's Qwen2.5-Omni runs on a phone. Nvidia's Cosmos 3 Edge similarly packs world modeling into 4 billion parameters for on-device tasks, ranking first on VANTAGE-Bench in its size class. Audio8's take is generative: it creates speech rather than transcribing it. At 0.6B parameters, it stays under the 1B mark where most open TTS models sit, and the zero-shot cloning works from English to Mandarin.

The bundled 44.1 kHz codec eliminates the need for a separate vocoder and delivers clean output by default. The

Get the tech essentials in 3 minutes every morning

One email, every weekday, with what actually matters in AI and tech.