V

VoxCPM2

Tokenizer-free TTS: clone any voice or design one from text, 30 languages, Apache 2.0

PriceFree / Open SourceReviewed2026-04-11

Expert verdict

Ship

3-1
3 Ships1 Skips
Visit github.com

The Panel's Take

VoxCPM2 is a 2B-parameter open-source text-to-speech model from OpenBMB that ditches the conventional approach of tokenizing speech into discrete units. Instead it models audio as continuous waveforms, producing 48kHz studio-quality output with an RTF of ~0.3 on an RTX 4090 — synthesizing 10 seconds of audio in about 3 seconds. It supports 30 languages and is released under Apache 2.0 for unrestricted commercial use. The standout capability is its dual voice creation modes: voice cloning from a short reference clip, and "voice design" where you describe a voice in plain text ("a calm middle-aged woman with a slight British accent") and the model generates a matching identity from scratch. This eliminates the dependency on reference audio for new character voices — a major workflow improvement for game devs, audiobook producers, and accessibility builders. VoxCPM2 is trending as one of the fastest-rising repositories on GitHub today, with over 9,300 stars since its recent release. A live HuggingFace demo is available for immediate testing. For developers building audio apps, games, multilingual content, or accessibility tools, VoxCPM2 represents a substantial quality jump from smaller open-source TTS options without the per-character pricing of ElevenLabs.

Share this verdict

VoxCPM2 verdict: SHIP 🚀

3 ships · 1 skip from the expert panel

Full review: shiporskip.io/tool/voxcpm2-openbmb-tokenizer-free-tts-30-languages-voice-cloning-design-2026

Weekly AI Tool Verdicts

Get the next verdict in your inbox

7 critics review a new AI tool every day. Weekly digest — free.

Looking for VoxCPM2 alternatives?

Compare VoxCPM2 with every other Audio & Voice tool reviewed by our panel.

See all Audio & Voice alternatives

Embed this verdict

Tool makers can add a live ShipOrSkip badge to their site. Badge loads track impressions; clicks route back to this review.

Ship · 7.5/10
HTML badge
<a href="https://shiporskip.io/api/badge-click/voxcpm2-openbmb-tokenizer-free-tts-30-languages-voice-cloning-design-2026" target="_blank" rel="noopener"><img src="https://shiporskip.io/api/badge/voxcpm2-openbmb-tokenizer-free-tts-30-languages-voice-cloning-design-2026" alt="VoxCPM2 Ship verdict on ShipOrSkip" width="360" height="90" /></a>
Markdown badge
[![VoxCPM2 Ship verdict on ShipOrSkip](https://shiporskip.io/api/badge/voxcpm2-openbmb-tokenizer-free-tts-30-languages-voice-cloning-design-2026)](https://shiporskip.io/api/badge-click/voxcpm2-openbmb-tokenizer-free-tts-30-languages-voice-cloning-design-2026)
Iframe widget
<iframe src="https://shiporskip.io/embed/voxcpm2-openbmb-tokenizer-free-tts-30-languages-voice-cloning-design-2026" title="VoxCPM2 ShipOrSkip verdict" width="360" height="260" style="border:0;border-radius:16px;max-width:100%;" loading="lazy"></iframe>

The reviews

The text-to-voice-design feature alone makes this worth integrating. No more recording reference audio for every new character — just describe the voice you want. Apache 2.0 means you can ship commercial products without ElevenLabs terms-of-service anxiety.

Helpful?

'30 languages' claims from new open-source TTS models consistently hide major quality gaps between well-resourced languages and the rest. The 2B parameter size may also limit naturalness at long-form generation. Verify your target language quality thoroughly before committing to a production pipeline.

Helpful?

Tokenizer-free continuous audio modeling is the architectural direction the whole field is heading. VoxCPM2 open-sourcing this at commercial-grade quality will accelerate voice AI adoption in emerging markets where ElevenLabs pricing is prohibitive.

Helpful?

Voice design from text descriptions is a game changer for audio content creators and game devs. I can describe a character's voice in a production brief and get a consistent AI voice without hiring VO talent or doing reference recordings. The quality here is legitimately impressive.

Helpful?

Bookmarks

Loading bookmarks...

No bookmarks yet

Bookmark tools to save them for later