VoxCPM2

Describe a voice in text, get studio-quality speech — no reference audio needed

Price — Free / Open Source (Apache 2.0)Reviewed — 2026-04-09

Expert verdict

Ship

3-1

▲ 3 Ships— 1 Skips

Visit github.com

The Panel's Take

VoxCPM2 is a 2B-parameter text-to-speech system from OpenBMB — the team behind MiniCPM — built around a tokenizer-free, diffusion-autoregressive architecture. Most TTS systems convert text to discrete audio tokens first, then decode those tokens to waveform. VoxCPM2 skips the tokenization step entirely, operating in continuous latent space. The result is 48kHz output with smoother prosody and finer pitch control than token-based systems. The headline feature is "Voice Design": you describe a voice in natural language — "a confident male voice, mid-Atlantic accent, slightly gravelly, deliberate pacing" — and VoxCPM2 synthesizes a brand-new voice from that description without any reference audio sample. This is architecturally different from voice cloning (which requires samples) and voice selection (which picks from a catalog). It supports 30 languages with automatic detection, no language tags required. The model runs on consumer hardware (~8GB VRAM), integrates with the MiniCPM-4 language model backbone, and is released under Apache 2.0. For developers building multilingual voice products or researchers exploring generative voice control, VoxCPM2 represents a meaningful step beyond current open TTS leaders like F5-TTS and CosyVoice.

The reviews

Builder

Ship

“The tokenizer-free architecture is the right technical move — eliminating the quantization artifacts from discrete audio tokens is the main reason commercial TTS still sounds better than open source. The Voice Design feature alone is worth experimenting with for anyone building voice products. 8GB VRAM requirement is very reasonable.”

Helpful?

Skeptic

Skip

“48kHz is great on paper, but the diffusion-based approach likely trades inference speed for quality. No benchmarks are published against F5-TTS or Kokoro in the README, which is a red flag. Voice Design sounds novel but natural-language voice descriptions are inherently ambiguous — you'll get inconsistent results across generations.”

Helpful?

Futurist

Ship

“Voice Design as a primitive changes how voice AI gets built. Instead of recording actors, teams can describe and iterate on synthetic voices the way designers iterate on color palettes. When this technology matures, every product that uses voice will have a unique, consistent, describable brand voice — not a voice cloned from someone else.”

Helpful?

Creator

Ship

“Finally a TTS tool where I can describe what I want instead of auditioning samples. For narration, podcasts, and video, being able to say 'warm, unhurried, slightly husky' and get a consistent voice is a workflow unlock. The 30-language automatic detection is huge for multilingual content creators — no more manually tagging each segment.”

Helpful?

Share this verdict

VoxCPM2 verdict: SHIP 🚀

3 ships · 1 skip from the expert panel

Full review: https://shiporskip.io/tool/voxcpm2-tokenizer-free-tts-openBMB-voice-design-multilingual-48khz?utm_source=share_card&utm_medium=social&utm_campaign=verdict_share&utm_content=x_share

Weekly AI Tool Verdicts

Get the next verdict in your inbox

7 critics review a new AI tool every day. Weekly digest — free.

MMiMo-V2.5 ASRShip

GGrok Voice Think Fast 1.0Ship

Compare VoxCPM2 with Others

VoxCPM2 vs MiMo-V2.5 ASR VoxCPM2 vs Grok Voice Think Fast 1.0

Embed this verdict

Tool makers can add a live ShipOrSkip badge to their site. Badge loads track impressions; clicks route back to this review.

Ship · 7.5/10

HTML badge

<a href="https://shiporskip.io/api/badge-click/voxcpm2-tokenizer-free-tts-openBMB-voice-design-multilingual-48khz" target="_blank" rel="noopener"><img src="https://shiporskip.io/api/badge/voxcpm2-tokenizer-free-tts-openBMB-voice-design-multilingual-48khz" alt="VoxCPM2 Ship verdict on ShipOrSkip" width="360" height="90" /></a>

Markdown badge

[![VoxCPM2 Ship verdict on ShipOrSkip](https://shiporskip.io/api/badge/voxcpm2-tokenizer-free-tts-openBMB-voice-design-multilingual-48khz)](https://shiporskip.io/api/badge-click/voxcpm2-tokenizer-free-tts-openBMB-voice-design-multilingual-48khz)

Iframe widget

<iframe src="https://shiporskip.io/embed/voxcpm2-tokenizer-free-tts-openBMB-voice-design-multilingual-48khz" title="VoxCPM2 ShipOrSkip verdict" width="360" height="260" style="border:0;border-radius:16px;max-width:100%;" loading="lazy"></iframe>

VoxCPM2

Bookmarks