Back
TechCrunchFundingTechCrunch2026-07-31

Smallest.ai Raises $13M to Make AI Phone Calls Pass the Turing Test

Smallest.ai has raised $13M to build ultra-low-latency voice AI models designed to make AI-driven phone calls indistinguishable from human ones. The startup is targeting the gap between current robotic-sounding voice AI and genuinely natural conversation.

Original source

Smallest.ai announced a $13 million funding round aimed at building voice models that can pass the Turing test in real phone call scenarios. The company's core technical bet is on latency and prosody — the two variables most responsible for the uncanny valley effect in current voice AI deployments. Most enterprise voice AI today fails not because it says the wrong thing, but because it sounds wrong: response gaps, flat intonation, and unnatural pacing that signals 'bot' within the first exchange.

The startup's approach centers on ultra-fast inference, targeting sub-200ms response latency end-to-end — a threshold the team argues is where human-like conversational rhythm begins. Getting there requires co-designing the model architecture with the inference stack rather than bolting a smaller model onto existing serving infrastructure, which is the more common approach taken by competitors wrapping existing TTS providers.

The voice AI market is increasingly crowded, with ElevenLabs, Cartesia, Deepgram, and a growing field of real-time voice API providers all competing on similar latency and quality claims. What distinguishes Smallest.ai's pitch is the emphasis on phone-call-specific use cases — a context with hard audio codec constraints (G.711, 8kHz sampling) and ambient noise variables that browser-based demos don't surface. Whether the models hold up under real PSTN conditions at scale is the question the funding will have to answer.

The $13M raise will reportedly go toward model training, inference infrastructure, and expanding the team. The company has not yet published a public benchmark methodology or third-party evaluation, which means the 'passes the Turing test' claim remains marketing until independently verified.

Panel Takes

The Builder

The Builder

Developer Perspective

The primitive here is a real-time voice inference API optimized for PSTN constraints — that's a specific enough problem that I'm not immediately dismissing it as a wrapper. The question I'd have in the first 10 minutes is whether their SDK abstracts the hard parts (codec negotiation, jitter buffers, SIP integration) or just wraps an HTTP endpoint and leaves the telephony plumbing to me. No public docs yet means I can't tell if the DX bet was the right one or if this is another 'request a demo' black box dressed up as an API.

The Skeptic

The Skeptic

Reality Check

'Passes the Turing test' is doing a lot of work in this pitch — it's a claim with no published methodology, no third-party eval, and a definition that conveniently shifts based on who's testing. The direct competitors here are Cartesia and ElevenLabs, both of which have public latency numbers and audio demos you can actually listen to; Smallest.ai has neither. My prediction for what kills this in 12 months: the underlying model providers ship native low-latency telephony endpoints, collapsing the differentiation to zero before the team can build enough workflow lock-in to survive.

The Futurist

The Futurist

Big Picture

The falsifiable thesis here is: 'by 2028, the majority of inbound and outbound business phone calls will be AI-handled, and the bottleneck is perceptual realism, not capability.' That's a real bet, and sub-200ms PSTN-native inference is exactly the infrastructure layer that bet requires — not a wrapper, but a rearchitecting of the stack for a world where phone calls are a compute problem. The second-order effect nobody's talking about: if AI calls become truly indistinguishable, call authentication and consent infrastructure becomes critical, and whoever owns the trust layer in that stack holds more power than the voice model provider.

The Founder

The Founder

Business & Market

The buyer here is clear — outbound sales and customer service ops teams with per-minute telephony budgets they're already spending on human agents or legacy IVR vendors — and that's a real budget line with real ROI math attached. The moat question is harder: co-designed inference and model architecture is defensible for maybe 18 months before the infrastructure commoditizes, so the real question is whether $13M is enough runway to accumulate enough customer workflow integration before ElevenLabs or a well-capitalized competitor closes the latency gap. If they can land a few enterprise accounts deep enough into their CRM and dialer stack, the switching cost story gets real fast.

Bookmarks

Loading bookmarks...

No bookmarks yet

Bookmark tools to save them for later