Miso Labs
Foundation voice models for building voice agents
Miso Labs provides Miso-TTS, a text-to-speech foundation model designed for voice agents that need fast, expressive speech output. It addresses latency in conversational AI, reporting 110ms real-time response, and supports one-shot voice cloning from a 10-second audio sample. The models are open source and can be deployed on-premises for teams that need local hosting.
Core capabilities include streaming and WebSocket SDKs, a public voice library, custom voice support, and dedicated voice fine-tuning. The platform targets builders and teams shipping voice agents to production, with plans for regulated workloads and high-volume use. Packaging is usage-based across Starter, Scale, and Enterprise editions, each with included monthly audio minutes and per-minute overage rates. Custom voices and on-premises deployment are available on Enterprise, which is offered under an annual contract with volume pricing.
12 alternatives to Miso Labs
Ranked by how well each tool replaces Miso Labs: shared features, audience, price and popularity.
Bringing technology to life
Covers 2 of 13 key features, has a free plan and starts cheaper.
Free plan66 out of 100 match$6/mo- 64 out of 100 match$9/mo
Turn text into natural-sounding speech with AI
Covers 1 of 13 key features and has a free plan.
Free plan63 out of 100 matchContact salesSpeech services for building apps that transcribe, translate, and synthesize speech
Covers 1 of 13 key features and has a free plan.
Free plan63 out of 100 matchUsage-basedSpeech-to-text, text-to-speech and voice agent APIs for real-time and batch audio, usable
Covers 1 of 13 key features and has a free plan.
Free plan63 out of 100 matchUsage-based- 63 out of 100 matchUsage-based
Speech-to-text and speech understanding APIs for voice AI applications
Covers 0 of 13 key features and has a free plan.
Free plan62 out of 100 matchContact salesCreate professional-quality voice overs in any dialect or production style with our secure
Covers 1 of 13 key features, has a free plan and starts cheaper.
Free plan62 out of 100 match$10/moPowerful Speech Platform – Text to Speech API, Speech Recognition API, Open Source SDKs
Covers 2 of 13 key features and has a free plan.
Free plan62 out of 100 matchUsage-basedStudio-grade AI text-to-speech and voice cloning with emotion control and a large voice
Covers 0 of 13 key features, has a free plan and starts cheaper.
Free plan61 out of 100 match$11/moFrontier voice AI company building the world's first Ensemble Listening Model that outperf
Covers 1 of 13 key features and has a free plan.
Free plan60 out of 100 matchUsage-basedThe data and evaluation layer for emotionally intelligent voice AI.
Covers 0 of 13 key features, has a free plan and starts cheaper.
Free plan60 out of 100 match$3/mo