Cartesia
by Cartesia
Delivers ultra-low-latency text-to-speech and speech-to-text through the Sonic model family, with a voice-agent framework and streaming WebSocket API.
Skills
Speech Synthesis
Generates natural speech from text with latency low enough for real-time conversational voice agents.
Voice Cloning
Creates a reusable custom voice from a short reference sample for consistent agent identity across calls.
Speech Recognition
Transcribes streaming audio to text, feeding live user speech into an agent loop as it is spoken.
Related Agents
AssemblyAI
Speech recognition API with real-time and batch transcription plus audio intelligence models for summarization, topic d…
Deepgram
Speech-to-text and voice AI API with real-time streaming transcription, speaker diarization, and text-to-speech for voi…
Bland AI
Deploys enterprise AI phone agents that handle inbound and outbound calls in 40+ languages with SOC2, HIPAA, and PCI DS…
Dograh
Builds and self-hosts voice AI agents with a visual workflow builder, telephony, and bring-your-own-key speech-to-speec…