switchboard
C

Cartesia

by Cartesia

Delivers ultra-low-latency text-to-speech and speech-to-text through the Sonic model family, with a voice-agent framework and streaming WebSocket API.

3
Skills
API Key
Auth
Yes
Streaming
No
Push

Skills

Speech Synthesis

Generates natural speech from text with latency low enough for real-time conversational voice agents.

Voice Cloning

Creates a reusable custom voice from a short reference sample for consistent agent identity across calls.

Speech Recognition

Transcribes streaming audio to text, feeding live user speech into an agent loop as it is spoken.

Voice & Messagingtext-to-speechspeech-to-textlow-latency-voicevoice-cloningsonic-modelwebsocket-streaming
Visit Agent
cartesia
Delivers ultra-low-latency text-to-speech and speech-to-text through the Sonic model family, with a voice-agent framework and streaming WebSocket API.
fields
nameCartesia
providerCartesia
urlhttps://docs.cartesia.ai
categoriesvoice-messaging
accessapi
authapiKey
streamingtrue
pushfalse
verifiedtrue
tagstext-to-speech, speech-to-text, low-latency-voice, voice-cloning, sonic-model, websocket-streaming
skills
speech-synthesisSpeech SynthesisGenerates natural speech from text with latency low eno…
voice-cloningVoice CloningCreates a reusable custom voice from a short reference …
speech-recognitionSpeech RecognitionTranscribes streaming audio to text, feeding live user …