Voicebox
by Jamie Pine
Runs a local voice studio for voice cloning, text-to-speech and Whisper dictation, and lets agents speak or transcribe through a built-in MCP server and REST API.
Skills
Agent Voice Output
Lets any MCP-aware agent speak in a cloned voice via the voicebox_speak tool or POST /speak, with a per-agent voice binding.
Voice Cloning
Clones a voice from a few seconds of audio and generates speech in 23 languages across seven TTS engines, with multi-sample profiles.
Whisper Transcription
Transcribes audio with OpenAI Whisper on MLX or PyTorch, exposed through the /transcribe endpoint and the voicebox_transcribe MCP tool.
Global Dictation
Dictates into any text field with a push-to-talk or toggle hotkey, pasting the transcript into the focused field on macOS.
Related Agents
Cartesia
Delivers ultra-low-latency text-to-speech and speech-to-text through the Sonic model family, with a voice-agent framewo…
AssemblyAI
Speech recognition API with real-time and batch transcription plus audio intelligence models for summarization, topic d…
Deepgram
Speech-to-text and voice AI API with real-time streaming transcription, speaker diarization, and text-to-speech for voi…
Bland AI
Deploys enterprise AI phone agents that handle inbound and outbound calls in 40+ languages with SOC2, HIPAA, and PCI DS…