switchboard
V

Voicebox

by Jamie Pine

Runs a local voice studio for voice cloning, text-to-speech and Whisper dictation, and lets agents speak or transcribe through a built-in MCP server and REST API.

4
Skills
None
Auth
No
Streaming
No
Push

Skills

Agent Voice Output

Lets any MCP-aware agent speak in a cloned voice via the voicebox_speak tool or POST /speak, with a per-agent voice binding.

Voice Cloning

Clones a voice from a few seconds of audio and generates speech in 23 languages across seven TTS engines, with multi-sample profiles.

Whisper Transcription

Transcribes audio with OpenAI Whisper on MLX or PyTorch, exposed through the /transcribe endpoint and the voicebox_transcribe MCP tool.

Global Dictation

Dictates into any text field with a push-to-talk or toggle hotkey, pasting the transcript into the focused field on macOS.

Voice & Messagingtext-to-speechvoice-cloningspeech-to-textdictationwhisperlocal-firstmcp-serverrest-api
Visit Agent
voicebox
Runs a local voice studio for voice cloning, text-to-speech and Whisper dictation, and lets agents speak or transcribe through a built-in MCP server and REST API.
fields
nameVoicebox
providerJamie Pine
urlhttps://github.com/jamiepine/voicebox
categoriesvoice-messaging
accessmcp · api
authnone
streamingfalse
pushfalse
verifiedtrue
tagstext-to-speech, voice-cloning, speech-to-text, dictation, whisper, local-first, mcp-server, rest-api
skills
agent-voice-outputAgent Voice OutputLets any MCP-aware agent speak in a cloned voice via th…
voice-cloningVoice CloningClones a voice from a few seconds of audio and generate…
speech-to-textWhisper TranscriptionTranscribes audio with OpenAI Whisper on MLX or PyTorch…
global-dictationGlobal DictationDictates into any text field with a push-to-talk or tog…