Zero-Shot Voice Cloning
Clone any voice from a few seconds of audio — no fine-tuning, no cloud. Works offline with Metal, CUDA, and ROCm acceleration.
The AI voice studio — clone any voice, generate speech, dictate anywhere. Free to download and use for a limited time, running locally.
v0.5.0 · First released August 1, 2026
Pick a preset from VoxLoom — Free AI Voice Studio. Full cloning and unlimited generation run in the desktop app — on your machine, not in the cloud.
"Welcome to VoxLoom — your voice, your machine, no cloud required."
Audio samples play here when available. Clone any voice in seconds inside the app.
VoxLoom is not a cloud freemium voice changer. It is a free, local-first workstation — cloning, dictation, TTS, and agent voices in one app.
| Feature | VoxLoom | ElevenLabs | WisprFlow | Dubbing AI |
|---|---|---|---|---|
| Pricing | Free (limited time) | $5–330/mo | $144/yr | Freemium + sub |
| Runs locally | ✓ | — | — | — |
| Voice cloning | Zero-shot, offline | Cloud | — | Cloud, real-time |
| Text-to-speech | 7 engines · 23 langs | Cloud | — | 500+ presets |
| System dictation | ✓ | — | Cloud | — |
| AI agent voices (MCP) | ✓ | — | — | — |
| Local REST API | No keys · no limits | Metered cloud | — | SDK only |
| Privacy | 100% on-device | Cloud processing | Cloud processing | Cloud processing |
| Best for | Creators · devs · privacy | Studio dubbing | Writing flow | Gaming · memes |
Six small steps from download to your first locally generated voice with VoxLoom — Free AI Voice Studio.
Grab the installer for macOS, Windows, or Linux. No account required.
Drop a short audio clip, record 10 seconds, or pick a preset voice.
Type any text, pick an engine, and hit Generate — all on your machine.
Assign a global hotkey. Hold, speak, release — text lands in any app.
Optional: point Discord, OBS, or your IDE MCP client at VoxLoom.
Clone, dictate, compose, and build — unlimited, offline, on your hardware.
Upload audio files, record with your microphone, or create a voice from a sample text. Every voice you import is yours — stored locally on your machine.
The voices included with Aispanvok were created with the same technology you get to use. Import your own and they are added to your library instantly.
No uploads. No cloud. Just your files, on your machine.
Click to record from your microphone. Maximum duration: 30 seconds.
Press a shortcut to start dictating into any text field — notes, editors, chats. Aispanvok transcribes with Whisper and types for you.
Hold, speak, release — from anywhere on your machine, into any app.
Raw and refined transcripts kept side by side, with the original audio forever.
Agents talk back through the same pill, in any cloned voice.
Base, Small, Medium, Large, and Turbo. Pick the size that fits your hardware — 99 languages across every tier, all running locally.
A local LLM cleans ums, self-corrections, and punctuation without rephrasing. Optional, toggleable, and never leaves your machine.
Any MCP-aware agent gets a voice with one tool call. The pill surfaces when an agent is speaking, so you always see what's coming out of your machine.
Give your AI agents a voice with the Aispanvok MCP server. One tool call and any MCP-aware agent — Claude Code, Cursor, Cline — speaks in a voice you cloned.
Transparent — every generation is logged locally.
Credible — clone your own voice for your assistant.
Customizable — pick any personality for your agent.
$ claude run ✓ Tests passing (42 files) ✓ Build succeeded in 12.4s → aispanvok.speak({ profile: "Morgan" })
Bind each MCP client to a voice profile. Claude Code in Morgan, Cursor in Scarlett — you know which agent is talking without looking.
Every agent-initiated speech surfaces the pill. No silent background TTS — you always see what's coming out of your machine.
MCP ships day one. ACP, A2A, and anything else built on a tool-call primitive slots into the same endpoint.
Give any voice profile a free-form personality. Then Rewrite your text in their voice, or let them Compose a fresh line of their own — your cloned voice, in full character.
🎭 Compose voices from scratch
✍️ Rewrite existing voices into new characters
✨ Apply voice effects on top
“1940s noir detective. World-weary, cynical, every situation a metaphor for the city's underbelly. Talks like he's seen one stack trace too many.”
“Build's wrapped, ship's left the dock. Another stack of code makes its way into prod, another row of green checks lining the wall.”
“Some days the city hums clean and the tests all pass. Not tonight. Tonight a flaky build walks into my CI pipeline, and nothing green comes out the other side.”
Restate your text in their voice while preserving every idea. Same content, their delivery — for scripts, dubs, and consistent character voice across long-form work.
No input needed — hit the button and the character improvises a fresh line of their own. Roll again for another take. Useful for game dialogue, narration cues, or character barks.
Every engine you download becomes a REST endpoint on your machine. Build apps, games, and voice tools with full programmatic control — no API keys, no rate limits, no per-character fees.
🛠️ Use it with any language or framework
⚙️ Automate your voice workflows
📴 Fully offline-capable
http://127.0.0.1:17493 /generateGenerate speech/generate/:id/cancelCancel a generation/profilesList voice profiles/profilesCreate a new profile/models/statusModel catalog & state/historyPast generations$ curl -X POST http://127.0.0.1:17493/generate \
-H "Content-Type: application/json" \
-d '{"text": "Welcome to the game, player one.", "profile": "morgan"}' \
--output line.wav Generate NPC dialogue on the fly, localize characters into new languages, or ship expressive voice lines without a studio.
Give your app or AI agent a voice. Real-time narration, accessibility readouts, voice replies — all running on the user's machine.
Batch-generate audiobook chapters, automate podcast intros, or wire Aispanvok into your Stream Deck. It's just a localhost URL.
VoxLoom bundles 7 TTS engines, Whisper, and a local LLM into one app — all interchangeable from the same interface, no separate installs.
Clone any voice from a few seconds of audio — no fine-tuning, no cloud. Works offline with Metal, CUDA, and ROCm acceleration.
Qwen3-TTS, Chatterbox, Kokoro, LuxTTS, and more — switch engines with one click and generate speech from English to Arabic, Japanese, and Hindi.
Hold a hotkey, talk, release: Whisper-based dictation lands in any text field, with raw and LLM-refined transcripts kept side by side.
One tool call and Claude Code, Cursor, or Cline speaks in a voice you cloned — each agent bound to its own voice profile.
A multi-track timeline for podcasts and narratives, plus pitch shift, reverb, delay, chorus, and compression built in.
No character limits, no API keys, no subscriptions. Auto-chunking with crossfade handles books and full chapters.
Clone any voice from a few seconds of audio — zero-shot, no fine-tuning — and generate speech across 7 TTS engines in 23 languages, all on your own hardware.
Dictate into any app with a global hotkey: hold, speak, release. Whisper-based STT keeps raw and refined transcripts side by side, forever searchable.
Give AI agents a voice: one MCP tool call (voxloom.speak) and Claude Code, Cursor, or Cline speaks in a voice you own.
Free to download and use for a limited time. No accounts, no API keys, no per-character fees, no rate limits — just install and run.
Built with Tauri (Rust) for native performance. Runs on macOS (Metal/MLX), Windows (CUDA), Linux, AMD ROCm, Intel Arc, and Docker.
Clone voices, generate narration, and edit multi-track stories — all locally. VoxLoom is the free voice studio for YouTubers, podcasters, and audiobook makers.
Learn more →Give Claude Code, Cursor, and Cline a voice you own. VoxLoom ships MCP and a local REST API — no API keys, no rate limits.
Learn more →Your voice is personal data. VoxLoom processes everything on-device — no accounts, no cloud uploads, no third-party training.
Learn more →Dictate into any app with a global hotkey. Whisper transcription plus optional LLM cleanup — all local, all private.
Learn more →Multi-track Stories editor, built-in effects, and unlimited local generation. Produce podcast episodes without a monthly voice bill.
Learn more →Generate speech in 23 languages with seven TTS engines. Localize characters and narration without re-recording or cloud API fees.
Learn more →Hosted on GitHub Releases — these links always point to the latest version.
VoxLoom is a local-first AI voice studio — a free alternative to ElevenLabs and WisprFlow in one native desktop app. Clone any voice from seconds of reference audio, generate speech in 23 languages across 7 TTS engines, dictate into any text field with a global hotkey, and give any MCP-aware AI agent a voice of your choosing.
The two cloud incumbents each own one half of the voice I/O loop — ElevenLabs on output, WisprFlow on input. VoxLoom does both, bridges them with a bundled local LLM for transcript refinement and per-profile voice personas, and runs the whole stack on your machine. Your voices, transcripts, and generated audio never leave your device.
Cloud voice services charge per character, rate-limit your generations, and process one of your most personal assets — your voice — on someone else’s servers. VoxLoom takes the opposite position:
Switch engines with one click — no separate installs, no Python environments.
| Engine | Best for |
|---|---|
| Qwen3-TTS (0.6B / 1.7B) | High-quality multilingual cloning with voice instructions |
| Qwen CustomVoice | 50+ curated preset voices, natural-language delivery control |
| LuxTTS | Ultra-lightweight (~1GB VRAM), 48kHz, CPU-friendly |
| Chatterbox Multilingual | Conversational multilingual speech |
| Chatterbox Turbo | Expressive speech with paralinguistic tags — [laugh], [sigh], [gasp] |
| HumeAI TADA | Emotionally expressive delivery |
| Kokoro | Compact, fast, expressive preset voices |
From English to Arabic, Japanese, Hindi, Swahili, and more — 23 languages across the engine catalog.
Clone a voice from a few seconds of reference audio — no fine-tuning required. Metal, CUDA, and ROCm hardware acceleration keep generation fast, and everything works fully offline.
A global dictation hotkey works system-wide: hold to talk, release, and your words land in whatever field is focused. Push-to-talk and toggle modes, accessibility-verified auto-paste on macOS, and an in-app mic on every text field. Whisper-based STT in five sizes — pick the one that fits your hardware.
Every capture keeps its original audio alongside raw and LLM-refined transcripts, archived and fully searchable.
VoxLoom ships with a built-in MCP server. One tool call — voxloom.speak — and any MCP-aware agent (Claude Code, Cursor, Cline) speaks to you in a voice you’ve cloned. Bind each agent client to a different voice profile so you always know who is talking.
Attach a free-form persona to any voice profile — a noir detective, a game character, your own brand voice. Then Compose fresh lines, Rewrite your text in character, or Respond to messages, powered by a bundled local LLM. Agents can invoke the same modes over MCP.
Auto-chunking with crossfade generates scripts, articles, and full chapters without interruption. Post-processing effects include pitch shift, reverb, delay, chorus, compression, and filters — a complete audio effects pipeline for professional results.
A multi-track timeline editor for conversations, podcasts, and narratives. Arrange voices, music, and sound effects on parallel tracks and export production-ready audio.
| VoxLoom | ElevenLabs | WisprFlow | |
|---|---|---|---|
| Voice cloning | ✅ Local, zero-shot | ✅ Cloud | ❌ |
| Speech generation | ✅ 7 engines, 23 languages | ✅ Cloud | ❌ |
| System-wide dictation | ✅ | ❌ | ✅ Cloud |
| AI agent voices (MCP) | ✅ | ❌ | ❌ |
| REST API | ✅ Local, no keys | ✅ Metered | ❌ |
| Pricing | Free (limited time) | $5–330/mo | $144/yr |
| Privacy | 100% local | Cloud | Cloud |
| Platform | Acceleration | Status |
|---|---|---|
| macOS (Apple Silicon / Intel) | Metal / MLX | Released |
| Windows x64 (MSI) | CUDA | Released |
| Linux x64 | CUDA / ROCm / Intel Arc | Build from source |
| Docker (any GPU) | CUDA / ROCm | docker compose up |
Installers are served from the edge with resumable downloads and permanent links. See the download page for all platforms.
VoxLoom is a free, local-first AI voice studio that runs locally on your machine. It combines zero-shot voice cloning, text-to-speech across 7 engines and 23 languages, system-wide dictation, and AI agent voices via MCP — a local alternative to both ElevenLabs and WisprFlow.
Yes — free to download and use for a limited time. No accounts, no subscriptions, no per-character fees, and no rate limits. Unlimited local generation on your own hardware.
Locally on your machine. Voices, transcripts, and generated audio never leave your device — privacy is the architecture, not a setting.
Seven engines ship in one app: Qwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox Multilingual, Chatterbox Turbo, HumeAI TADA, and Kokoro — plus Whisper for speech-to-text and a bundled local LLM for transcript refinement.
Yes. VoxLoom ships with a built-in MCP server — a single voxloom.speak tool call lets any MCP-aware agent (Claude Code, Cursor, Cline) speak in a voice you own, with each agent bound to its own voice profile.