VoxLoom — Free AI Voice Studio Flagship

The AI voice studio — clone any voice, generate speech, dictate anywhere. Free to download and use for a limited time, running locally.

v0.5.0 · First released August 1, 2026

macOS Windows Linux
Try a voice

Hear the local difference

Pick a preset from VoxLoom — Free AI Voice Studio. Full cloning and unlimited generation run in the desktop app — on your machine, not in the cloud.

"Welcome to VoxLoom — your voice, your machine, no cloud required."

Audio samples play here when available. Clone any voice in seconds inside the app.

Compare

The local alternative

VoxLoom is not a cloud freemium voice changer. It is a free, local-first workstation — cloning, dictation, TTS, and agent voices in one app.

Feature VoxLoomElevenLabsWisprFlowDubbing AI
Pricing Free (limited time) $5–330/mo $144/yr Freemium + sub
Runs locally
Voice cloning Zero-shot, offline Cloud Cloud, real-time
Text-to-speech 7 engines · 23 langs Cloud 500+ presets
System dictation Cloud
AI agent voices (MCP)
Local REST API No keys · no limits Metered cloud SDK only
Privacy 100% on-device Cloud processing Cloud processing Cloud processing
Best for Creators · devs · privacy Studio dubbing Writing flow Gaming · memes
Quick start

Up and running in five minutes

Six small steps from download to your first locally generated voice with VoxLoom — Free AI Voice Studio.

1

Download & install

Grab the installer for macOS, Windows, or Linux. No account required.

~1 min
2

Import or clone a voice

Drop a short audio clip, record 10 seconds, or pick a preset voice.

~30 sec
3

Generate your first line

Type any text, pick an engine, and hit Generate — all on your machine.

~10 sec
4

Set up dictation

Assign a global hotkey. Hold, speak, release — text lands in any app.

~1 min
5

Route to your apps

Optional: point Discord, OBS, or your IDE MCP client at VoxLoom.

~30 sec

You're creating

Clone, dictate, compose, and build — unlimited, offline, on your hardware.

Done
Built for real work

Local power, zero compromise

7 TTS engines One-click switch
23 Languages Multilingual cloning
0 API keys Local REST endpoint
100% On-device Privacy by design
Import Voices

Import a voice in 3 ways

Upload audio files, record with your microphone, or create a voice from a sample text. Every voice you import is yours — stored locally on your machine.

The voices included with Aispanvok were created with the same technology you get to use. Import your own and they are added to your library instantly.

No uploads. No cloud. Just your files, on your machine.

Upload a clip Microphone System Audio
🎵 Drag & drop any audio file — WAV, MP3, FLAC, or WebM
my_voice_sample.mp3 00:32 · 44.1kHz
narration.wav 01:15 · 48kHz
Import voice
Dictation

Dictate anywhere in your system

Press a shortcut to start dictating into any text field — notes, editors, chats. Aispanvok transcribes with Whisper and types for you.

Hold, speak, release — from anywhere on your machine, into any app.

Raw and refined transcripts kept side by side, with the original audio forever.

Agents talk back through the same pill, in any cloned voice.

Hold on macOS, CtrlAlt on Windows
Recording 0:00
Whisper Base 74M Small 244M Medium 769M Large 1.5B Turbo 809M 99 langs

Whisper, sized for every machine

Base, Small, Medium, Large, and Turbo. Pick the size that fits your hardware — 99 languages across every tier, all running locally.

Refined transcripts

A local LLM cleans ums, self-corrections, and punctuation without rephrasing. Optional, toggleable, and never leaves your machine.

Agents speak in voices you own

Any MCP-aware agent gets a voice with one tool call. The pill surfaces when an agent is speaking, so you always see what's coming out of your machine.

Agent Integration

Bring your agent to life

Give your AI agents a voice with the Aispanvok MCP server. One tool call and any MCP-aware agent — Claude Code, Cursor, Cline — speaks in a voice you cloned.

Transparent — every generation is logged locally.

Credible — clone your own voice for your assistant.

Customizable — pick any personality for your agent.

Terminal
$ claude run
✓ Tests passing (42 files)
✓ Build succeeded in 12.4s
→ aispanvok.speak({ profile: "Morgan" })

Per-agent voice

Bind each MCP client to a voice profile. Claude Code in Morgan, Cursor in Scarlett — you know which agent is talking without looking.

Always visible

Every agent-initiated speech surfaces the pill. No silent background TTS — you always see what's coming out of your machine.

Open protocols

MCP ships day one. ACP, A2A, and anything else built on a tool-call primitive slots into the same endpoint.

Personalities

Voices with a personality

Give any voice profile a free-form personality. Then Rewrite your text in their voice, or let them Compose a fresh line of their own — your cloned voice, in full character.

🎭  Compose voices from scratch

✍️  Rewrite existing voices into new characters

✨  Apply voice effects on top

🕵️
Marlowe Voice profile · cloned from a 12s sample
Personality

“1940s noir detective. World-weary, cynical, every situation a metaphor for the city's underbelly. Talks like he's seen one stack trace too many.”

Rewrite Compose
In character · Marlowe

“Build's wrapped, ship's left the dock. Another stack of code makes its way into prod, another row of green checks lining the wall.”

Rewrite

Restate your text in their voice while preserving every idea. Same content, their delivery — for scripts, dubs, and consistent character voice across long-form work.

Compose

No input needed — hit the button and the character improvises a fresh line of their own. Roll again for another take. Useful for game dialogue, narration cues, or character barks.

Built-in REST API

Your local voice API

Every engine you download becomes a REST endpoint on your machine. Build apps, games, and voice tools with full programmatic control — no API keys, no rate limits, no per-character fees.

🛠️  Use it with any language or framework

⚙️  Automate your voice workflows

📴  Fully offline-capable

API Reference http://127.0.0.1:17493
POST/generateGenerate speech
POST/generate/:id/cancelCancel a generation
GET/profilesList voice profiles
POST/profilesCreate a new profile
GET/models/statusModel catalog & state
GET/historyPast generations
$ curl -X POST http://127.0.0.1:17493/generate \
  -H "Content-Type: application/json" \
  -d '{"text": "Welcome to the game, player one.", "profile": "morgan"}' \
  --output line.wav
No API keys No rate limits No per-character fees Works offline Your audio, your machine

Games

Generate NPC dialogue on the fly, localize characters into new languages, or ship expressive voice lines without a studio.

Apps & agents

Give your app or AI agent a voice. Real-time narration, accessibility readouts, voice replies — all running on the user's machine.

Scripts & tools

Batch-generate audiobook chapters, automate podcast intros, or wire Aispanvok into your Stream Deck. It's just a localhost URL.

Supported Models

Your favorite models, out of the box

VoxLoom bundles 7 TTS engines, Whisper, and a local LLM into one app — all interchangeable from the same interface, no separate installs.

🔊 Text-to-Speech
Qwen3-TTS Multilingual cloning, voice instructions
Qwen CustomVoice 50+ preset voices, delivery control
Chatterbox Multilingual Conversational multilingual speech
Chatterbox Turbo Expressive tags — [laugh], [sigh]
LuxTTS Ultra-light, 48kHz, CPU-friendly
HumeAI TADA Emotionally expressive delivery
Kokoro Compact and expressive voices
🎙️ Speech-to-Text
Whisper Base → Large, Turbo — 99 languages
🧠 LLM
Qwen3 1.7B Bundled local LLM for refinement & personas
Why VoxLoom — Free AI Voice Studio

Everything you need, on your machine

Zero-Shot Voice Cloning

Clone any voice from a few seconds of audio — no fine-tuning, no cloud. Works offline with Metal, CUDA, and ROCm acceleration.

7 TTS Engines, 23 Languages

Qwen3-TTS, Chatterbox, Kokoro, LuxTTS, and more — switch engines with one click and generate speech from English to Arabic, Japanese, and Hindi.

System-Wide Dictation

Hold a hotkey, talk, release: Whisper-based dictation lands in any text field, with raw and LLM-refined transcripts kept side by side.

AI Agent Voices via MCP

One tool call and Claude Code, Cursor, or Cline speaks in a voice you cloned — each agent bound to its own voice profile.

Stories Editor & Effects

A multi-track timeline for podcasts and narratives, plus pitch shift, reverb, delay, chorus, and compression built in.

Unlimited Local Generation

No character limits, no API keys, no subscriptions. Auto-chunking with crossfade handles books and full chapters.

Highlights

Clone any voice from a few seconds of audio — zero-shot, no fine-tuning — and generate speech across 7 TTS engines in 23 languages, all on your own hardware.

Dictate into any app with a global hotkey: hold, speak, release. Whisper-based STT keeps raw and refined transcripts side by side, forever searchable.

Give AI agents a voice: one MCP tool call (voxloom.speak) and Claude Code, Cursor, or Cline speaks in a voice you own.

Free to download and use for a limited time. No accounts, no API keys, no per-character fees, no rate limits — just install and run.

Built with Tauri (Rust) for native performance. Runs on macOS (Metal/MLX), Windows (CUDA), Linux, AMD ROCm, Intel Arc, and Docker.

Download

Hosted on GitHub Releases — these links always point to the latest version.

What is VoxLoom?

VoxLoom is a local-first AI voice studio — a free alternative to ElevenLabs and WisprFlow in one native desktop app. Clone any voice from seconds of reference audio, generate speech in 23 languages across 7 TTS engines, dictate into any text field with a global hotkey, and give any MCP-aware AI agent a voice of your choosing.

The two cloud incumbents each own one half of the voice I/O loop — ElevenLabs on output, WisprFlow on input. VoxLoom does both, bridges them with a bundled local LLM for transcript refinement and per-profile voice personas, and runs the whole stack on your machine. Your voices, transcripts, and generated audio never leave your device.

Why local-first matters

Cloud voice services charge per character, rate-limit your generations, and process one of your most personal assets — your voice — on someone else’s servers. VoxLoom takes the opposite position:

  • Complete privacy — models, voice data, and captures never leave your machine
  • No meter running — unlimited local generation, no API keys, no quotas
  • Free for a limited time — download and run on your own machine, no account or subscription required

7 TTS engines, 23 languages

Switch engines with one click — no separate installs, no Python environments.

EngineBest for
Qwen3-TTS (0.6B / 1.7B)High-quality multilingual cloning with voice instructions
Qwen CustomVoice50+ curated preset voices, natural-language delivery control
LuxTTSUltra-lightweight (~1GB VRAM), 48kHz, CPU-friendly
Chatterbox MultilingualConversational multilingual speech
Chatterbox TurboExpressive speech with paralinguistic tags — [laugh], [sigh], [gasp]
HumeAI TADAEmotionally expressive delivery
KokoroCompact, fast, expressive preset voices

From English to Arabic, Japanese, Hindi, Swahili, and more — 23 languages across the engine catalog.

Core capabilities

Zero-shot voice cloning

Clone a voice from a few seconds of reference audio — no fine-tuning required. Metal, CUDA, and ROCm hardware acceleration keep generation fast, and everything works fully offline.

Dictate anywhere

A global dictation hotkey works system-wide: hold to talk, release, and your words land in whatever field is focused. Push-to-talk and toggle modes, accessibility-verified auto-paste on macOS, and an in-app mic on every text field. Whisper-based STT in five sizes — pick the one that fits your hardware.

Every capture keeps its original audio alongside raw and LLM-refined transcripts, archived and fully searchable.

Give your AI agents a voice

VoxLoom ships with a built-in MCP server. One tool call — voxloom.speak — and any MCP-aware agent (Claude Code, Cursor, Cline) speaks to you in a voice you’ve cloned. Bind each agent client to a different voice profile so you always know who is talking.

Voice personalities

Attach a free-form persona to any voice profile — a noir detective, a game character, your own brand voice. Then Compose fresh lines, Rewrite your text in character, or Respond to messages, powered by a bundled local LLM. Agents can invoke the same modes over MCP.

Unlimited length, professional effects

Auto-chunking with crossfade generates scripts, articles, and full chapters without interruption. Post-processing effects include pitch shift, reverb, delay, chorus, compression, and filters — a complete audio effects pipeline for professional results.

Stories editor

A multi-track timeline editor for conversations, podcasts, and narratives. Arrange voices, music, and sound effects on parallel tracks and export production-ready audio.

VoxLoom vs. cloud services

VoxLoomElevenLabsWisprFlow
Voice cloning✅ Local, zero-shot✅ Cloud
Speech generation✅ 7 engines, 23 languages✅ Cloud
System-wide dictation✅ Cloud
AI agent voices (MCP)
REST API✅ Local, no keys✅ Metered
PricingFree (limited time)$5–330/mo$144/yr
Privacy100% localCloudCloud

Platform support

PlatformAccelerationStatus
macOS (Apple Silicon / Intel)Metal / MLXReleased
Windows x64 (MSI)CUDAReleased
Linux x64CUDA / ROCm / Intel ArcBuild from source
Docker (any GPU)CUDA / ROCmdocker compose up

Download & updates

Installers are served from the edge with resumable downloads and permanent links. See the download page for all platforms.

FAQ

Frequently asked questions

What is VoxLoom?

VoxLoom is a free, local-first AI voice studio that runs locally on your machine. It combines zero-shot voice cloning, text-to-speech across 7 engines and 23 languages, system-wide dictation, and AI agent voices via MCP — a local alternative to both ElevenLabs and WisprFlow.

Is VoxLoom free?

Yes — free to download and use for a limited time. No accounts, no subscriptions, no per-character fees, and no rate limits. Unlimited local generation on your own hardware.

Where is my voice data stored?

Locally on your machine. Voices, transcripts, and generated audio never leave your device — privacy is the architecture, not a setting.

Which TTS engines are supported?

Seven engines ship in one app: Qwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox Multilingual, Chatterbox Turbo, HumeAI TADA, and Kokoro — plus Whisper for speech-to-text and a bundled local LLM for transcript refinement.

Can AI agents like Claude Code speak with my cloned voice?

Yes. VoxLoom ships with a built-in MCP server — a single voxloom.speak tool call lets any MCP-aware agent (Claude Code, Cursor, Cline) speak in a voice you own, with each agent bound to its own voice profile.