Local voice for AI coding agents: text-to-speech and speech-to-text in one Rust binary.
No Python and no cloud service: synthesis and transcription run on your machine. Several TTS backends, Whisper speech-to-text, an MCP server with 14 tools, and one command that configures the AI tools installed on your machine (14 supported).
English • Français • 中文 • 日本語 • 한국어 • Español
vox
|
+-----------------+------------------+
| |
speak (TTS) hear (STT)
| |
+--------+---------+----------+ Whisper
| | | | (candle)
pocket piper qwen-native say 99 languages
(candle) (ONNX) (candle) (macOS) CPU, Metal or CUDA
| | | |
+---- rodio (playback) ----+ cpal (microphone)
say plays through the macOS say command. A fifth backend, kokoro, exists
only in a build compiled with --features kokoro (see below).
| Backend | Engine | Languages | Voice cloning | GPU | Where |
|---|---|---|---|---|---|
pocket |
Kyutai pocket-tts (100M) on candle | English | Yes¹ | No | All platforms. Default for English and when no language is given |
piper |
Piper on ONNX Runtime | 11² | No | No | All platforms. Default for every other language |
qwen-native |
Qwen3-TTS (0.6B) on candle | 10³ | Yes | In a Metal or CUDA build | All platforms |
say |
macOS /usr/bin/say |
System voices | No | No | macOS only |
kokoro |
Kokoro on ONNX Runtime | See --list-voices |
No | No | Only in a build compiled with --features kokoro |
¹
alba,marius,javert,jean,fantine,cosette,eponine,azelma) that need no setup: the public weights (~226 MB) are downloaded on first use. Voice cloning from a reference WAV needsHF_TOKENand the license of the gated kyutai/pocket-tts checkpoint accepted. The bundled checkpoint is English-only (all 8 voices are English speakers); Kyutai's per-language checkpoints (fr/de/es/it/pt) need upstream support in the pocket-tts crate and are not wired up.²
piperhas one default voice for each of en, fr, es, de, it, pt, zh, ko, ru, ar and nl, chosen by-l.-vwith the name of another voice of rhasspy/piper-voices selects that voice, for example-v fr_FR-siwis-low; with a value that is not a piper voice name, vox prints a note and uses the default voice of the language. A voice is downloaded the first time it is used (~60 MB for amediumvoice).piperhas no Japanese voice, so Japanese usesqwen-nativeby default;-b piper -l jafails at once with a message that says so.³
qwen-nativeaccepts en, fr, es, de, it, pt, zh, ja, ko and ru. It loads the Qwen3-TTS Base model, which has no preset voices: it is the backend for voice clones. Without a clone,-lhas no effect (the model infers the language from the text). The first use downloads the model and its tokenizer, about 2.5 GB.
kokoro is not in the release binaries: there, vox -b kokoro answers
Unknown backend: kokoro.
Time from launching vox to the first sound, for a sentence that lasts about
3 seconds. Measured in October 2026 on an Apple A18 Pro laptop (8 GB, low-power
mode) with the release Metal build:
| Backend | Without the daemon | With the daemon |
|---|---|---|
pocket (English) |
0.18 to 0.55 s | not measured |
piper (French) |
about 0.75 s | 0.26 to 0.32 s |
A long English text (15 s of audio) starts after 0.81 s: pocket plays while
it is still generating. piper and pocket send their samples straight to the
audio device, which is opened while the model loads. The first use of a
backend also downloads its model.
vox daemon start keeps the models loaded between calls (vox daemon status,
vox daemon stop). While it runs, vox "..." calls that use pocket, piper,
qwen-native or kokoro go through it, except calls with -o. It is not
started automatically, and it stops by itself after 300 seconds without a
request (vox daemon start --idle-timeout <seconds>; 0 means no timeout).
The MCP server does not use the daemon: it keeps the models loaded in its own
process.
VOX_TIMINGS=1 prints the time of each step of an utterance on stderr.
Time of a whole vox -o file.wav "..." command: process start, model loaded
from disk, synthesis, file written. Nothing is played, so the figures do not
depend on an audio device. Measured in October 2026 with the v0.17.0 release
binaries, each command in a new process, models already downloaded. Each figure
is the median of 5 runs (3 for qwen-native) after one run that is not counted.
| macOS, Metal build | Linux, CPU build | Linux, CUDA build | Windows | |
|---|---|---|---|---|
pocket, English, one sentence (about 3 s of audio) |
0.8 s | 3.7 s | 3.8 s | 2.5 s |
pocket, English, four sentences (about 14 s) |
3.3 s | 13.5 s | 13.3 s | 9.5 s |
piper, French, one sentence (about 2.6 s) |
0.6 s | 1.2 s | 1.2 s | 1.6 s |
piper, French, four sentences (about 13 s) |
1.6 s | 2.2 s | 2.2 s | 2.5 s |
qwen-native, cloned voice, one sentence (about 2.4 s) |
31 s | 56 s | 8.7 s | 46 s |
Whisper base (vox hear --file), 13.8 s of audio |
1.7 s | 15.1 s | 1.1 s | 5.6 s |
- macOS: Apple A18 Pro laptop, 8 GB, macOS 27.
- Linux: Ubuntu 26.04 under WSL2 on an Intel Core i7-8700 (6 cores) with an RTX 4070 Ti SUPER (16 GB).
- Windows: Windows 11 on that same machine. There is one Windows build and it uses the CPU. The commands were started from WSL, which adds about 50 ms to each.
What the figures show:
pocketandpipertake the same time in the CPU and the CUDA build: they run on the CPU in both.- The CUDA build is what makes
qwen-nativeand Whisper fast on Linux: 6 times and 14 times faster here than the CPU build. pocketgenerates about as fast as it plays on this Linux machine, and four times faster on the Mac. It plays while it generates, so the sound starts long before the times above; see Time to first sound.- On the same processor, the CPU build is slower under WSL2 than the Windows
build for
pocket,qwen-nativeand Whisper, and faster forpiper. We have not looked into why. - The first
qwen-nativecall after a download took longer: 83 s on Linux with the CPU build, 98 s on Windows, 9 s with the CUDA build.
The cloned voice used a 13.8 s reference recording at 24 kHz with its transcription. Linux on ARM64 has not been measured.
The minimum and maximum of each figure, the first-call times, the texts used and what was not measured are in docs/BENCHMARKS.md.
say and kokoro have not been measured again. The figures below are older:
end-to-end times for one sentence of about 50 characters (model load, synthesis
and playback) on an M2 Pro using the CPU.
| Backend | Time |
|---|---|
say |
3 s |
kokoro |
under 1 s |
# Quick install (macOS Apple Silicon, Linux x86_64 and ARM64, WSL2)
curl -fsSL https://github.com/_p/raw/rtk-ai/vox/main/install.sh | sh
# Custom install dir (no sudo needed)
curl -fsSL https://github.com/_p/raw/rtk-ai/vox/main/install.sh | VOX_INSTALL_DIR=~/.local/bin sh
# Homebrew (macOS Apple Silicon, Linux)
brew install rtk-ai/tap/voxThe installer defaults to /usr/local/bin (with sudo if needed). When sudo is
not usable (CI, agents, no TTY), it falls back to ~/.local/bin automatically.
Most users need nothing special. The two default voices, pocket (English) and
piper (other languages), run on the CPU in every build. A GPU build speeds up
only two things: Whisper transcription (vox hear) and the qwen-native
backend (voice cloning). When the GPU device cannot be opened at run time,
both fall back to the CPU.
| Platform | Build | How to get it |
|---|---|---|
| macOS, Apple Silicon | Metal (GPU) | Installer or Homebrew |
| macOS, Intel | none | Not supported, see below |
| Linux x86_64 | CPU, or CUDA with an NVIDIA card | Installer (it asks, see below). Homebrew, .deb and .rpm are CPU |
| Linux ARM64 | CPU | Installer, Homebrew, .deb or .rpm |
| Windows x86_64 | CPU | The .zip from GitHub Releases. With an NVIDIA card: WSL2 and the Linux installer |
To check which build is installed, run vox config show. Its acceleration:
line reads Metal (GPU): used by Whisper and qwen-native,
CUDA (NVIDIA GPU): used by Whisper and qwen-native or CPU only. It reports
how the binary was built, not which device is in use.
There is no Intel Mac build (x86_64-apple-darwin): the ONNX runtime that
piper depends on ships no prebuilt binary for Intel Macs. The last release
with an Intel Mac binary is v0.10.0. The installer stops there with a message
saying so, before downloading anything.
There is no Windows CUDA build either. With an NVIDIA card on Windows, run the Linux installer inside WSL2.
The Linux binaries need glibc 2.39 or newer: Ubuntu 24.04, Debian 13, Fedora 40
and later. On Ubuntu 22.04 or Debian 12 the installer stops with a message
before downloading anything. Building from source does not help there: the
ONNX runtime used by the piper voices needs glibc 2.38 itself. They also need
ALSA (libasound2t64, or alsa-lib on Fedora) and OpenSSL 3 at run time.
On Linux x86_64, when nvidia-smi runs successfully, the installer asks
whether to install the CUDA build. The question says what it changes (faster
voice cloning with qwen-native and faster transcription with Whisper; the
default voices run on the CPU either way) and what it needs. The default
answer is No.
VOX_GPU answers the question in advance:
VOX_GPU |
Effect on Linux x86_64 with an NVIDIA card |
|---|---|
auto (default) |
Asks when a terminal is reachable. Otherwise installs the CPU build and prints how to get the CUDA one |
cuda |
Installs the CUDA build without asking |
cpu |
Installs the CPU build without asking |
curl -fsSL https://github.com/_p/raw/rtk-ai/vox/main/install.sh | VOX_GPU=cuda shAny other value is an error. VOX_GPU=cuda on a platform it does not apply to
(macOS, Linux ARM64) prints a warning and installs the normal build for that
platform. macOS always gets the Metal build, the only macOS build, so
VOX_GPU=cpu only prints a warning there too. On Linux x86_64 without a
working nvidia-smi, VOX_GPU=cuda prints a warning and still tries the CUDA
build.
The CUDA build needs, on the machine that runs it:
- the NVIDIA driver (it provides
libcuda.so.1); - the CUDA 12 runtime libraries cuBLAS (
libcublas.so.12,libcublasLt.so.12) and cuRAND (libcurand.so.10). This is whatlddlists on a real CUDA build.
These libraries are linked dynamically: without them the CUDA binary does not
start at all. So the installer runs the downloaded binary (vox --version)
before installing it. If it does not start, the installer lists the missing
libraries and installs the CPU build instead. It also installs the CPU build,
with a warning, when the release has no CUDA binary or its download fails. Its
last lines say which
build was installed (CPU, CUDA or Metal) and, for a CPU build on a machine with
an NVIDIA card, how to switch.
The CUDA build is compiled for CUDA 12.6 and GPU compute capability 8.0, so it needs an RTX 30 series card or newer (or a data-center card of the same generations). Older cards are not supported: the GPU kernels do not compile below 8.0.
What has been checked on real hardware, an RTX 4070 Ti SUPER under WSL2 with CUDA 12.4: a CUDA build compiled there starts, the installer installs it and falls back to the CPU build in a container without the CUDA libraries, and Whisper transcribes 24.6 s of audio in about 1.4 s against about 26 s for the CPU build on the same machine, with the same text. The binary the release workflow produces has not been run yet: CI compiles it on runners that have no GPU.
| Platform | Binary | Build |
|---|---|---|
| macOS (Apple Silicon) | vox-aarch64-apple-darwin.tar.gz |
Metal |
| Linux x86_64 | vox-x86_64-unknown-linux-gnu.tar.gz |
CPU |
| Linux x86_64, NVIDIA | vox-x86_64-unknown-linux-gnu-cuda.tar.gz |
CUDA 12. Built by the release workflow, not in v0.16.0 or earlier. Installed with VOX_GPU=cuda |
| Linux ARM64 | vox-aarch64-unknown-linux-gnu.tar.gz |
CPU |
| Linux (Debian/Ubuntu) | vox-{x86_64,aarch64}-unknown-linux-gnu.deb |
CPU |
| Linux (Fedora/RHEL) | vox-{x86_64,aarch64}-unknown-linux-gnu.rpm |
CPU |
| Windows x86_64 | vox-x86_64-pc-windows-msvc.zip |
CPU |
Download from GitHub Releases. The CUDA job of the release workflow is allowed to fail, so a release can go out without the CUDA binary: check the asset list of the release you install.
cargo install --git https://github.com/rtk-ai/vox # CPU only
cargo install --git https://github.com/rtk-ai/vox --features metal # macOS Apple Silicon (Metal)
cargo install --git https://github.com/rtk-ai/vox --features cuda # Linux x86_64, NVIDIA (CUDA)
# From a clone of this repository
cargo install --path . # same --features flagsDo not run cargo install vox: on crates.io that name belongs to an unrelated
crate (github.com/bearcove/vox).
Build prerequisites:
- Linux:
sudo apt install build-essential cmake pkg-config clang libclang-dev libssl-dev libasound2-dev. The piper voices build espeak-ng with cmake and generate their bindings with libclang; model downloads link OpenSSL. --features cuda: the CUDA 12 toolkit withnvccon thePATH(CI uses 12.6.3; 12.4 is known to work). The GPU kernels are compiled for one compute capability, read fromnvidia-smion the build machine, and run on that generation and newer ones. SetCUDA_COMPUTE_CAPwhen building for another machine or wherenvidia-smiis not on thePATH(under WSL it lives in/usr/lib/wsl/lib). The lowest value that compiles is80(8.0, RTX 30 series), which CI and the release build use.
The default backend depends on the language, not on the platform:
| Language | Default backend | First use |
|---|---|---|
English, or no -l |
pocket |
Downloads the public weights (~226 MB) |
Japanese (-l ja) |
qwen-native |
Downloads Qwen3-TTS (about 2.5 GB). piper has no Japanese voice. Slower than the other two: 14 s for a short sentence, model load included, on an Apple A18 Pro laptop with the Metal build |
| Any other language | piper |
Downloads the voice of that language (~60 MB). The pocket checkpoint is English-only |
A stored backend preference (vox config set backend ...) replaces both
defaults.
vox "Hello, world." # Speak with the default backend (pocket)
vox -l fr "Bonjour" # French: piper
vox -b piper "Hello from piper." # Choose a backend
vox -b qwen-native "Hello from Qwen3." # Qwen3-TTS
vox --volume 2.0 "Louder!" # 2x volume (range: 0.0-5.0)
echo "Piped text" | vox # Read from stdin
vox -o note.wav "Saved, not spoken." # Write a WAV file instead of playing
vox --list-voices # Voices of the selected backend
vox setup # Interactive TUI configurationvox setup opens a terminal interface to choose the backend, voice, language,
style and volume, and to test the result:
┌ Backend ─────┐┌ Voice ────┐┌ Language ┐┌ Style ─────┐┌ Volume ┐┌ Config ──────────────────────┐
│> pocket ││> alba ││> en ││> (default) ││ 0.5 ││ Backend: pocket │
│ piper ││ marius ││ fr ││ calm ││ 0.75 ││ Voice: alba │
│ qwen-native ││ javert ││ es ││ energetic ││> 1.0 ││ Lang: en │
│ say ││ jean ││ de ││ warm ││ 1.25 ││ [T] Test [S] Save [Q] Quit │
└──────────────┘└───────────┘└──────────┘└────────────┘└────────┘└──────────────────────────────┘
The backend list holds the backends of the build (say on macOS only).
Up/Down or j/k move in a list, Tab, Left/Right or h/l change panel, T speaks a
test sentence, S saves, Q or Esc quits.
S saves the backend, the language, the voice and the style. The volume is used
for the test only: there is no volume preference, pass --volume on each call.
AI agents use the command-line flags instead: vox -l fr "text".
One command writes the vox MCP server into the configuration of the AI tools found on your machine. It knows 14 AI tools: Claude Code, Claude Desktop, Cursor, Windsurf, VS Code / Copilot, Zed, Codex, OpenCode, Gemini, Amazon Q, Cline, Roo Code, Kilo Code and Amp.
vox init # MCP server (default), for the installed tools among the 14
vox init -m cli # CLAUDE.md block + Stop hook, in the current directory
vox init -m skill # /speak slash command for Claude Code
vox init -m all # all of the above| Mode | What it writes | What the agent gets |
|---|---|---|
mcp |
A vox serve entry in the MCP configuration of each tool found on the machine, in your home directory. A tool counts as installed when its configuration file or its own directory exists; the others are reported as not installed, skipped and nothing is written for them |
14 tools: vox_speak, vox_hear, vox_list_voices, clones, preferences, statistics and sound packs |
cli |
A block in CLAUDE.md and a Stop hook in .claude/settings.json, both in the current directory |
An instruction to run vox "..." after a significant task. The hook says a short phrase ("Done.") when Claude Code stops |
skill |
~/.claude/commands/speak.md |
The /speak command |
Running vox init again is safe: what is already configured is left as it is.
After its report, vox init names the tools to restart, or says that nothing
changed.
The MCP server starts with vox serve (stdio); vox init writes that command
for you. With the MCP tools an agent can hold a voice conversation by itself:
vox_hear, think, vox_speak, with no API key.
The generated instructions and the Stop hook use your language. vox takes
--lang, then your vox config set lang preference, then the system locale.
With none of them it tells the agent to match the language you write in.
vox init -m cli --lang de # agent summaries in German, the hook says "Fertig."
vox init -m cli --lang es # Spanish, the hook says "Listo."--lang accepts en, fr, es, de, it, pt, zh, ja, ko, ru, ar and nl. The hook
runs vox -l <lang> "...", which uses piper for every language but English.
vox ships a Claude Code plugin that draws the spectrum of the voice above the prompt while vox speaks:
▂ ▃ ▃ ▄ ▃ vox · speaking
▃ ▅ ▄ ▂ ▄ ▃ █ █ ▅ ▁ ▂ ▅ █ █ █ ▆ ▃ ▁ Done. The tests pass.
The bars are the spectrum of the audio, not an animation. While it plays, vox
writes the spectrum of the sound to now-playing.json in its config directory
and removes the file afterwards; the plugin reads that file. It works with the
MCP tools, with vox ... run from the shell, and with the Stop hook. At the
default volume, the say backend plays outside vox and writes nothing, so the
plugin shows no spectrum for it.
In a Claude Code session (2.1.287 or later):
/plugin marketplace add rtk-ai/vox
/plugin install vox@vox
Then pick your colors, kept from one session to the next:
/vox-wave # preview, no audio needed
/vox-wave color ocean # sunset, ocean, forest, fire, violet, rainbow, mono
/vox-wave color #00ff00 #0000ff # your own gradient, one to three stops
vox init tells the agent that the plugin exists (in the CLAUDE.md block and
in the MCP instructions), so you can also ask Claude how to get the visualizer.
Details in plugins/vox.
vox clone add patrick --audio ~/voice.wav --text "Transcription"
vox clone record myvoice --duration 10
vox -v patrick "This speaks with your voice."
vox clone list
vox clone remove patrickCloning works with qwen-native, with no setup, and with pocket, which needs
HF_TOKEN (see Backends). A 3-second reference clip is enough.
Without -b, vox -v <clone> uses pocket when pocket is the selected
backend and HF_TOKEN is set, and qwen-native in every other case. A -b on
the command line is respected: -b pocket without HF_TOKEN stops with an
error that asks for the token, and a backend that cannot clone (piper or
say, for example) prints a note and ignores the clone.
clone add reads a WAV, MP3, FLAC or Ogg file, converts it to WAV and keeps
that copy in the clones/ folder of the config directory: the original file
can be moved or deleted afterwards. A .m4a file is refused. So is a name that
is already taken, whatever its case: remove that clone first. clone record
saves its recording, a WAV, in the same folder.
vox config show
vox config set backend qwen-native
vox config set lang fr
vox config set voice alba
vox config set stt_model openai/whisper-base
vox config resetThe keys are backend, voice, lang, rate, gender, style, model,
stt_model and pack. A flag on the command line wins over a preference.
langaccepts en, fr, es, de, it, pt, zh, ja, ko, ru, ar and nl.rate(words per minute) applies tosayonly,modeltoqwen-nativeonly.genderandstyleare accepted and stored, but no backend of this version uses them.
vox config show ends with an acceleration: line that says how the binary
was built (Metal, CUDA or CPU).
vox pack list # Installed packs, and a few names to install
vox pack install peon # Download a pack
vox pack set peon # Make it the active pack
vox pack play greeting # Play a random sound of a category
vox pack remove peon # Delete itvox pack install takes the name of a pack of the peon-ping registry (browse
them at https://openpeon.com/packs); vox pack list suggests a few of them.
Packs are published in the CESP format (openpeon.json), which vox converts
when it installs them. The categories are greeting, acknowledge,
complete, error, permission, resource_limit and annoyed. A pack can
carry other categories, which keep their CESP name (session.end,
task.progress).
vox -o note.wav "Text to save" # default backend
vox -l fr -o note.wav "Texte a enregistrer" # piper
vox -b qwen-native -v myvoice -o note.wav "..." # with a voice clone
ffmpeg -i note.wav -c:a libopus -b:a 32k note.ogg # convert it afterwards, here to Opus-o writes a WAV file instead of playing, on every backend and every platform.
A call with -o does not go through the daemon.
--output exists on the command line only. The MCP vox_speak tool has no
such parameter, so an agent cannot use it to write to a path of its choice.
Whisper on candle, 99 languages. The model is downloaded from Hugging Face on
first use. It runs on the GPU in a Metal or CUDA build. The MCP server keeps it
loaded between two vox_hear calls; vox hear from the shell loads it at each
call.
vox hear # Listen, auto-detect language, print text
vox hear -l fr -t 60 -s 3.0 # French, max 60s, stop after 3s of silence
vox hear -m openai/whisper-large-v3-turbo # Larger model (GPU + 16 GB RAM recommended)
vox hear -f recording.wav # Transcribe a WAV file instead of the micA short beep marks the start of the recording and a lower one its end
(VOX_CUES=0 turns them off).
Model size is a tradeoff. Time and RAM were measured warm on 11.85 s of French
speech, on an Apple M2 (tiny, base, small). The size on disk is that of
the files in the Hugging Face cache:
| Model | Time | RAM | On disk | Quality |
|---|---|---|---|---|
openai/whisper-tiny |
0.93 s | 350 MB | 154 MB | roughest |
openai/whisper-base |
1.52 s | 631 MB | 293 MB | the default |
openai/whisper-small |
5.07 s | 1.98 GB | 970 MB | one fewer mistake in 30 words |
openai/whisper-large-v3-turbo |
not measured | ~3.5 GB | not measured | not measured, GPU recommended |
base is the default because, in that test, it was 3x faster and 3x lighter
than small for one more mistake on a 30-word sentence.
What a GPU build changes: on a 12-core x86_64 Linux machine (WSL2) with an RTX
4070 Ti SUPER, base transcribes 24.6 s of audio in about 26 s with the CPU
build and in about 1.4 s with the CUDA build, with the same text.
To use another model, in order of precedence:
vox hear -m openai/whisper-small # 1. this call only
export VOX_STT_MODEL=openai/whisper-small # 2. this shell
vox config set stt_model openai/whisper-small # 3. stored preference
# 4. [whisper] model_id in models.toml, in the config directory| Env var | Description |
|---|---|
VOX_STT_MODEL |
Whisper repo. Wins over the stored preference and models.toml (default openai/whisper-base, ~630 MB RAM; openai/whisper-large-v3-turbo ~3.5 GB) |
VOX_VAD_THRESHOLD |
Minimum RMS speech threshold, between 0 and 1 (default 0.0125; adapts to ambient noise) |
VOX_VAD_DEBUG |
Set to 1 to print RMS levels and the chosen threshold |
VOX_CUES |
Set to 0 to turn off the beeps at the start and the end of a recording |
No external tool is needed: the microphone is read with cpal.
There are two ways to talk with an assistant.
Through the MCP server, on every platform and with no API key: the agent calls
vox_hear, thinks, and answers with vox_speak. Ask it to "chat" or "talk".
With vox chat, on macOS only. vox calls the Claude API itself, so it needs
ANTHROPIC_API_KEY:
export ANTHROPIC_API_KEY=sk-...
vox chat -l fr # Talk with Claudevox chat records until you press Enter, transcribes with Whisper, sends the
text to the Claude API and speaks the answer. Its prompts and its system prompt
follow -l, or the stored lang preference: English by default, French for
fr. For another language they stay in English and Claude is asked to answer
in that language. The model is claude-haiku-4-5; VOX_CHAT_MODEL names
another one.
What you say and what vox reads aloud stay on your machine, except with
vox chat, which sends the transcript to the Claude API. Models and voices are
downloaded from Hugging Face on first use; vox pack install downloads a sound
pack from GitHub.
~/.config/vox/ # Linux; ~/Library/Application Support/vox/ on macOS
vox.db # SQLite: preferences, voice clones, usage log
clones/ # Reference audio of the voice clones (WAV)
packs/ # Installed sound packs
piper/ # Piper voices
pocket/ # Configuration of the pocket model
models.toml # Optional: your own model ids
now-playing.json # Only while vox plays: the spectrum for the visualizer
daemon.pid # Only while the daemon runs
daemon.log # Output of the daemon, emptied at each start
The Whisper, pocket and Qwen3-TTS weights are in the Hugging Face cache
(~/.cache/huggingface/hub), not in this directory.
| Env var | Description |
|---|---|
VOX_CONFIG_DIR |
Override config directory |
VOX_DB_PATH |
Override database path |
VOX_TIMINGS |
Set to 1 to print, on stderr, when each step of an utterance finished (model load, synthesis, device open, first sample) |
VOX_DAEMON_PORT |
Port of the daemon on 127.0.0.1 (default 19876) |
HF_TOKEN |
Hugging Face token, needed only for voice cloning with pocket |
ANTHROPIC_API_KEY |
Needed only by vox chat |
VOX_CHAT_MODEL |
Claude model used by vox chat (default claude-haiku-4-5) |
These documents are written in French.
| Document | Description |
|---|---|
| Architecture | Technical architecture, backends, DB schema, MCP protocol, security |
| Features | All commands and features documented |
| Guide | Installation, quick start, troubleshooting |
| Benchmarks | Generation and transcription times on macOS, Linux and Windows, with the method |
