Voice on one GPU: Whisper and Chatterbox#
A spoken conversation needs ears and a voice that answer in a fraction of a second. On one consumer GPU — an 8 GB RTX 4060 is enough — Whisper (speech to text, served by whisper.cpp) and Chatterbox Multilingual (text to speech, from Resemble AI) run side by side, each taking the GPU memory it needs: they share the GPU. The chat model can run on another machine. The app's voice mode uses all three.
What you need:
- A machine with an NVIDIA GPU offering about 7 GB (an 8 GB card whose desktop holds 1 GB is fine), an NVIDIA driver of the 580 series or newer (whisper.cpp's CUDA build is CUDA 13; Chatterbox's is CUDA 12.8), 8 GB of RAM free, x86-64, and a data location. A Windows PC under WSL2 works.
- A chat deployment anywhere in the workspace (see Serve an open chat model on one GPU machine).
What each takes of the GPU#
| Model | Variant | Weights | GPU memory taken | RAM |
|---|---|---|---|---|
whisper:base |
q8_0 |
82 MB | 0.7 GB | 2 GiB |
whisper:small |
q8_0 |
264 MB | 0.9 GB | 2 GiB |
whisper:large-v3-turbo |
q5_0 |
574 MB | 1.2 GB | 2 GiB |
whisper:large-v3-turbo |
f16 |
1.6 GB | 2.2 GB | 3 GiB |
chatterbox:500m |
fp32 |
3.2 GB | 4.2 GB | 6 GiB |
The GPU memory is the weights, the engine's working buffers and the CUDA
context of its process (about 0.4 GB each). Large-v3-turbo f16 and
Chatterbox together take 6.4 GB: they fit a GPU offering 7. On a GPU
offering less, take turbo q5_0 (5.4 GB together).
The GPU does not hold an engine to its share: each sizes itself (whisper.cpp allocates its model and buffers; Chatterbox its model in full precision). Something else on the GPU — a game, another engine — can leave them short.
Chatterbox: what to know#
- Licence: MIT (the model,
ResembleAI/chatterbox, and the server that serves it, Chatterbox TTS Server2.0.0). - Languages: 23, Portuguese among them. A deployment speaks one
language, set with the engine argument
--language(defaulten); deploy one per language you need. - Voices: a voice is a short reference recording Chatterbox imitates.
The server comes with 28 English speakers —
Olivia.wav,Gabriel.wav,Emily.wav… — and speaks the deployment's language in their voice, with an accent from the recording. Voices of your own (cloning from your recording) are not offered yet.
1. Deploy them#
- Eos → Models → Library, Category: transcription, open
Whisper, click Large v3 Turbo · f16, press Deploy.
Runs on: A GPU (whisper.cpp: CUDA, or Vulkan on AMD). Share
the GPU with other models is on: the line beside it says takes
2.2 GB of each GPU. Name:
listen. Clear Scale to zero when idle, press Deploy. - Category: speech, open Chatterbox Multilingual, click
500M · fp32, press Deploy. Name:
voice. Engine arguments:--language pt. Clear Scale to zero when idle, press Deploy.
$ curl -fsS -X POST "$ASTRALYX_API/eos/deployments" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"metadata": {"name": "listen"}, "spec": {"model": "whisper-large-v3-turbo-f16", "gpu_share": true, "replicas": {"min": 1, "max": 1}}}' > /dev/null
$ curl -fsS -X POST "$ASTRALYX_API/eos/deployments" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"metadata": {"name": "voice"}, "spec": {"model": "chatterbox-500m-fp32", "engine_args": ["--language", "pt"], "replicas": {"min": 1, "max": 1}}}' > /dev/null
A Whisper deployment runs on CPUs unless it asks for a GPU:
gpu_share, gpus or gpu_vendor. Chatterbox shares its GPU by
default.
The first start downloads the weights (1.6 GB and 3.2 GB) and pulls the engines' images (about 2 GB and 5.4 GB). The machine's page then shows its GPU shared: listen 2.2 GB, voice 4.2 GB — 6.4 GB taken.
A third model that does not fit waits, and says why: GPU 0 of majin: 6.4 of 7 GB taken by shared work; it needs 2.2.
2. Talk#
Open the app, choose the chat model, and press the
wave button: what you say goes to listen, the answer is spoken by
voice — see Talk to your models. Settings →
Voice chooses the voice.
From a program, they are OpenAI's audio endpoints:
import os
from openai import OpenAI
client = OpenAI(base_url=os.environ["EOS_URL"], api_key=os.environ["ASTRAEUS_API_KEY"])
with client.audio.speech.with_streaming_response.create(
model="voice", voice="Olivia.wav", input="Olá! O treinamento terminou sem erros.", response_format="mp3"
) as r:
r.stream_to_file("aviso.mp3")
with open("aviso.mp3", "rb") as f:
print(client.audio.transcriptions.create(model="listen", file=f, language="pt").text)
Chatterbox answers wav, mp3 and opus, and takes speed.