Skip to content

Voice on one GPU: Whisper and Chatterbox#

A spoken conversation needs ears and a voice that answer in a fraction of a second. On one consumer GPU — an 8 GB RTX 4060 is enough — Whisper (speech to text, served by whisper.cpp) and Chatterbox Multilingual (text to speech, from Resemble AI) run side by side, each taking the GPU memory it needs: they share the GPU. The chat model can run on another machine. The app's voice mode uses all three.

What you need:

  • A machine with an NVIDIA GPU offering about 7 GB (an 8 GB card whose desktop holds 1 GB is fine), an NVIDIA driver of the 580 series or newer (whisper.cpp's CUDA build is CUDA 13; Chatterbox's is CUDA 12.8), 8 GB of RAM free, x86-64, and a data location. A Windows PC under WSL2 works.
  • A chat deployment anywhere in the workspace (see Serve an open chat model on one GPU machine).

What each takes of the GPU#

Model Variant Weights GPU memory taken RAM
whisper:base q8_0 82 MB 0.7 GB 2 GiB
whisper:small q8_0 264 MB 0.9 GB 2 GiB
whisper:large-v3-turbo q5_0 574 MB 1.2 GB 2 GiB
whisper:large-v3-turbo f16 1.6 GB 2.2 GB 3 GiB
chatterbox:500m fp32 3.2 GB 4.2 GB 6 GiB

The GPU memory is the weights, the engine's working buffers and the CUDA context of its process (about 0.4 GB each). Large-v3-turbo f16 and Chatterbox together take 6.4 GB: they fit a GPU offering 7. On a GPU offering less, take turbo q5_0 (5.4 GB together).

The GPU does not hold an engine to its share: each sizes itself (whisper.cpp allocates its model and buffers; Chatterbox its model in full precision). Something else on the GPU — a game, another engine — can leave them short.

Chatterbox: what to know#

  • Licence: MIT (the model, ResembleAI/chatterbox, and the server that serves it, Chatterbox TTS Server 2.0.0).
  • Languages: 23, Portuguese among them. A deployment speaks one language, set with the engine argument --language (default en); deploy one per language you need.
  • Voices: a voice is a short reference recording Chatterbox imitates. The server comes with 28 English speakers — Olivia.wav, Gabriel.wav, Emily.wav… — and speaks the deployment's language in their voice, with an accent from the recording. Voices of your own (cloning from your recording) are not offered yet.

1. Deploy them#

  1. Eos → Models → Library, Category: transcription, open Whisper, click Large v3 Turbo · f16, press Deploy. Runs on: A GPU (whisper.cpp: CUDA, or Vulkan on AMD). Share the GPU with other models is on: the line beside it says takes 2.2 GB of each GPU. Name: listen. Clear Scale to zero when idle, press Deploy.
  2. Category: speech, open Chatterbox Multilingual, click 500M · fp32, press Deploy. Name: voice. Engine arguments: --language pt. Clear Scale to zero when idle, press Deploy.
$ astra eos add whisper:large-v3-turbo-f16
$ astra eos deploy whisper-large-v3-turbo-f16 --name listen --min 1 --share-gpu
$ astra eos add chatterbox:500m
$ astra eos deploy chatterbox-500m-fp32 --name voice --min 1 --engine-arg=--language --engine-arg=pt
$ curl -fsS -X POST "$ASTRALYX_API/eos/deployments" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"metadata": {"name": "listen"}, "spec": {"model": "whisper-large-v3-turbo-f16", "gpu_share": true, "replicas": {"min": 1, "max": 1}}}' > /dev/null
$ curl -fsS -X POST "$ASTRALYX_API/eos/deployments" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"metadata": {"name": "voice"}, "spec": {"model": "chatterbox-500m-fp32", "engine_args": ["--language", "pt"], "replicas": {"min": 1, "max": 1}}}' > /dev/null

A Whisper deployment runs on CPUs unless it asks for a GPU: gpu_share, gpus or gpu_vendor. Chatterbox shares its GPU by default.

The first start downloads the weights (1.6 GB and 3.2 GB) and pulls the engines' images (about 2 GB and 5.4 GB). The machine's page then shows its GPU shared: listen 2.2 GB, voice 4.2 GB — 6.4 GB taken.

A third model that does not fit waits, and says why: GPU 0 of majin: 6.4 of 7 GB taken by shared work; it needs 2.2.

2. Talk#

Open the app, choose the chat model, and press the wave button: what you say goes to listen, the answer is spoken by voice — see Talk to your models. Settings → Voice chooses the voice.

From a program, they are OpenAI's audio endpoints:

voice.py
import os
from openai import OpenAI

client = OpenAI(base_url=os.environ["EOS_URL"], api_key=os.environ["ASTRAEUS_API_KEY"])

with client.audio.speech.with_streaming_response.create(
    model="voice", voice="Olivia.wav", input="Olá! O treinamento terminou sem erros.", response_format="mp3"
) as r:
    r.stream_to_file("aviso.mp3")

with open("aviso.mp3", "rb") as f:
    print(client.audio.transcriptions.create(model="listen", file=f, language="pt").text)

Chatterbox answers wav, mp3 and opus, and takes speed.