A Brazilian Portuguese voice: Qwen3-TTS#
This recipe picks a text-to-speech model for Brazilian Portuguese and serves it beside Whisper on one consumer NVIDIA GPU, for spoken conversations in the app or from your own programs. The pick is Qwen3-TTS 0.6B (Alibaba's Qwen team, Apache 2.0), which speaks in the voice of a Brazilian recording that every deployment ships with.
Which voice model for Portuguese#
| Model | Engine | Runs on | Brazilian Portuguese | First sound | GPU memory |
|---|---|---|---|---|---|
Kokoro 82M (pf_dora, pm_alex…) |
speaches | CPUs | Native voices, but flat and robotic | Fastest: faster than real time on a few cores | none |
Piper (pt-br-faber) |
speaches | CPUs | Native, clearly synthetic | Fast | none |
| Chatterbox Multilingual 500M | Chatterbox | NVIDIA GPU | Speaks Portuguese in the voice it imitates; the general multilingual model leaves artefacts and an accent | After its first sentence (WAV) or the whole text (MP3) | 3.7 GB (bf16) |
| Qwen3-TTS 0.6B | vLLM-Omni | NVIDIA GPU | Clones a Brazilian speaker in context: accent and rhythm, not only timbre | WAV streams from the first audio frame | about 5.8 GB |
| Qwen3-TTS 1.7B | vLLM-Omni | NVIDIA GPU | The same, from a larger model | The same | about 7.7 GB |
Why Qwen3-TTS for Portuguese:
- A native accent. Qwen3-TTS speaks ten languages, Portuguese among
them, and its Base models speak in the voice of a reference recording
given with its transcript. Every deployment ships a Brazilian one,
pt-BR-1, so the voice, the accent and the rhythm are Brazilian. Chatterbox's general model and its English voices give Portuguese an English accent. - Streaming. Asked for WAV or PCM, it sends audio as it is decoded, the first piece after one audio frame; it does not wait for a sentence.
- Licence. Apache 2.0 for the model and for vLLM-Omni, its server.
What it costs:
- GPU memory: about 5.8 GB for 0.6B — it runs two engines (one writes the speech codes, one turns them into sound), each in its own process. Beside Whisper on an 8 GB card, take a small Whisper (below).
- Start time: about one to two and a half minutes for a replica to
start (two engines load, then compile and capture their GPU kernels);
the first start also downloads 2.5 GB of weights and pulls a 9.4 GB
image. Keep it running (
min: 1) rather than scaling it to zero. - NVIDIA only: vLLM-Omni's CUDA image, with an NVIDIA driver of the 580 series or newer (CUDA 13).
Measured figures
The GPU memory above is computed from Qwen3-TTS's weights and a published measurement of vLLM-Omni on a 16 GB GPU, held to one request at a time. It has not been measured on an 8 GB RTX 4060 yet; check the machine's page after the first start.
Before you begin#
- A machine with an NVIDIA GPU offering about 6.5 GB for both models (an 8 GB card whose desktop holds 1 GB offers 7), an NVIDIA driver of the 580 series or newer, 12 GB of RAM free, x86-64, and a data location. A Windows PC under WSL2 works.
- A chat deployment anywhere in the workspace (see Serve an open chat model on one GPU machine).
What each takes of the GPU#
| Model | Variant | Weights | GPU memory taken | RAM |
|---|---|---|---|---|
whisper:base |
q8_0 |
82 MB | 0.7 GB | 2 GiB |
whisper:small |
q8_0 |
264 MB | 0.9 GB | 2 GiB |
whisper:large-v3-turbo |
q5_0 |
574 MB | 1.2 GB | 2 GiB |
qwen3-tts:0.6b |
bf16 |
2.5 GB | 5.8 GB | 9 GiB |
qwen3-tts:1.7b |
bf16 |
4.5 GB | 7.7 GB | 11 GiB |
Qwen3-TTS 0.6B with Whisper small takes 6.7 GB; with large-v3-turbo
q5_0, 7.0 GB — the most an 8 GB card with a desktop offers. The 1.7B
model needs a larger GPU (12 GB or more) beside Whisper.
On Linux each model is held to its GPU memory. On Windows (WSL2) nothing holds them to it — each engine's own sizing does — and the machine's page says fractions not enforced here. Qwen3-TTS's speech engine keeps a fixed cache for one request; it does not grow with the GPU.
1. Deploy them#
- Eos → Models → Library, Category: transcription, open
Whisper, click Small · q8_0, press Deploy. Runs on:
A GPU (whisper.cpp: CUDA, or Vulkan on AMD). Name:
listen. Clear Scale to zero when idle, press Deploy. - Category: speech, open Qwen3-TTS, click 0.6B · bf16,
press Deploy. Runs on is An NVIDIA GPU (text to speech:
vLLM-Omni). Name:
voice. Clear Scale to zero when idle, press Deploy.
$ curl -fsS -X POST "$ASTRALYX_API/eos/deployments" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"metadata": {"name": "listen"}, "spec": {"model": "whisper-small-q8-0", "gpu_memory_gb": 0.9, "replicas": {"min": 1, "max": 1}}}' > /dev/null
$ curl -fsS -X POST "$ASTRALYX_API/eos/deployments" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"metadata": {"name": "voice"}, "spec": {"model": "qwen3-tts-0-6b-bf16", "replicas": {"min": 1, "max": 1}}}' > /dev/null
Qwen3-TTS takes the GPU memory it needs by default; a Whisper
deployment runs on CPUs unless it asks for a GPU (gpu_memory_gb,
gpus or gpu_vendor).
The replica's log shows voice pt-BR-1: 7.2 s at 24000 Hz as it starts,
then vLLM-Omni loading its two stages. It takes traffic once
GET /health answers.
2. Talk#
Open the app, choose the chat model, and press the
wave button: what you say goes to listen, the answer is spoken by
voice in pt-BR-1 — see Talk to your models.
From a program, they are OpenAI's audio endpoints. vLLM-Omni takes two
fields of its own: language (Portuguese, English, Spanish,
French, German, Italian, Russian, Chinese, Japanese,
Korean; Auto by default — name it: a short sentence can be taken for
another language) and stream_format: "audio" (WAV or PCM sent as it is
decoded, at speed 1 only).
import os
from openai import OpenAI
client = OpenAI(base_url=os.environ["EOS_URL"], api_key=os.environ["ASTRAEUS_API_KEY"])
with client.audio.speech.with_streaming_response.create(
model="voice", voice="pt-BR-1", input="Olá! O treinamento terminou sem erros.",
response_format="wav", extra_body={"language": "Portuguese", "stream_format": "audio"},
) as r:
r.stream_to_file("aviso.wav")
with open("aviso.wav", "rb") as f:
print(client.audio.transcriptions.create(model="listen", file=f, language="pt").text)
Through the Playground and the app, the language is the one the request
names (pt-BR), else the voice's, and WAV is streamed by itself.
How quickly it speaks#
- WAV or PCM, streamed: the first sound after the text is read and its first audio frame (an eighth of a second of speech) is made and decoded.
- MP3, FLAC: once the whole text is made. The app asks for MP3 a sentence at a time, its first piece short, so the first sound comes after one short piece is made.
- A replica answers one request at a time; more replicas
(
replicas.max) answer more, each with its own GPU memory.
To see it on your machine, time a request:
$ curl -sS -o /dev/null -w 'first byte %{time_starttransfer}s, all %{time_total}s\n' "$EOS_URL/audio/speech" \
-H "Authorization: Bearer $ASTRAEUS_API_KEY" -H 'content-type: application/json' \
-d '{"model": "voice", "voice": "pt-BR-1", "input": "Bom dia! Como posso ajudar?", "language": "Portuguese", "response_format": "wav", "stream_format": "audio"}'
The voice#
pt-BR-1 is 7.2 s of a public-domain recording — chapter one of Machado
de Assis' Dom Casmurro (1899), read for
LibriVox by the
volunteer Leni (Internet Archive item dom_casmurro_2102_librivox,
Public Domain Mark 1.0) — and what is said in it, "E eles, por graça,
chamam-me assim, alguns em bilhetes: Dom Casmurro, domingo vou jantar com
você." Qwen3-TTS reads both before each request and continues in that
voice.
A deployment has no other voice: Qwen3-TTS's Base models have no
speakers of their own, and the server's voice upload is not reachable
through Eos. A request may carry its own reference recording (ref_audio
as a data: URL, with ref_text); a URL to fetch is refused.
Clone only voices you have the right to use
A recording you send as ref_audio is cloned. Send your own voice,
or a voice whose owner gave you permission to clone it for this use.
Engine arguments#
engine_args may add --enforce-eager (no captured GPU kernels: a
quicker start, slower speech), --seed, --disable-log-stats,
--disable-uvicorn-access-log and --uvicorn-log-level. The rest is set
by the deployment.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
| The replica waits, saying how much of the GPU is taken and that it needs 5.8 | Whisper or another model holds too much of the GPU. | Take a smaller Whisper (small, base), or move it to CPUs. |
| The replica restarts during its start, out of GPU memory | Something outside Eos (a game, the desktop) holds more of the GPU than when it was placed. | Free the GPU, or give it more: gpu_memory_gb. |
400 Invalid language 'Pt-br' |
A program sent a language code to the endpoint. | Send "language": "Portuguese". |
400 streaming requires … |
stream_format: "audio" with MP3, or with another speed. |
Ask for WAV or PCM at speed 1, or leave stream_format out. |