Skip to content

A Brazilian Portuguese voice: Qwen3-TTS#

This recipe picks a text-to-speech model for Brazilian Portuguese and serves it beside Whisper on one consumer NVIDIA GPU, for spoken conversations in the app or from your own programs. The pick is Qwen3-TTS 0.6B (Alibaba's Qwen team, Apache 2.0), which speaks in the voice of a Brazilian recording that every deployment ships with.

Which voice model for Portuguese#

Model Engine Runs on Brazilian Portuguese First sound GPU memory
Kokoro 82M (pf_dora, pm_alex…) speaches CPUs Native voices, but flat and robotic Fastest: faster than real time on a few cores none
Piper (pt-br-faber) speaches CPUs Native, clearly synthetic Fast none
Chatterbox Multilingual 500M Chatterbox NVIDIA GPU Speaks Portuguese in the voice it imitates; the general multilingual model leaves artefacts and an accent After its first sentence (WAV) or the whole text (MP3) 3.7 GB (bf16)
Qwen3-TTS 0.6B vLLM-Omni NVIDIA GPU Clones a Brazilian speaker in context: accent and rhythm, not only timbre WAV streams from the first audio frame about 5.8 GB
Qwen3-TTS 1.7B vLLM-Omni NVIDIA GPU The same, from a larger model The same about 7.7 GB

Why Qwen3-TTS for Portuguese:

  • A native accent. Qwen3-TTS speaks ten languages, Portuguese among them, and its Base models speak in the voice of a reference recording given with its transcript. Every deployment ships a Brazilian one, pt-BR-1, so the voice, the accent and the rhythm are Brazilian. Chatterbox's general model and its English voices give Portuguese an English accent.
  • Streaming. Asked for WAV or PCM, it sends audio as it is decoded, the first piece after one audio frame; it does not wait for a sentence.
  • Licence. Apache 2.0 for the model and for vLLM-Omni, its server.

What it costs:

  • GPU memory: about 5.8 GB for 0.6B — it runs two engines (one writes the speech codes, one turns them into sound), each in its own process. Beside Whisper on an 8 GB card, take a small Whisper (below).
  • Start time: about one to two and a half minutes for a replica to start (two engines load, then compile and capture their GPU kernels); the first start also downloads 2.5 GB of weights and pulls a 9.4 GB image. Keep it running (min: 1) rather than scaling it to zero.
  • NVIDIA only: vLLM-Omni's CUDA image, with an NVIDIA driver of the 580 series or newer (CUDA 13).

Measured figures

The GPU memory above is computed from Qwen3-TTS's weights and a published measurement of vLLM-Omni on a 16 GB GPU, held to one request at a time. It has not been measured on an 8 GB RTX 4060 yet; check the machine's page after the first start.

Before you begin#

  • A machine with an NVIDIA GPU offering about 6.5 GB for both models (an 8 GB card whose desktop holds 1 GB offers 7), an NVIDIA driver of the 580 series or newer, 12 GB of RAM free, x86-64, and a data location. A Windows PC under WSL2 works.
  • A chat deployment anywhere in the workspace (see Serve an open chat model on one GPU machine).

What each takes of the GPU#

Model Variant Weights GPU memory taken RAM
whisper:base q8_0 82 MB 0.7 GB 2 GiB
whisper:small q8_0 264 MB 0.9 GB 2 GiB
whisper:large-v3-turbo q5_0 574 MB 1.2 GB 2 GiB
qwen3-tts:0.6b bf16 2.5 GB 5.8 GB 9 GiB
qwen3-tts:1.7b bf16 4.5 GB 7.7 GB 11 GiB

Qwen3-TTS 0.6B with Whisper small takes 6.7 GB; with large-v3-turbo q5_0, 7.0 GB — the most an 8 GB card with a desktop offers. The 1.7B model needs a larger GPU (12 GB or more) beside Whisper.

On Linux each model is held to its GPU memory. On Windows (WSL2) nothing holds them to it — each engine's own sizing does — and the machine's page says fractions not enforced here. Qwen3-TTS's speech engine keeps a fixed cache for one request; it does not grow with the GPU.

1. Deploy them#

  1. Eos → Models → Library, Category: transcription, open Whisper, click Small · q8_0, press Deploy. Runs on: A GPU (whisper.cpp: CUDA, or Vulkan on AMD). Name: listen. Clear Scale to zero when idle, press Deploy.
  2. Category: speech, open Qwen3-TTS, click 0.6B · bf16, press Deploy. Runs on is An NVIDIA GPU (text to speech: vLLM-Omni). Name: voice. Clear Scale to zero when idle, press Deploy.
$ astra eos add whisper:small-q8_0
$ astra eos deploy whisper-small-q8-0 --name listen --min 1 --gpu-memory 0.9
$ astra eos add qwen3-tts:0.6b
$ astra eos deploy qwen3-tts-0-6b-bf16 --name voice --min 1
$ curl -fsS -X POST "$ASTRALYX_API/eos/deployments" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"metadata": {"name": "listen"}, "spec": {"model": "whisper-small-q8-0", "gpu_memory_gb": 0.9, "replicas": {"min": 1, "max": 1}}}' > /dev/null
$ curl -fsS -X POST "$ASTRALYX_API/eos/deployments" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"metadata": {"name": "voice"}, "spec": {"model": "qwen3-tts-0-6b-bf16", "replicas": {"min": 1, "max": 1}}}' > /dev/null

Qwen3-TTS takes the GPU memory it needs by default; a Whisper deployment runs on CPUs unless it asks for a GPU (gpu_memory_gb, gpus or gpu_vendor).

The replica's log shows voice pt-BR-1: 7.2 s at 24000 Hz as it starts, then vLLM-Omni loading its two stages. It takes traffic once GET /health answers.

2. Talk#

Open the app, choose the chat model, and press the wave button: what you say goes to listen, the answer is spoken by voice in pt-BR-1 — see Talk to your models.

From a program, they are OpenAI's audio endpoints. vLLM-Omni takes two fields of its own: language (Portuguese, English, Spanish, French, German, Italian, Russian, Chinese, Japanese, Korean; Auto by default — name it: a short sentence can be taken for another language) and stream_format: "audio" (WAV or PCM sent as it is decoded, at speed 1 only).

voz.py
import os
from openai import OpenAI

client = OpenAI(base_url=os.environ["EOS_URL"], api_key=os.environ["ASTRAEUS_API_KEY"])

with client.audio.speech.with_streaming_response.create(
    model="voice", voice="pt-BR-1", input="Olá! O treinamento terminou sem erros.",
    response_format="wav", extra_body={"language": "Portuguese", "stream_format": "audio"},
) as r:
    r.stream_to_file("aviso.wav")

with open("aviso.wav", "rb") as f:
    print(client.audio.transcriptions.create(model="listen", file=f, language="pt").text)

Through the Playground and the app, the language is the one the request names (pt-BR), else the voice's, and WAV is streamed by itself.

How quickly it speaks#

  • WAV or PCM, streamed: the first sound after the text is read and its first audio frame (an eighth of a second of speech) is made and decoded.
  • MP3, FLAC: once the whole text is made. The app asks for MP3 a sentence at a time, its first piece short, so the first sound comes after one short piece is made.
  • A replica answers one request at a time; more replicas (replicas.max) answer more, each with its own GPU memory.

To see it on your machine, time a request:

$ curl -sS -o /dev/null -w 'first byte %{time_starttransfer}s, all %{time_total}s\n' "$EOS_URL/audio/speech" \
    -H "Authorization: Bearer $ASTRAEUS_API_KEY" -H 'content-type: application/json' \
    -d '{"model": "voice", "voice": "pt-BR-1", "input": "Bom dia! Como posso ajudar?", "language": "Portuguese", "response_format": "wav", "stream_format": "audio"}'

The voice#

pt-BR-1 is 7.2 s of a public-domain recording — chapter one of Machado de Assis' Dom Casmurro (1899), read for LibriVox by the volunteer Leni (Internet Archive item dom_casmurro_2102_librivox, Public Domain Mark 1.0) — and what is said in it, "E eles, por graça, chamam-me assim, alguns em bilhetes: Dom Casmurro, domingo vou jantar com você." Qwen3-TTS reads both before each request and continues in that voice.

A deployment has no other voice: Qwen3-TTS's Base models have no speakers of their own, and the server's voice upload is not reachable through Eos. A request may carry its own reference recording (ref_audio as a data: URL, with ref_text); a URL to fetch is refused.

Clone only voices you have the right to use

A recording you send as ref_audio is cloned. Send your own voice, or a voice whose owner gave you permission to clone it for this use.

Engine arguments#

engine_args may add --enforce-eager (no captured GPU kernels: a quicker start, slower speech), --seed, --disable-log-stats, --disable-uvicorn-access-log and --uvicorn-log-level. The rest is set by the deployment.

Troubleshooting#

Symptom Cause Fix
The replica waits, saying how much of the GPU is taken and that it needs 5.8 Whisper or another model holds too much of the GPU. Take a smaller Whisper (small, base), or move it to CPUs.
The replica restarts during its start, out of GPU memory Something outside Eos (a game, the desktop) holds more of the GPU than when it was placed. Free the GPU, or give it more: gpu_memory_gb.
400 Invalid language 'Pt-br' A program sent a language code to the endpoint. Send "language": "Portuguese".
400 streaming requires … stream_format: "audio" with MP3, or with another speed. Ask for WAV or PCM at speed 1, or leave stream_format out.