Speech on CPUs: transcribe and speak#
You give a workspace ears and a voice without a GPU: Whisper turns speech into text (served by whisper.cpp) and Kokoro — or a Piper voice — turns text into speech (served by speaches), both on the CPUs of one machine. The app's microphone and its voice mode use them, the Playground's Transcribe and Speak tabs try them, and programs call them with the OpenAI API. Recordings and texts stay on your machines.
What you need:
- A machine of the cluster with 4 CPU cores and 4 GB of RAM free (GPU or not), x86-64 (whisper.cpp's image is built for x86-64 only), with a data location.
- For calls from programs: a key and a gateway, and the variables in How the recipes are written.
Choose the sizes#
Speech to text — the library's Whisper family, its ggml variants
(whisper.cpp; the safetensors ones are vLLM's, for GPUs):
| Model | Variant (default first) | Weights | RAM asked | CPU cores (default) | A 5-second phrase on 4 cores, roughly | For |
|---|---|---|---|---|---|---|
whisper:tiny |
q8_0, q5_1, f16 |
44 MB | 2 GiB | 2 | 0.2–0.4 s | Short English commands; weakest elsewhere |
whisper:base |
q8_0, q5_1, f16 |
82 MB | 2 GiB | 2 | 0.4–0.8 s | Dictation: the best speed for its accuracy |
whisper:small |
q8_0, q5_1, f16 |
264 MB | 2 GiB | 4 | 1–2 s | Better accuracy, other languages, names |
whisper:medium |
q5_0, q8_0, f16 |
539 MB | 2 GiB | 4 | 3–5 s | Recordings transcribed after the fact |
whisper:large-v3-turbo |
q5_0, q8_0, f16 |
574 MB | 2 GiB | 4 | 4–8 s | The most accurate on CPUs; for recordings, not conversation |
Whisper listens in 30-second windows: a short phrase costs about as much
as half a minute of audio, and a longer recording proportionally more. The
times are a guide, from whisper.cpp's own benchmarks scaled to four recent
x86 cores; yours depend on the CPU (AVX2 or AVX-512 helps a lot), on its
memory bandwidth and on what else runs. Quantized variants (q5, q8)
are smaller and faster than f16 for nearly the same text.
Text to speech — Kokoro (one model, 54 voices: American and British English, Spanish, French, Hindi, Italian, Japanese, Chinese and Brazilian Portuguese) or a Piper voice (one voice per model: Faber, Brazilian Portuguese; LJSpeech, American English):
| Model | Variant | Weights | RAM asked | CPU cores (default) | First audio of a sentence, roughly |
|---|---|---|---|---|---|
kokoro:82m |
int8 (default), fp16, fp32 |
121–354 MB | 2 GiB | 2 | 0.3–1 s; faster than real time |
piper:pt-br-faber, piper:en-us-ljspeech |
medium |
63 MB | 2 GiB | 2 | 0.1–0.3 s; many times faster than real time |
Kokoro sounds more natural; Piper is lighter and quicker.
Both on one machine. Whisper base and Kokoro together ask for
2 + 2 CPU cores and about 2.7 GiB of RAM (each its weights plus its
engine's own: 1 GiB for whisper.cpp, 1.5 GiB for speaches; the console and
the plan round them to whole GiB), and use about 1.5 GB of it in practice.
CPU replicas use the machine's idle cores beyond what they ask for, so both
get faster on a bigger machine. To also chat on CPUs, add the chat model's
own needs (a 1.5B model at Q4_K_M, about 2 GB more): 8 GB of RAM in all
is comfortable.
1. Deploy them#
- Eos → Models → Library, Category: transcription, open
Whisper, click Base · q8_0 and press Deploy. Runs on
says CPUs (speech to text: whisper.cpp): there is nothing to
choose. Name:
listen. Clear Scale to zero when idle so it answers at once, and press Deploy. - Category: speech, open Kokoro, click 82M · int8, press
Deploy. Name:
voice. Clear Scale to zero when idle, press Deploy.
$ astra eos add whisper:base
$ astra eos deploy whisper-base-q8-0 --name listen --min 1
$ astra eos add kokoro:82m
$ astra eos deploy kokoro-82m-int8 --name voice --min 1
No --cpu is needed: a speech model runs on CPUs.
$ for m in whisper:base kokoro:82m; do
curl -fsS -X POST "$ASTRALYX_API/eos/catalog-models" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
-H 'content-type: application/json' -d "{\"catalog\": \"$m\"}" | jq -r .metadata.name
done
whisper-base-q8-0
kokoro-82m-int8
$ curl -fsS -X POST "$ASTRALYX_API/eos/deployments" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"metadata": {"name": "listen"}, "spec": {"model": "whisper-base-q8-0", "replicas": {"min": 1, "max": 1}}}' > /dev/null
$ curl -fsS -X POST "$ASTRALYX_API/eos/deployments" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
-d '{"metadata": {"name": "voice"}, "spec": {"model": "kokoro-82m-int8", "replicas": {"min": 1, "max": 1}}}' > /dev/null
The spec names no accelerator: a deployment of a speech model that
asks for no GPUs is stored as "accelerator": "cpu".
Each first start downloads the weights to the machine (82 MB and 121 MB) and pulls the engine's image. The deployments' pages show Engine whisper.cpp on CPUs and speaches on CPUs, and Serves speech to text and text to speech. Kokoro is loaded with its first request (a second or two) and then kept.
2. Try them#
- Eos → Playground → Transcribe, pick
listen, press Record, say something, Stop, then Transcribe. - Eos → Playground → Speak, pick
voice, type a sentence, choose a voice (pf_dora · pt-BRfor Portuguese), press Speak. Under the audio: how long the first audio took.
3. Use them in the app#
Open the app on your phone. With listen and voice
running, the microphone writes down what you say with listen, the
speaker under an answer reads it with voice, and the wave button starts
a spoken conversation with the chat's model — see
Talk to your models. Settings → Voice chooses
among several.
4. Call them from a program#
import os
from openai import OpenAI
client = OpenAI(base_url=os.environ["EOS_URL"], api_key=os.environ["ASTRAEUS_API_KEY"])
# Text to speech: an MP3, in Kokoro's Brazilian Portuguese voice.
with client.audio.speech.with_streaming_response.create(
model="voice", voice="pf_dora", input="Olá! O treinamento terminou sem erros."
) as r:
r.stream_to_file("aviso.mp3")
# Speech to text: that recording, back as text.
with open("aviso.mp3", "rb") as f:
print(client.audio.transcriptions.create(model="listen", file=f, language="pt").text)
whisper.cpp takes any recording — WAV, MP3, FLAC, OGG, WebM, MP4/M4A — and
converts it on the machine. Ask response_format="verbose_json" for
segments with their times, or srt / vtt for subtitles.
Clean up#
The models and their weights stay until you delete them too.