Skip to content

Speech on CPUs: transcribe and speak#

You give a workspace ears and a voice without a GPU: Whisper turns speech into text (served by whisper.cpp) and Kokoro — or a Piper voice — turns text into speech (served by speaches), both on the CPUs of one machine. The app's microphone and its voice mode use them, the Playground's Transcribe and Speak tabs try them, and programs call them with the OpenAI API. Recordings and texts stay on your machines.

What you need:

  • A machine of the cluster with 4 CPU cores and 4 GB of RAM free (GPU or not), x86-64 (whisper.cpp's image is built for x86-64 only), with a data location.
  • For calls from programs: a key and a gateway, and the variables in How the recipes are written.

Choose the sizes#

Speech to text — the library's Whisper family, its ggml variants (whisper.cpp; the safetensors ones are vLLM's, for GPUs):

Model Variant (default first) Weights RAM asked CPU cores (default) A 5-second phrase on 4 cores, roughly For
whisper:tiny q8_0, q5_1, f16 44 MB 2 GiB 2 0.2–0.4 s Short English commands; weakest elsewhere
whisper:base q8_0, q5_1, f16 82 MB 2 GiB 2 0.4–0.8 s Dictation: the best speed for its accuracy
whisper:small q8_0, q5_1, f16 264 MB 2 GiB 4 1–2 s Better accuracy, other languages, names
whisper:medium q5_0, q8_0, f16 539 MB 2 GiB 4 3–5 s Recordings transcribed after the fact
whisper:large-v3-turbo q5_0, q8_0, f16 574 MB 2 GiB 4 4–8 s The most accurate on CPUs; for recordings, not conversation

Whisper listens in 30-second windows: a short phrase costs about as much as half a minute of audio, and a longer recording proportionally more. The times are a guide, from whisper.cpp's own benchmarks scaled to four recent x86 cores; yours depend on the CPU (AVX2 or AVX-512 helps a lot), on its memory bandwidth and on what else runs. Quantized variants (q5, q8) are smaller and faster than f16 for nearly the same text.

Text to speech — Kokoro (one model, 54 voices: American and British English, Spanish, French, Hindi, Italian, Japanese, Chinese and Brazilian Portuguese) or a Piper voice (one voice per model: Faber, Brazilian Portuguese; LJSpeech, American English):

Model Variant Weights RAM asked CPU cores (default) First audio of a sentence, roughly
kokoro:82m int8 (default), fp16, fp32 121–354 MB 2 GiB 2 0.3–1 s; faster than real time
piper:pt-br-faber, piper:en-us-ljspeech medium 63 MB 2 GiB 2 0.1–0.3 s; many times faster than real time

Kokoro sounds more natural; Piper is lighter and quicker.

Both on one machine. Whisper base and Kokoro together ask for 2 + 2 CPU cores and about 2.7 GiB of RAM (each its weights plus its engine's own: 1 GiB for whisper.cpp, 1.5 GiB for speaches; the console and the plan round them to whole GiB), and use about 1.5 GB of it in practice. CPU replicas use the machine's idle cores beyond what they ask for, so both get faster on a bigger machine. To also chat on CPUs, add the chat model's own needs (a 1.5B model at Q4_K_M, about 2 GB more): 8 GB of RAM in all is comfortable.

1. Deploy them#

  1. Eos → Models → Library, Category: transcription, open Whisper, click Base · q8_0 and press Deploy. Runs on says CPUs (speech to text: whisper.cpp): there is nothing to choose. Name: listen. Clear Scale to zero when idle so it answers at once, and press Deploy.
  2. Category: speech, open Kokoro, click 82M · int8, press Deploy. Name: voice. Clear Scale to zero when idle, press Deploy.
$ astra eos add whisper:base
$ astra eos deploy whisper-base-q8-0 --name listen --min 1
$ astra eos add kokoro:82m
$ astra eos deploy kokoro-82m-int8 --name voice --min 1

No --cpu is needed: a speech model runs on CPUs.

$ for m in whisper:base kokoro:82m; do
    curl -fsS -X POST "$ASTRALYX_API/eos/catalog-models" -H "Authorization: Bearer $ASTRALYX_TOKEN" \
      -H 'content-type: application/json' -d "{\"catalog\": \"$m\"}" | jq -r .metadata.name
  done
whisper-base-q8-0
kokoro-82m-int8
$ curl -fsS -X POST "$ASTRALYX_API/eos/deployments" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"metadata": {"name": "listen"}, "spec": {"model": "whisper-base-q8-0", "replicas": {"min": 1, "max": 1}}}' > /dev/null
$ curl -fsS -X POST "$ASTRALYX_API/eos/deployments" -H "Authorization: Bearer $ASTRALYX_TOKEN" -H 'content-type: application/json' \
    -d '{"metadata": {"name": "voice"}, "spec": {"model": "kokoro-82m-int8", "replicas": {"min": 1, "max": 1}}}' > /dev/null

The spec names no accelerator: a deployment of a speech model that asks for no GPUs is stored as "accelerator": "cpu".

Each first start downloads the weights to the machine (82 MB and 121 MB) and pulls the engine's image. The deployments' pages show Engine whisper.cpp on CPUs and speaches on CPUs, and Serves speech to text and text to speech. Kokoro is loaded with its first request (a second or two) and then kept.

2. Try them#

  • Eos → Playground → Transcribe, pick listen, press Record, say something, Stop, then Transcribe.
  • Eos → Playground → Speak, pick voice, type a sentence, choose a voice (pf_dora · pt-BR for Portuguese), press Speak. Under the audio: how long the first audio took.

3. Use them in the app#

Open the app on your phone. With listen and voice running, the microphone writes down what you say with listen, the speaker under an answer reads it with voice, and the wave button starts a spoken conversation with the chat's model — see Talk to your models. Settings → Voice chooses among several.

4. Call them from a program#

speech.py
import os
from openai import OpenAI

client = OpenAI(base_url=os.environ["EOS_URL"], api_key=os.environ["ASTRAEUS_API_KEY"])

# Text to speech: an MP3, in Kokoro's Brazilian Portuguese voice.
with client.audio.speech.with_streaming_response.create(
    model="voice", voice="pf_dora", input="Olá! O treinamento terminou sem erros."
) as r:
    r.stream_to_file("aviso.mp3")

# Speech to text: that recording, back as text.
with open("aviso.mp3", "rb") as f:
    print(client.audio.transcriptions.create(model="listen", file=f, language="pt").text)
$ pip install openai
$ python speech.py
Olá! O treinamento terminou sem erros.

whisper.cpp takes any recording — WAV, MP3, FLAC, OGG, WebM, MP4/M4A — and converts it on the machine. Ask response_format="verbose_json" for segments with their times, or srt / vtt for subtitles.

Clean up#

Eos → Deployments, open each of listen and voice, Delete.

$ astra eos deployments delete listen
$ astra eos deployments delete voice
$ for d in listen voice; do
    curl -fsS -X DELETE "$ASTRALYX_API/eos/deployments/$d" -H "Authorization: Bearer $ASTRALYX_TOKEN"
  done

The models and their weights stay until you delete them too.