Skip to content

Talk to your models#

In a chat, you can speak instead of typing and hear the answer instead of reading it. With the workspace's own speech models running — a Whisper that transcribes and a Kokoro or Piper voice, which run on CPUs (Speech on CPUs) — what you say and what you hear stay on your machines. Without them, the app uses the phone's own: the browser's speech recognition (in Chrome and Safari, the browser maker's service) and the phone's voice.

Dictate and read aloud#

  • Dictate: with the message box empty, tap the microphone, speak, tap Done. The text lands in the box for you to check and send.
  • Read aloud: tap the speaker under an answer. With a voice deployment, the first sentence plays while the next are being made; tap again to stop.

Voice mode#

Tap the wave button beside the microphone (in a chat with a model picked): a panel opens at the bottom of the chat.

  1. Tap the big button and speak. The turn ends by itself when you pause (about a second of silence), or tap again when you finish. Or hold the button while you speak and let go to send.
  2. What you said is written down by the workspace's Whisper and shown (“…”), then sent to the chat's model as an ordinary message: it appears in the chat, and its answer is written there as usual.
  3. As soon as the answer's first sentence is complete, it is spoken; the next sentences are made while it plays and follow in order.
  4. When it has said everything, it listens again: carry on talking.
  5. Tap while it thinks or speaks to interrupt: it stops at once and listens to you.

Under the button, how quickly the first word came — from the moment you stopped speaking to the first sound (first word in 1.4 s).

The chat is kept on the phone like any other; voice mode adds nothing else to it. Close the panel with ×.

How quickly it answers#

The first word comes after: writing down what you said, the model's first sentence, and making that sentence's sound. On CPUs only, with Whisper base, a small chat model and Kokoro, expect roughly 2 to 4 seconds:

Step On 4 recent x86 cores, roughly
Uploading your phrase (Opus or AAC, a few KB a second) under 0.1 s
Whisper base writing down a 5-second phrase 0.4–0.8 s
A 1–2B chat model reading the conversation and writing its first sentence 1–2 s
Kokoro making that sentence's sound (Piper: 0.1–0.3 s) 0.3–1 s

A chat model on a GPU, Piper instead of Kokoro, or more cores bring it under 2 seconds. A deployment that was scaled to zero adds its start the first time: keep the speech deployments at one replica at least (Scale to zero when idle off) for conversation. Long chats take longer to read: start a new chat for a new topic.

The app does its part: your phrase is recorded small (Opus in WebM on Android and desktop browsers, AAC in MP4 on iPhone) and sent as it is, nothing waits for the whole answer, the first sentence is at least a few words long (never a lone Sure.), sentences are cut at their real end — not at Dr., e.g., 3.14 or a numbered line — and the voice is warmed up when the panel opens.

Choose who listens and who speaks#

Settings → Voice, for the current workspace, on this phone:

Setting Choices
Transcribe with Auto — the first of the workspace's speech-to-text deployments that is running (else one asleep, which wakes); one of them by name; or This browser's recognition.
Speak with Auto — the first text-to-speech deployment running; one by name; or This device's voice.
Voice With Kokoro: For my language (the phone's language: pf_dora for Portuguese, af_heart for English…) or one of its voices.

Only the workspace's own deployments are used for your voice, never one shared with it.

On an iPhone#

Add the app to the Home Screen and open it from there. The first tap on the wave button (or the speaker) lets the app play sound: iOS allows audio only after a touch. Safari asks once for the microphone. If the answer is silent, check the ring/silent switch and the volume.