Skip to content
Agent

Voice

Speak instead of typing and hear replies read aloud — locally or in the cloud.

You can talk to the agent instead of typing, and have it read replies aloud. Both recognition and read-aloud come in two kinds: local (offline, free) and cloud (no model download).

Voice input

Turn on "Speak in the composer" under Settings → Preferences → Voice (turn on "Show advanced settings" at the bottom of the settings sidebar first), and a microphone button appears in the input box; ⌘⇧M (Ctrl+Shift+M on Windows / Linux) also starts / stops recording.

LocalCloud
HowDownload a recognition model; text appears as you speakPick a provider and add your own key
OptionsLive model plus a whole-utterance model to finalize; you can also import sherpa-onnx modelsOpenAI or a compatible API (Whisper), Alibaba DashScope, Volcengine, iFlytek
PlatformsApple silicon Macs onlyAll

Local recognition supports a vocabulary (names you say often, one per line) and corrections (e.g. DeepSeek: deep seek, deep sick) to cut down on mistakes.

Read aloud

With "Read replies aloud" on, the agent reads each reply to you when it finishes.

OptionDetails
EdgeMicrosoft Edge online voices; free, no setup, needs internet
Neox CloudIncluded in your plan, no key
OpenAI / compatible APIYour own key; choose model and voice
Alibaba CosyVoiceYour own key
Local modelOffline, fastest first sentence; Apple silicon Macs only

Speed is adjustable (0.5–2×). "Shorten long replies first" is on by default and condenses long replies to a few points before reading.

CLI

The CLI supports read-aloud only: /tts chooses Edge, OpenAI or off. There is no voice input.

Availability

Voice inputRead aloud
Desktop✓✓
CLI✗✓

Voice is a way to give input and receive output, not a tool the agent can call.