Voice
Speak instead of typing and hear replies read aloud — locally or in the cloud.
You can talk to the agent instead of typing, and have it read replies aloud. Both recognition and read-aloud come in two kinds: local (offline, free) and cloud (no model download).
Voice input
Turn on "Speak in the composer" under Settings → Preferences → Voice (turn on "Show advanced settings" at the bottom of the settings sidebar first), and a microphone button appears in the input box; ⌘⇧M (Ctrl+Shift+M on Windows / Linux) also starts / stops recording.
| Local | Cloud | |
|---|---|---|
| How | Download a recognition model; text appears as you speak | Pick a provider and add your own key |
| Options | Live model plus a whole-utterance model to finalize; you can also import sherpa-onnx models | OpenAI or a compatible API (Whisper), Alibaba DashScope, Volcengine, iFlytek |
| Platforms | Apple silicon Macs only | All |
Local recognition supports a vocabulary (names you say often, one per line) and corrections (e.g. DeepSeek: deep seek, deep sick) to cut down on mistakes.
Read aloud
With "Read replies aloud" on, the agent reads each reply to you when it finishes.
| Option | Details |
|---|---|
| Edge | Microsoft Edge online voices; free, no setup, needs internet |
| Neox Cloud | Included in your plan, no key |
| OpenAI / compatible API | Your own key; choose model and voice |
| Alibaba CosyVoice | Your own key |
| Local model | Offline, fastest first sentence; Apple silicon Macs only |
Speed is adjustable (0.5–2×). "Shorten long replies first" is on by default and condenses long replies to a few points before reading.
CLI
The CLI supports read-aloud only: /tts chooses Edge, OpenAI or off. There is no voice input.
Availability
| Voice input | Read aloud | |
|---|---|---|
| Desktop | ✓ | ✓ |
| CLI | ✗ | ✓ |
Voice is a way to give input and receive output, not a tool the agent can call.

