// docs

Voice

Why type every prompt? Explaining a goal out loud is quicker than typing it, and easier to get into the detail. Speak your prompt and S.E.N.T.R.I answers back once the task is done.

Speak your prompt

Press the microphone in the prompt line, or push to talk with Shift+Space, and say what you want. Your voice shows live in the chat as a line that moves while you speak. The recording is turned into text on your own machine and sent as your prompt, exactly as if you had typed it. Every agent in the Agent Workspace has a microphone of its own.

Speech-to-text models

Speech-to-text runs locally, so your recordings are not sent to a provider to be transcribed. Choose the model in Settings:

ModelNotes
Parakeet V3The model voice input launched with in v5.1.
Nemotron 3.5 ASR 0.6BA compact model covering 40 languages, added in v5.2.3.
Whisper V3 turboWhisper's turbo model, added in v5.2.4.

Hear it answer

When a prompt was spoken, the answer is read back to you, voiced by Grok Voice TTS through OpenRouter. The line in the chat turns purple while S.E.N.T.R.I is the one talking, and the speaker button in the prompt line mutes it.

A spoken turn gets an answer written to be listened to:

  • One to three short sentences: the outcome, and the one detail that matters.
  • No tables, lists, headings or code — unless you asked for one, and then you get it.
  • No hostnames, paths or numbers you did not ask for, and no running commentary about what it is about to do.

Only the final answer changes. The work in between — every skill call, every approval — is exactly what it would be for a typed prompt.

Voice from your phone

Voice notes sent to S.E.N.T.R.I on WhatsApp are transcribed and handed to the main agent like any other prompt, and it can answer with a voice reply. See Chat Integrations.

Which plan

Local speech-to-text and the voice agent are part of the S.E.N.T.R.I plan. See Plans & Features.