Cue · Glossary · Speech to Text

Speech to Text

Speech to text is the conversion of spoken audio into written words by software. It is the capability underneath dictation, voice notes, and transcription, and what happens to the text afterwards depends on the product you use.

Definition

Speech to text, also written as speech-to-text and shortened to STT, is the process of turning spoken audio into written words. A speech recognition model receives audio from a microphone or a recorded file and returns a transcript. The term names the conversion itself rather than any single app, so two tools can both do speech to text and still feel very different to use.

Speech to text vs dictation vs voice AI

These three labels overlap, which is why they are easy to confuse. Speech to text is the underlying conversion. Dictation is one product use of it, where the resulting text is placed into the field you are typing in. Voice AI is the broadest of the three and may go beyond a transcript to interpret a request, use context you have allowed, or prepare a next step.

How it works, and where the text ends up

Audio is captured, split into manageable segments, and mapped to words by a recognition model that usually adds punctuation and capitalization. Some tools stream a partial transcript while you speak; others return the full text once the audio ends. The output then goes somewhere specific: into the text field you have focused, onto the clipboard, into a saved transcript, or into the tool's own notes view. Check that destination before you commit to a workflow, because it decides how much copying and cleanup you do afterwards.

What varies between products

Supported languages, handling of names and technical vocabulary, speaker labels, timestamps, recording length limits, offline support, and whether processing happens on your device or through a service all differ. Cost models differ too, from included desktop features to per-minute transcription. Test with your own microphone, accent, and subject matter, since general claims rarely describe every setup.

Permissions and review

Speech to text needs microphone access, and capturing a meeting or a call can also involve system audio and the awareness of the other people present. Look at what a tool records, whether audio and transcripts are sent to a service, how long they are retained, and whether you can delete them. Read the transcript before you forward or publish it, because recognition errors tend to land in names, numbers, and specialist terms where they are hardest to spot.

Speech to text in Cue

Cue is a desktop voice AI for macOS 13 or later and Windows 10 or later. For dictation, focus a text field in a supported app, use the Dictation shortcut shown in your installed version, speak, then review the inserted text. For a longer session you can start Cue's desktop recorder yourself; it shows a progressive transcript while the meeting runs, returns structured notes and action items after you stop it, and nothing joins the participant list. Cue is free to start; Cue Plus is $19.99/month.

Want to try a voice AI agent? Try Cue free — Mac & Windows, $19.99/mo Plus tier.