What Is Speech to Text? How Voice Recognition Actually Works
Speech to text is software that turns spoken words into written text. It works by recording your voice, converting the sound into a map of its frequencies, and running that through a model trained on huge amounts of transcribed speech, which predicts the most likely words you said.
You'll also see it called automatic speech recognition (ASR), voice recognition, voice typing or dictation. They all describe the same basic job. The one exception worth knowing: "voice recognition" sometimes means identifying who is speaking, which is a different technology called speaker recognition.
What speech to text actually does
At its simplest, a speech-to-text system takes audio in and gives text out. That audio might be you talking into a MacBook mic, a recorded interview, a Zoom call, or a doctor speaking into a headset between appointments.
What happens next depends on the use. Dictation apps type the text wherever your cursor is, while captioning tools put it on screen under a video and meeting tools save it as a transcript with speaker labels. Voice assistants pass it on to something else that decides what to do with your request.
The recognition step underneath is broadly the same in all of them.
How does voice recognition work?
Modern systems follow roughly four steps. Older ones split the work into more pieces, but the shape is similar.
1. Capture the audio
A microphone turns air pressure into a digital signal: thousands of numbers per second describing how loud the sound is at each instant. Mic quality matters more than most people expect. A cheap headset close to your mouth often beats an expensive mic on the far side of a desk, because the model gets more of your voice and less of the room.
2. Turn sound into a spectrogram
Raw audio is awkward for a model to work with, so it gets converted into a spectrogram, a picture of which frequencies are present over time. Most systems use a version called a log-Mel spectrogram, which squeezes frequencies to roughly match how human hearing works.
OpenAI's description of its Whisper model is a good concrete example. Incoming audio is split into 30-second chunks, each chunk becomes a log-Mel spectrogram, and that goes into the model's encoder.
3. Predict the words
This is where the systems have changed most.
Older speech recognition used a chain of separate parts. An acoustic model guessed which speech sounds were present, a pronunciation dictionary mapped sounds to words, and a language model scored which word sequences were plausible. "Recognise speech" and "wreck a nice beach" sound almost identical; the language model is what picks the sensible one.
Most current systems are end-to-end neural networks. Whisper, for instance, is an encoder-decoder Transformer: the encoder turns the spectrogram into an internal representation, and the decoder writes out text one token at a time, using everything it has heard so far as context. OpenAI says it was trained on 680,000 hours of multilingual audio paired with transcripts collected from the web. That scale is a big part of why modern models cope with accents and background noise far better than the dictation software of fifteen years ago.
4. Clean up the output
The raw output then gets tidied. Depending on the system, that means adding punctuation and capitals, turning "twenty five dollars" into "$25", removing filler words like "um", or applying a custom vocabulary so your product names come out right. Some apps now run a second language model over the text to fix grammar or reformat it as an email.
This step is optional, and it's where apps differ most. Two tools using the same recognition model can produce very different-looking text.
Cloud vs on-device speech to text
The biggest practical choice is where the recognition happens.
| Cloud | On-device | |
|---|---|---|
| Where your audio goes | To the provider's servers | Stays on your computer or phone |
| Internet needed | Yes | No |
| Model size | Can be very large | Has to fit your hardware |
| Speed | Depends on your connection | Depends on your chip |
| Cost to the provider | Pays for servers, so usually a subscription | Your hardware does the work |
Cloud speech to text sends your audio to a server, runs it through a large model, and sends text back. Wispr Flow works this way, and its help docs walk users through "No internet" errors when they try to dictate offline. Google Docs voice typing depends on your browser's speech service: Google's help page says the browser "controls the speech-to-text service" and decides how your speech is processed. Cloud models can be enormous and updated constantly, which is a genuine advantage for difficult audio.
On-device speech to text runs the model locally. Apple Silicon Macs include a Neural Engine for machine-learning work, so a local model can keep up with normal speech on a recent laptop. Apple Dictation can process general dictation on your Mac (you can check in Keyboard settings), and apps like Blazing Transcribe, VoiceInk and Superwhisper's free tier run open models such as Whisper locally. Nothing leaves the machine, and it keeps working on a plane.
Which is better depends on what you're dictating. For client notes, medical details or anything under an NDA, keeping audio on your own hardware removes a whole category of risk. For transcribing a noisy multi-speaker recording, a large cloud model may still have the edge.
Where speech to text is used
A short and incomplete list:
- Dictation: writing emails, documents and messages by voice instead of typing. Our guide on how to use speech to text on Mac covers the built-in options.
- Transcribing recordings such as interviews, podcasts, lectures and voice memos. See how to transcribe audio to text if that's your use.
- Meetings, where a tool records the call and produces notes with names attached.
- Healthcare, where clinicians dictate notes instead of typing them into the record afterwards. There's a fuller breakdown in our piece on medical speech to text.
- Live captions on video calls, TV and phones.
- Accessibility, for people who can't type comfortably because of RSI, dyslexia or a motor disability.
- Voice assistants and phone menus, where the transcript is only the first step.
The limits of speech recognition
It's good now. It isn't perfect, and the failure modes are predictable.
Accents and dialects. Models do best on the kinds of speech they saw most of in training. A 2020 study in PNAS tested five major commercial systems and found an average word error rate of 0.35 for Black speakers compared with 0.19 for white speakers. Models have improved since, but the underlying problem of uneven training data hasn't gone away.
Names and jargon. A model has never heard your colleague's surname or your company's internal product names. This is why custom vocabulary features exist, and why medical and legal tools train on domain-specific language.
Noise and crosstalk. Background chatter, music and two people talking at once all hurt accuracy. Honestly, moving the mic closer helps more than any software setting.
Homophones and punctuation. "Their", "there" and "they're" sound identical. The model has to guess from context, and it's often right, but not always.
Made-up text. Large generative models occasionally produce words that weren't said, particularly over long silences or very noisy sections. If a transcript matters, someone should read it.
Accuracy is usually measured as word error rate (WER): the share of words that were substituted, dropped or inserted compared with a human transcript. So a WER of 0.05 is about one mistake in 20 words. We go into what affects that number, and what you can do about it, in our guide to speech to text accuracy.
Trying it on a Mac
The quickest way to see speech to text for yourself is the dictation already built into macOS. Turn it on in System Settings, then Keyboard, then Dictation, and start it with the Dictation shortcut or Edit, then Start Dictation.
If you'd rather hold a key and talk into any app, Blazing Transcribe is our free, open-source Mac app that does exactly that. Hold fn, speak, let go, and the text appears at your cursor. It runs on-device, so your audio never leaves the Mac, and it transcribes at 155× real-time, according to our homepage. It needs macOS 15 on Apple Silicon. You can download it here.
If something isn't working, our speech to text troubleshooting guide covers the usual culprits, from mic permissions to the wrong input device.
FAQ
Is speech to text the same as voice recognition?
Mostly, yes. In everyday use both mean turning speech into text. Strictly speaking, "voice recognition" can also mean speaker recognition, which identifies who is talking rather than what they said. Speech to text, ASR and dictation all refer to the words.
Is speech to text AI?
Modern speech to text is a form of AI. It uses neural networks trained on large collections of recorded speech and matching transcripts. Older systems used statistical models and needed you to train them on your own voice first; current models generally work well out of the box.
Does speech to text work offline?
It depends on the app. On-device tools run the model on your computer and work without a connection. Cloud tools, including Wispr Flow and browser-based voice typing, need the internet because recognition happens on a server.
How accurate is speech to text?
On clear audio from a single speaker with a decent mic, modern systems get most words right. Accuracy drops with heavy accents, background noise, overlapping speakers and unusual vocabulary. Word error rate is the standard measure, and it varies a lot by audio quality, so test with your own voice and your own mic before trusting any headline number.