How Speech-to-Text Works: From Sound Waves to Text

Speech-to-text turns a continuous sound wave into discrete words, and it does it in a fixed series of steps. The microphone signal is sampled into numbers, the numbers are turned into a spectrogram, a neural network maps the spectrogram to likely text, and a formatting step makes that text readable. Every modern system — a phone keyboard, a meeting notetaker, a dictation app — follows some version of this pipeline.

The short answer to "how does it know what I said" is: it does not know, it estimates. The model outputs the most probable text for the sound it received, based on what it saw during training. That single fact explains most of the behaviour you notice in practice: why common phrases come out perfectly, why a colleague's surname does not, and why a noisy pause can turn into a sentence nobody said.

This article walks through the pipeline stage by stage, without the maths and without any particular vendor's internals.

Step 1: How does a microphone turn sound into numbers?

Sound is a pressure wave. The microphone converts it into a voltage, and an analogue-to-digital converter measures that voltage many thousands of times per second. Each measurement is a sample.

Music is usually recorded at 44,100 or 48,000 samples per second. Speech recognition rarely needs that much: the frequencies that carry speech sit mostly below 8 kHz, so most recognisers work at 16,000 samples per second (16 kHz). The Whisper paper, for example, states that all audio is resampled to 16,000 Hz before anything else happens. Telephone audio is often 8 kHz, which is one reason phone recordings are harder to transcribe — part of the speech signal was never captured.

At this stage there are no words, no letters and no silence detection. There is only a long list of numbers: one minute of 16 kHz audio is 960,000 of them.

Step 2: What is a spectrogram and why not use the raw waveform?

The raw waveform says how loud the signal is at each instant, but the information that separates one speech sound from another is in its frequencies. An "s" is a burst of high-frequency noise; a vowel like "a" has strong, stable bands of energy lower down.

So the recogniser cuts the audio into short overlapping windows — typically about 25 milliseconds each, shifted by 10 milliseconds — and measures how much energy each window has at each frequency. Stacked side by side, these measurements form a spectrogram: time on one axis, frequency on the other, loudness as brightness.

Most systems use a log-Mel spectrogram. "Mel" means the frequency bands are spaced the way human hearing perceives pitch — dense at low frequencies, sparse at high ones. "Log" compresses loudness the way the ear does. Whisper's original models use 80 Mel channels on 25 ms windows with a 10 ms stride; its large-v3 model moved to 128. Either way, one second of audio becomes about 100 columns of numbers, a far more compact and informative input than 16,000 raw samples.

Step 3: How does the model map sound to words?

This is where the approaches have changed most over the last fifteen years.

The classic pipeline: acoustic model, dictionary, language model

For decades, recognisers were built from separate parts. An acoustic model estimated which speech sounds (phonemes) were likely in each frame, a pronunciation dictionary listed how each word is built from phonemes, and a language model scored which word sequences are plausible. Hidden Markov models tied it together over time. Around 2012, deep neural networks replaced the older Gaussian mixture models inside this pipeline and cut error rates substantially — the turning point is usually dated to the joint paper by Hinton and colleagues from four research groups.

End-to-end models

Modern systems mostly train one neural network to go straight from spectrogram to text. Three designs dominate:

Approach How it aligns sound and text Typical strength
CTC (Graves et al., 2006) Emits a label or a "blank" for every frame, then collapses repeats Fast, simple, good for streaming
Transducer, RNN-T (Graves, 2012) Decides at each step whether to emit a token or read the next frame Streaming with low latency
Attention encoder-decoder An encoder reads the whole segment, a decoder writes text token by token High accuracy on complete utterances

Whisper is an example of the third kind: an encoder-decoder Transformer that reads 30-second chunks and writes text. Streaming transcription — words appearing while you are still talking — favours the first two, because they can emit output before the sentence is finished.

The output is usually not whole words but subword tokens: pieces like "trans", "cript", "ion". That lets a model spell words it never saw as a whole, which is why a new product name can come out as a plausible but wrong combination of fragments.

Where the language knowledge comes from

An end-to-end model has no separate dictionary, but it still carries a strong sense of which word sequences are likely — learned from the text paired with its training audio. Scale matters here: Whisper's original models were trained on 680,000 hours of labelled audio, of which 117,000 hours covered 96 languages other than English; large-v3 used 1 million hours of weakly labelled audio plus 4 million hours of pseudo-labelled audio. Languages and accents that are rare in the training data are recognised less accurately, for the same reason rare names are.

Step 4: Decoding — choosing the final text

The network does not output one answer. At each step it outputs a probability for every possible token. Decoding is the search for the best overall sequence: greedy decoding takes the top token each time, while beam search keeps several candidate sentences alive and picks the best complete one.

This is also where vocabulary hints act. Many engines accept a list of expected terms or a short text prompt; OpenAI's prompting guide for Whisper, for instance, shows a prompt that lists product and company names to steer spelling. A hint shifts probabilities toward those words but cannot rescue audio the model genuinely cannot hear.

Step 5: Punctuation and formatting

Raw recognition output tends to look like "so the meeting is at three thirty on march fifth". Turning it into "So the meeting is at 3:30 on March 5th." is a separate job:

Some models, Whisper included, learn to output punctuated text directly because their training transcripts were punctuated. Others run a dedicated model after recognition. Either way, formatting errors are a different class of error from recognition errors: the words can be right and the text still hard to read.

Why does speech-to-text make mistakes?

Knowing the pipeline makes the typical errors predictable:

  1. Rare words lose to common ones. A surname, an internal project code or a brand-new tool sounds like something more frequent, and the more frequent thing wins. Vocabulary hints are the direct fix.
  2. Missing signal cannot be recovered. A laptop microphone across the room, 8 kHz phone audio or two people talking at once remove information before the model sees it.
  3. Generative models can hallucinate. A decoder that writes text token by token sometimes produces fluent sentences during silence or noise. The 2024 FAccT paper Careless Whisper found entire invented phrases or sentences in roughly 1% of the Whisper transcriptions it examined. Removing silence before recognition reduces this.
  4. Who spoke is a separate problem. Recognition produces words; assigning them to speakers is speaker diarization, a different model with its own errors.

Accuracy is usually reported as word error rate (WER): substitutions, deletions and insertions divided by the number of words actually spoken. It is useful for comparing systems on the same audio, but figures from different vendors are not directly comparable: the test recordings differ, and so do the normalisation rules that decide whether "3:30" and "three thirty" count as the same.

What this means when you choose a tool

For everyday use, a few practical conclusions follow from the pipeline:

In Speak-Y, dictation works in 67+ languages with automatic punctuation, and meeting mode records calls locally without a bot joining, transcribes by speaker and keeps transcripts on your device by default — the trade-offs of bot-free recording are covered in recording meetings without bots. Once transcripts exist, the next question is what you do with them: a local MCP server lets AI assistants such as Claude or Cursor search and read them, which is explained from the ground up in what an MCP server is.

FAQ

How does speech-to-text work, in one paragraph?

A microphone turns air pressure into a stream of numbers, typically 16,000 samples per second for speech recognition. The software converts short overlapping slices of that signal into a spectrogram — a picture of which frequencies are loud at each moment. A neural network reads the spectrogram and predicts the most likely sequence of words or word pieces, and a final step adds punctuation, capitalisation and number formatting.

What is a spectrogram and why does speech recognition use it?

A spectrogram shows how the energy of a sound is spread across frequencies over time. Speech recognisers usually use a log-Mel spectrogram, which spaces frequencies the way human hearing does. OpenAI's Whisper, for example, computes an 80-channel log-Mel spectrogram on 25-millisecond windows with a 10-millisecond stride, which gives the model 100 feature frames per second of audio.

Why does speech-to-text get names and technical terms wrong?

The model picks the most probable text for the sound it hears, and probability comes from training data. A rare surname or a new product name appeared seldom or never in that data, so a common word that sounds similar wins. Vocabulary hints help: many engines accept a list of expected terms or a short text prompt that biases the output toward those spellings.

Can speech-to-text invent words that were never said?

Yes. Models that generate text token by token can produce fluent phrases during silence or noise. A 2024 FAccT study, Careless Whisper, found that roughly 1% of Whisper transcriptions in its sample contained entire hallucinated phrases or sentences. Trimming silence before recognition and reviewing important transcripts reduce the risk.

What is word error rate?

Word error rate (WER) is the standard accuracy metric: substitutions plus deletions plus insertions, divided by the number of words in the reference transcript. Lower is better, and it can exceed 100% when a system inserts many extra words. WER figures are only comparable when they come from the same test audio and the same text normalisation rules.