Why doesn't a podcast at 2x speed turn shrill, while a song played twice as fast turns the singer into a chipmunk? Changing speed with and without preserving pitch are two completely different technical paths.
Listening to podcasts at 1.5x speed is a daily habit for many commuters. Play a song twice as fast, and the vocalist turns into a squeaky chipmunk. Yet in music-learning apps, a guitar solo can be slowed to half speed while staying perfectly in tune, making it easy to transcribe note by note. All of these are "speed changes"—so why do the outcomes differ so dramatically?
The answer: there are two fundamentally different technical routes for changing audio speed—one brutally simple (resampling), one ingeniously complex (time-stretching). Understand the difference, and you'll know exactly what the "speed," "pitch," and "preserve pitch" options in audio tools actually do.
Why Speed Changes Pitch: Starting with Physics
A sound's pitch is determined by its frequency: the higher the frequency, the higher the pitch. An audio file stores the waveform of a sound as it varies over time.
The most direct way to change speed is resampled playback: reading the waveform data at a different rate. Playing audio at twice the speed compresses every wave cycle to half its duration in time—and halving the period doubles the frequency.
The math of pitch follows the twelve-tone equal temperament: every doubling of frequency raises the pitch by one octave (12 semitones); each semitone up multiplies the frequency by the twelfth root of two (about 1.0595). That gives us:
| Playback speed | Frequency change | Pitch change | How it sounds |
|---|---|---|---|
| 0.5x | halved | down 12 semitones (one octave) | deep, "monster-like" |
| 0.75x | × 0.75 | down about 5 semitones | noticeably muffled |
| 1.5x | × 1.5 | up about 7 semitones | shrill and comical |
| 2.0x | doubled | up 12 semitones (one octave) | the "chipmunk" effect |
So any speed change without special processing inevitably changes pitch—that's physics, not a software flaw. In the vinyl and tape era, this was simply how things worked: spin the record faster, and the sound had to get squeakier.
Time-Stretching: Speed Without the Pitch Shift
Time-stretching tackles a specific problem: change an audio signal's duration while keeping its frequency (pitch) intact. Mathematically, this is a nontrivial task—duration and frequency are entangled within the waveform.
Mainstream algorithms follow a "chop it up and reassemble" approach:
- Frame slicing: cut the audio into many very short frames, each roughly 20–40 milliseconds;
- Respacing: to slow down, overlap adjacent frames more; to speed up, spread them further apart—thereby changing the total duration;
- Overlap-Add (OLA): crossfade at frame boundaries to smooth over the seams.
Because the waveform inside each tiny frame is left untouched, the frequency content survives—pitch stays the same. Meanwhile, the density of the frames changes the duration—speed changes.
Algorithm Evolution and Limits
The problem with basic OLA is that arbitrarily cut frame boundaries create phase discontinuities, which listeners hear as warble, echo-like artifacts, and metallic ringing. Smarter variants were developed to address this:
- WSOLA (Waveform Similarity Overlap-Add): instead of cutting at fixed positions, the algorithm searches nearby for the spot where the waveform is "most similar," so the spliced waveform continues naturally. This is the core idea behind today's mainstream consumer speed-change algorithms, and it works exceptionally well on speech.
- Phase vocoder: operates in the frequency domain—the audio is first transformed into a spectrum, each frequency component's phase is handled separately, and then everything is reassembled. More stable on music (especially harmonically rich content), but prone to a "swimmy" or "hollow" character when pushed.
- Transient preservation: drums and percussion are the natural enemy of speed-change algorithms—transients are so brief that chopping and respacing them blurs and softens their attack. Advanced algorithms detect transients and treat them specially.
The Boundaries of Speed Factors
Time-stretching is not magic—the more extreme the factor, the more audible the artifacts:
| Speed range | Speech quality | Music quality | Typical use |
|---|---|---|---|
| 0.9x – 1.5x | nearly transparent | nearly transparent | fine-tuning audiobooks and lectures |
| 0.75x – 2.0x | good | acceptable | language learning, fast podcast listening |
| 0.5x – 2.5x | clearly processed | noticeable degradation | transcription, extreme speed listening |
| < 0.5x or > 2.5x | severe artifacts | not recommended | special effects |
Speech is the easiest content to handle (a single source with a relatively simple spectrum); multi-instrument music with strong rhythms is far harder. At the same 2x speed, a podcast merely sounds like "fast talking," while a symphony may already have smeared into mush.
Pitch-Shifting: The Inverse of Time-Stretching
The counterpart of time-stretching is pitch-shifting: changing the pitch while keeping the duration intact. Interestingly, the two are technically two sides of the same coin—pitch-shifting can be achieved by combining "time-stretch + resample":
Raise pitch without changing speed = resample up (which also speeds it up)
→ then time-stretch back to the original duration (keeping the new pitch)
Pitch shift is usually expressed in semitones: +1 semitone multiplies the frequency by about 1.0595, and +12 semitones is one octave up. Finer adjustments use cents, where 1 semitone = 100 cents—under ideal conditions, the human ear can distinguish differences of roughly 5–6 cents.
Typical applications:
- Key-changing accompaniment: raise a backing track by 2 semitones to fit a singer's range;
- Sound design: pitch voices down for weight, up for cartoonish characters;
- Pitch correction: nudge a vocal by a few dozen cents to bring an off-key performance back in line (Auto-Tune used gently).
Formants: The Real Culprit Behind the "Chipmunk"
If you pitch a vocal up by 7 semitones, what you get isn't just "the same voice, higher"—it's a thin, cartoonish squeak, even when the time-stretching algorithm is flawless. The reason is formants.
A human voice is shaped by two components working together:
- Vocal cord vibration: produces the fundamental frequency, which determines pitch;
- Vocal tract resonance: the oral and nasal cavities form a resonating chamber that amplifies specific frequency regions—those boosted regions are the formants. Their distribution defines timbre: whether a voice sounds like an adult or a child, whether a vowel sounds like "ah" or "ee."
A straightforward pitch shift moves the fundamental and the formants together: shifting up 7 semitones also pushes the vocal-tract resonance regions higher, effectively swapping an "adult resonance chamber" for a "child-sized" one—hence the chipmunk. Shift far down instead, and you get the deep "monster voice."
Professional pitch-shifting tools separate and preserve the formants: they move only the fundamental (changing pitch) while pulling the formants back to their original positions (keeping timbre). That's the technical dividing line between "pitch change with timbre preserved" and "pitch change with timbre shifted."
The Full Map: Four Combinations
Combining "change speed or not" with "change pitch or not" covers every audio speed/pitch scenario:
| Goal | Technique | Result | Typical use |
|---|---|---|---|
| Speed + pitch change | resampling | fast & shrill / slow & deep | creative effects, tape-style speedups |
| Speed change, pitch preserved | time-stretching | tempo changes, voice stays | audiobooks, lectures, transcription |
| Pitch change, speed preserved | pitch-shifting (stretch + resample combo) | key changes, rhythm intact | key-changing backing tracks, tuning |
| Pitch + timbre preserved, speed intact | formant-preserving shift | natural-sounding transposition | professional vocal processing |
Common Misconceptions
- "Speed change without pitch shift is the higher-quality, lossless option." Quite the opposite—resampling is nearly lossless (only the pitch changes), while time-stretching introduces processing artifacts. Which route to choose depends on whether you want the pitch to change, not on which is "more advanced."
- "Listening at 2x saves half the time with no comprehension loss." Most people comprehend fine at 1.5–2x, but beyond 2.5x, information retention drops noticeably—and algorithm artifacts start competing for your attention.
- "Slowing down music for practice costs nothing in quality." Below 0.5x, transient smearing and harmonic drift become clearly audible; for transcribing harmony, staying at or above 0.6x is advisable.
- "Pitch-shifting just makes a voice squeakier or deeper." That's actually the sound of formants shifting along with the pitch; formant-preserving shift changes only the note, not the character of the voice.
Practical Tips
- Podcasts and audiobooks: start at 1.25x and adapt gradually; 1.5x is the sweet spot between efficiency and comprehension. Always use the "preserve pitch" mode.
- Language learning: 0.75–0.9x slow playback for shadowing works better than an extreme 0.5x—less distortion, more natural intonation.
- Music practice: use time-stretch mode for slow transcription; use resampling when practicing rhythmic feel (hearing pitch rise with speed actually reinforces the tempo-pitch relationship).
- Creative effects: for tape-fast-forward or vintage cartoon sounds, use resampling directly—a "physically correct" time-stretch would defeat the purpose.
- Check the hardest passages: after processing, listen specifically to percussion-heavy sections and rapid speech—that's where artifacts show up first.
Further Reading
- Video Speed Change: The Secrets of Fast-Forward and Slow Motion — frame interpolation and slow motion on the video side
- Volume, Gain, and Loudness: Making Audio Sound Just Right — the other most commonly adjusted audio parameter
- Audio Cutting and Merging: Trimming and Joining Sound — the fundamentals of working on a timeline
This site's audio speed tool offers both "preserve pitch" and "pitch follows speed" modes—and after reading this article, you know they correspond to the time-stretching and resampling routes respectively, and which scenario calls for which.