Audio Speed and Pitch: The Art of Time-Stretching

Why doesn't a podcast at 2x speed turn shrill, while a song played twice as fast turns the singer into a chipmunk? Changing speed with and without preserving pitch are two completely different technical paths.

Listening to podcasts at 1.5x speed is a daily habit for many commuters. Play a song twice as fast, and the vocalist turns into a squeaky chipmunk. Yet in music-learning apps, a guitar solo can be slowed to half speed while staying perfectly in tune, making it easy to transcribe note by note. All of these are "speed changes"—so why do the outcomes differ so dramatically?

The answer: there are two fundamentally different technical routes for changing audio speed—one brutally simple (resampling), one ingeniously complex (time-stretching). Understand the difference, and you'll know exactly what the "speed," "pitch," and "preserve pitch" options in audio tools actually do.

Why Speed Changes Pitch: Starting with Physics

A sound's pitch is determined by its frequency: the higher the frequency, the higher the pitch. An audio file stores the waveform of a sound as it varies over time.

The most direct way to change speed is resampled playback: reading the waveform data at a different rate. Playing audio at twice the speed compresses every wave cycle to half its duration in time—and halving the period doubles the frequency.

The math of pitch follows the twelve-tone equal temperament: every doubling of frequency raises the pitch by one octave (12 semitones); each semitone up multiplies the frequency by the twelfth root of two (about 1.0595). That gives us:

Playback speedFrequency changePitch changeHow it sounds
0.5xhalveddown 12 semitones (one octave)deep, "monster-like"
0.75x× 0.75down about 5 semitonesnoticeably muffled
1.5x× 1.5up about 7 semitonesshrill and comical
2.0xdoubledup 12 semitones (one octave)the "chipmunk" effect

So any speed change without special processing inevitably changes pitch—that's physics, not a software flaw. In the vinyl and tape era, this was simply how things worked: spin the record faster, and the sound had to get squeakier.

Time-Stretching: Speed Without the Pitch Shift

Resampling changes speed and pitch together; time-stretch changes speed only

Time-stretching tackles a specific problem: change an audio signal's duration while keeping its frequency (pitch) intact. Mathematically, this is a nontrivial task—duration and frequency are entangled within the waveform.

Mainstream algorithms follow a "chop it up and reassemble" approach:

  1. Frame slicing: cut the audio into many very short frames, each roughly 20–40 milliseconds;
  2. Respacing: to slow down, overlap adjacent frames more; to speed up, spread them further apart—thereby changing the total duration;
  3. Overlap-Add (OLA): crossfade at frame boundaries to smooth over the seams.

Because the waveform inside each tiny frame is left untouched, the frequency content survives—pitch stays the same. Meanwhile, the density of the frames changes the duration—speed changes.

Algorithm Evolution and Limits

The problem with basic OLA is that arbitrarily cut frame boundaries create phase discontinuities, which listeners hear as warble, echo-like artifacts, and metallic ringing. Smarter variants were developed to address this:

  • WSOLA (Waveform Similarity Overlap-Add): instead of cutting at fixed positions, the algorithm searches nearby for the spot where the waveform is "most similar," so the spliced waveform continues naturally. This is the core idea behind today's mainstream consumer speed-change algorithms, and it works exceptionally well on speech.
  • Phase vocoder: operates in the frequency domain—the audio is first transformed into a spectrum, each frequency component's phase is handled separately, and then everything is reassembled. More stable on music (especially harmonically rich content), but prone to a "swimmy" or "hollow" character when pushed.
  • Transient preservation: drums and percussion are the natural enemy of speed-change algorithms—transients are so brief that chopping and respacing them blurs and softens their attack. Advanced algorithms detect transients and treat them specially.

The Boundaries of Speed Factors

Time-stretching is not magic—the more extreme the factor, the more audible the artifacts:

Speed rangeSpeech qualityMusic qualityTypical use
0.9x – 1.5xnearly transparentnearly transparentfine-tuning audiobooks and lectures
0.75x – 2.0xgoodacceptablelanguage learning, fast podcast listening
0.5x – 2.5xclearly processednoticeable degradationtranscription, extreme speed listening
< 0.5x or > 2.5xsevere artifactsnot recommendedspecial effects

Speech is the easiest content to handle (a single source with a relatively simple spectrum); multi-instrument music with strong rhythms is far harder. At the same 2x speed, a podcast merely sounds like "fast talking," while a symphony may already have smeared into mush.

Pitch-Shifting: The Inverse of Time-Stretching

The counterpart of time-stretching is pitch-shifting: changing the pitch while keeping the duration intact. Interestingly, the two are technically two sides of the same coin—pitch-shifting can be achieved by combining "time-stretch + resample":

Raise pitch without changing speed = resample up (which also speeds it up)
                                   → then time-stretch back to the original duration (keeping the new pitch)

Pitch shift is usually expressed in semitones: +1 semitone multiplies the frequency by about 1.0595, and +12 semitones is one octave up. Finer adjustments use cents, where 1 semitone = 100 cents—under ideal conditions, the human ear can distinguish differences of roughly 5–6 cents.

Typical applications:

  • Key-changing accompaniment: raise a backing track by 2 semitones to fit a singer's range;
  • Sound design: pitch voices down for weight, up for cartoonish characters;
  • Pitch correction: nudge a vocal by a few dozen cents to bring an off-key performance back in line (Auto-Tune used gently).

Formants: The Real Culprit Behind the "Chipmunk"

If you pitch a vocal up by 7 semitones, what you get isn't just "the same voice, higher"—it's a thin, cartoonish squeak, even when the time-stretching algorithm is flawless. The reason is formants.

A human voice is shaped by two components working together:

  1. Vocal cord vibration: produces the fundamental frequency, which determines pitch;
  2. Vocal tract resonance: the oral and nasal cavities form a resonating chamber that amplifies specific frequency regions—those boosted regions are the formants. Their distribution defines timbre: whether a voice sounds like an adult or a child, whether a vowel sounds like "ah" or "ee."

A straightforward pitch shift moves the fundamental and the formants together: shifting up 7 semitones also pushes the vocal-tract resonance regions higher, effectively swapping an "adult resonance chamber" for a "child-sized" one—hence the chipmunk. Shift far down instead, and you get the deep "monster voice."

Professional pitch-shifting tools separate and preserve the formants: they move only the fundamental (changing pitch) while pulling the formants back to their original positions (keeping timbre). That's the technical dividing line between "pitch change with timbre preserved" and "pitch change with timbre shifted."

The Full Map: Four Combinations

Combining "change speed or not" with "change pitch or not" covers every audio speed/pitch scenario:

GoalTechniqueResultTypical use
Speed + pitch changeresamplingfast & shrill / slow & deepcreative effects, tape-style speedups
Speed change, pitch preservedtime-stretchingtempo changes, voice staysaudiobooks, lectures, transcription
Pitch change, speed preservedpitch-shifting (stretch + resample combo)key changes, rhythm intactkey-changing backing tracks, tuning
Pitch + timbre preserved, speed intactformant-preserving shiftnatural-sounding transpositionprofessional vocal processing

Common Misconceptions

  1. "Speed change without pitch shift is the higher-quality, lossless option." Quite the opposite—resampling is nearly lossless (only the pitch changes), while time-stretching introduces processing artifacts. Which route to choose depends on whether you want the pitch to change, not on which is "more advanced."
  2. "Listening at 2x saves half the time with no comprehension loss." Most people comprehend fine at 1.5–2x, but beyond 2.5x, information retention drops noticeably—and algorithm artifacts start competing for your attention.
  3. "Slowing down music for practice costs nothing in quality." Below 0.5x, transient smearing and harmonic drift become clearly audible; for transcribing harmony, staying at or above 0.6x is advisable.
  4. "Pitch-shifting just makes a voice squeakier or deeper." That's actually the sound of formants shifting along with the pitch; formant-preserving shift changes only the note, not the character of the voice.

Practical Tips

  • Podcasts and audiobooks: start at 1.25x and adapt gradually; 1.5x is the sweet spot between efficiency and comprehension. Always use the "preserve pitch" mode.
  • Language learning: 0.75–0.9x slow playback for shadowing works better than an extreme 0.5x—less distortion, more natural intonation.
  • Music practice: use time-stretch mode for slow transcription; use resampling when practicing rhythmic feel (hearing pitch rise with speed actually reinforces the tempo-pitch relationship).
  • Creative effects: for tape-fast-forward or vintage cartoon sounds, use resampling directly—a "physically correct" time-stretch would defeat the purpose.
  • Check the hardest passages: after processing, listen specifically to percussion-heavy sections and rapid speech—that's where artifacts show up first.

Further Reading

This site's audio speed tool offers both "preserve pitch" and "pitch follows speed" modes—and after reading this article, you know they correspond to the time-stretching and resampling routes respectively, and which scenario calls for which.

Related Tools