Hearing a moment
Press the waveform and hold still. What should one instant of a sound sound like? Eight answers, each live, timed and measured.
Drag across the waveform to scrub; hold still to hear the moment. The lighter band is what the method listens to. Frame, overlap and spread shape the spectral methods (random phase and the vocoders); bank, read and lines the noiscillators. With the waveform focused: ← → move the caret (Shift: ten times further), Space keeps playing, 1–8 pick a method. Drop a sound file anywhere on the page.
Cost and fidelity
Every method plays one gesture at the caret: a 1.5 s hold, a 1 s drag at half speed, a 0.5 s hold. Cost is the time to render a block of 128 samples, the fastest of four turns in a worker; the budget at 44.1 kHz is 2,902 µs. Fidelity compares the first hold with the sound around the caret.
| Method | µs per block | Real time | Level | Distance | Flutter | Repeats | Listen |
|---|
- Levelthe hold's power against the source's
- Distancethe mean dB difference per 21.5 Hz bin, 50 Hz–16 kHz, levels aligned. A frozen noise differs from any one moment of noise, so noise never scores 0.
- Flutterhow much the hold's loudness moves: the spread of its 10 ms RMS, dB
- Repeatshow periodically it moves: the highest correlation of its loudness with itself 20 to 400 ms later. A frame played again and again scores near 1; so does a sound that beats, as the chime does.
The noisc bank costs ns per noiscillator per sample: one core plays about of them in real time. Random phase is a bank too, a noise line on each of the frame's bins, and costs µs a block against the bank's : an inverse FFT plays every line at once. A GPU could run the oscillators in parallel, but an AudioWorklet cannot reach WebGPU, and the FFT already makes a dense bank cheap.
The methods
- TapeThe read head chases the caret, so pitch follows speed and a still caret is silence. Audacity's Scrub.1
- LoopThe 80 ms around the caret, again and again, each pass from where the caret is. It buzzes at its own rate, and pitch holds only to the nearest of its harmonics, 13.9 Hz apart. Reaper's looped-segment scrub.2
- Grains80 ms grains, a new one every 20 ms, each from the caret ± 15 ms.3 Steady, chorused, a tone blurred to the grain's resolution.
- Noisc bankNoiscillators4 on a fixed log grid, one per band, set every hop to the power the caret's spectrum holds there. Nothing is fitted. Read from the bins, it is a noise vocoder:5 a partial's leakage becomes noise in every band it touches, and pitch survives only where bands are narrower than the harmonics' spacing. Reassigned,13 each bin's power moves to its instantaneous frequency first: a partial lands in one band, whose line sits at its frequency, a sine as far as it is one.
- Noisc linesTonal peaks become lines, as wide as each peak is wider than a steady sine's;8,9 third-octave noise bands carry the rest. Sines plus noise,6,7 drawn with noiscillators. A line is a noise band (its envelope wanders), an FM line (its level steady, its frequency jittering, a Lorentzian line), a drifting line (its level steady, its frequency wandering off and back, a nearly Gaussian line)14 or a sine and noise in the measured share.
- Random phaseThe caret's spectrum with new random phases every hop: Paulstretch.10 Noise stays noise; a tone widens to a band about two bins wide, centred on it, and each hold peaks somewhere within half a bin: with 2,048-sample frames, a held 220 Hz sine up to 83 cents off.
- VocoderEach peak's phase advances as it did between two frames a hop apart; the bins around it keep their offsets to it.11 Tones hold exactly; noise freezes into a metallic ring, and a held frame plays its own shape again every hop. More overlap only repeats it faster; a spread overlap-adds frames from around the caret instead, and the repeating goes.
- Vocoder + noiseThe vocoder for the bins a tonal peak's own leakage explains, random phase for the rest. Melodyne's engine has the same parts: frequencies from phase differences, a harmonic model, a residual.12
- Audacity Manual, Scrubbing and Seeking: Scrub plays at the speed of the mouse.
- REAPER, looped-segment scrub at the edit cursor (Preferences › Audio › Playback); Scrub and jog, The REAPER Blog (2017).
- C. Roads, Microsound, MIT Press (2001).
- D. Iv, noiscillators: oscillation with uncertain frequency, a line of any width from a sine to noise, as a noise band (quadrature), a frequency walk (Wiener, or Ornstein–Uhlenbeck) or a sine and noise (rice).
- R. V. Shannon, F.-G. Zeng, V. Kamath, J. Wygonski & M. Ekelid, “Speech recognition with primarily temporal cues,” Science 270 (1995).
- R. J. McAulay & T. F. Quatieri, “Speech analysis/synthesis based on a sinusoidal representation,” IEEE Trans. ASSP 34 (1986).
- X. Serra & J. O. Smith, “Spectral modeling synthesis: a sound analysis/synthesis system based on a deterministic plus stochastic decomposition,” Computer Music Journal 14 (1990).
- J. O. Smith & X. Serra, “PARSHL: an analysis/synthesis program for non-harmonic sounds based on a sinusoidal representation,” ICMC (1987): a peak's frequency from the parabola through its log magnitudes.
- F. J. Harris, “On the use of windows for harmonic analysis with the discrete Fourier transform,” Proc. IEEE 66 (1978): Hann's half-power width, 1.44 bins.
- Nasca Octavian Paul, Paulstretch (2006).
- J. Laroche & M. Dolson, “Improved phase vocoder time-scale modification of audio,” IEEE Trans. Speech and Audio Processing 7 (1999).
- P. Neubäcker (Celemony), US 8,022,286 B2, “Sound-object oriented analysis and note-object oriented processing of polyphonic sound recordings” (2011).
- K. Kodera, R. Gendrin & C. de Villedary, “Analysis of time-varying signals with small BT values,” IEEE Trans. ASSP 26 (1978): reassignment.
- P. W. Anderson, “A mathematical model for the narrowing of spectral lines by exchange or motion,” J. Phys. Soc. Japan 9 (1954); R. Kubo, “Note on the stochastic theory of resonance absorption,” J. Phys. Soc. Japan 9 (1954): a line whose frequency walks slowly is its frequencies' Gaussian; walking fast, it narrows to a Lorentzian.