Lab 01 · Typesetting speech

How do you set words under a waveform?

A waveform is time drawn along a line. A transcript is text, and text keeps its own pace. To edit speech by its words the two have to agree. This is how we got them to, and what typesetting and music engraving already knew about it.

September 202614 min readFrom waveareaFigures run on a NASA podcast

01Two clocks

Open a podcast in an audio editor and you get its waveform: bars running left to right, each one a few hundredths of a second of sound. Open its transcript and you get words. Both describe the same two minutes, and they disagree about how long two minutes is.

The waveform keeps the sound's clock. At 160 pixels a second, a word said in a third of a second gets 53 pixels, however many letters it has. The text keeps the reader's clock: at this size the is 20 pixels wide and astronauts 65, however fast they were said.

Set every word under the moment it starts and both clocks show at once. Where the speaker hurries, words pile onto each other. Where they slow down, words drift apart and hang in the air.

FIG. 1.1Seven seconds of the podcast, each word set where its sound starts. Drag the scale: narrow, the words pile up; wide, they drift apart.

This matters for wavearea, the editor we are building for long speech: podcasts, lectures, interviews. It has to show both clocks at once. You read the words to find a place and cut where the waveform shows a pause. So the question here is small and stubborn: where should each word sit?

02Three ways that fail

We tried the obvious answers first, on two public-domain recordings that come with published transcripts: an episode of NASA's Houston We Have a Podcast and a LibriVox reading of O. Henry's The Gift of the Magi. Each answer gives one of the clocks away.

Squeeze the words. Keep each word under its own sound and condense its letters until it fits. At our editor's default zoom, 97 words in 100 needed squeezing and 85 still ended up away from their sound. Text whose width changes from word to word stops being text.

Make room in the waveform. HTML has a standard for small text attached to other text: ruby, used for pronunciation guides over Japanese and Chinese characters. Make each word's bars the base and the word its annotation, and the browser widens the base wherever the word is longer than its sound. The words read perfectly. The waveform breaks into islands, and a pixel no longer stands for the same span of time.

Put the text first. Set the transcript as a document and hang each word's bars under it, drawn as wide as the word. It reads like a book, but the waveform under it keeps neither the sound's timing nor its shape: every word gets as much waveform as it has letters.

FIG. 2.1The same ten seconds three ways. Squeezed: the time is kept and the text is lost. Room made: the text is kept and the time breaks. Text first: the waveform becomes decoration.

The clock that must not give is the sound's. An editor cuts time, so time on the screen has to be true.

03Time holds, text gives

So the waveform stays linear: one pixel always stands for the same span of sound. Whatever bends has to bend in the text, and text has more give than it seems. It has the spaces between words.

Typesetters have bent those spaces for five centuries; it is how justified text gets a straight right edge. Our edge is the sound itself. Before bending anything, though, it pays to choose a scale at which little bending is needed.

Measure the transcript's width and divide it by the time its words take to say, leaving the pauses out. The result is the scale at which text and speech keep the same pace, on average: for this podcast, in this article's type, about 125 pixels a second. Much wider, and most spaces need more stretch than looks right, so the words gather in the middle of their sound. Much narrower, and they crowd.

FIG. 3.1The first minute, laid out as this article goes on to describe. The scale starts at the fit. Wider, spaces run out of stretch and words drift from their sound; narrower, they crowd and push.

The fit scale also says how zoom should work. Audio editors usually zoom by samples per bar: in ours, 1,024 samples to a bar 1.25 pixels wide. Then the same zoom shows speech recorded at 44.1 kHz at 54 pixels a second and speech recorded at 24 kHz at 29. Zooming by pixels per second, starting from the fit, keeps the words readable whatever the file.

04Where the phrases are

Every word has a time

To put a word under its sound we need to know when it was said. A speech recognizer can guess both the words and their times. When the words are already known, as they are with a published transcript, forced alignment does better, because it only has to answer when.

We used wav2vec 2.0, a speech model trained with connectionist temporal classification (CTC). For every 20 milliseconds of audio it gives the probability of each letter, and of a blank that means no letter. Aligning is finding the most probable path through those probabilities that spells the transcript in order, one column of dynamic programming per 20 ms. Where the transcript leaves out something the speaker said, a catch-all state absorbs it, so it doesn't drag the words after it late.

CTC has one flaw worth knowing. It lets blanks sit between a word's letters, so a faint first letter can land seconds early, in the silence or music before the word. Where a word's letters fall more than half a second apart and the model hears no speech between them, we keep the part nearest the words around it.

FIG. 4.1Forced alignment on the opening of the podcast. Each bracket is one word's sound as the aligner found it, to the nearest 20 ms. Brackets touch where the speaker runs words together.

Pauses make phrases

People speak in runs, not words: stretches between breaths and hesitations. On a waveform they are clusters of bars with quiet between them. We call a stretch quiet when its peaks stay below a threshold set a quarter of the way from the recording's noise floor to its speech level, and a quiet stretch of 180 ms or more a pause.

The pauses mark the phrases the recording itself contains. In the narration, 84% of them fall at punctuation: the reader breathes where the author put a comma. Conversation is looser; one cluster in the podcast runs 78 words without a pause. A cluster too long for a line splits where it pauses most.

FIG. 4.2Fourteen seconds as peak levels in decibels, with the silence threshold drawn in blue. Shaded stretches stay under it long enough to count as pauses; the brackets are the clusters they leave. Drag the minimum pause to watch phrases merge and split.

05Boxes and glue

TeX, the typesetting system Donald Knuth wrote for his own books, sets a line from boxes and glue. Words are boxes of fixed width. The spaces between them are glue: each has a natural width and may stretch or shrink within limits, and a paragraph is set so that its glue bends as evenly as possible.

We keep the glue and change what pulls on it. In TeX the pull comes from the right margin. Here every word has its own target: the center of its sound. Words go where the sum of squared distances between word centers and sound centers is smallest, while every space stays within its limits: between ¾ and 2½ of a normal space inside a cluster, at least 2½ between clusters, with no upper limit there.

FIG. 5.1One line, as boxes and glue. Ticks mark the center of each word's sound; dashed lines show how far each word sits from its tick. Centering moves a cluster as one block. Least squares lets each space take up the tempo inside it.

Three behaviours fall out of that one rule, with no special cases. A cluster whose text is shorter than its sound spreads its spaces and centers on the sound. One whose text is longer overflows evenly on both sides. When two clusters collide, both move, the one with more words less. Words stay near their sound and phrases stay phrases.

06Where lines end

Nothing inside a waveform says where a line should end, so audio editors wrap it every so many seconds. Text does say. A line should end between phrases, the way a subtitle or a line of verse does. Ending lines in pauses also means a word's sound is never cut in two, so hyphens are never needed. A fixed grid, which cuts wherever the seconds run out, would need them.

Choosing the breaks is Knuth and Plass again. Rather than filling each line as far as it will go, their algorithm chooses all of a paragraph's breaks together, minimizing a cost: each line's unfilled share squared, plus a penalty for each break. Ours are nothing between clusters, 8 after punctuation and 40 inside a cluster. The time between two lines is shared: the boundary falls inside the pause, and the next line starts just before its first word. A pause longer than two lines can hold gets lines of its own, and a sound longer than a line runs on into lines without words, like the extender under a held note.

FIG. 6.1A paragraph three ways. Every few seconds: the grid cuts wherever the time runs out. Greedy: each line takes all it can and leaves the rest to chance. Chosen together: the breaks that cost least overall. Each line's end is labelled with its kind of break and its penalty.

07One selection

An editor like this gives you two surfaces to select on, the words and the bars, and it has to feel like one. We make the browser's own selection the only one. The view is editable, with typing blocked, so it has a real caret and a real selection. Wherever a selection lands, in words or in bars, it becomes a span of time, and that span is painted on both. A selection made in words snaps to whole words; a double click on bars selects the word spoken there.

The caret is a moment in time, so it shows twice: as the text caret where you clicked and as a thin mark on the bars. Playback moves the mark and underlines the word being said, so playing never looks like selecting.

FIG. 7.1Two minutes of the podcast. Click to place the caret, drag across words or bars to select, double-click on bars to select a word, press Space or Play to hear the selection, or from the caret on.

08Drawing the bars

The bars are text too. Wavefont is a font whose glyphs are bars: a character for each height, and accents that shift a bar up or down, one step or ten at a time. A block of samples becomes the character for the distance from its lowest level to its highest, shifted to their middle. Because they are text, the bars take a real caret and a native selection. Letter spacing makes the gap between them, and browsers space a character together with its accents, so every bar keeps the same pitch.

Bars here keep a fixed width and gap, as in a voice message. Zoom changes how many samples a bar covers, never the bar. One detail matters on high-density screens. A bar 1.25 CSS pixels wide is 2.5 device pixels, so every other edge between bars falls in the middle of a pixel. Each bar paints that pixel half dark, and two half-coverages blend to three quarters, not to one. A faint seam shows. Bars whose edges fall on whole device pixels stay solid.

FIG. 8.1Device pixels under touching bars, magnified. Numbers are the coverage of pixels that two bars share. Each bar covers half; blended, 1 − ½ × ½ = ¾, lighter than the bars around it.

Bars are one look. The other is a continuous outline, the minimum and maximum of every pixel column, which gl-waveform draws on the GPU. Laid under the same bars made transparent, it looks different and behaves the same: the text and the selection don't change.

FIG. 8.2The same lines drawn with wavefont bars and as a gl-waveform outline, with the bars kept, invisible, above it for selection.

09Fast enough

Laying out eight minutes of the podcast, 1,316 words, takes 15 to 20 milliseconds in Chrome, and 21 to 26 with the page built, because the browser lays out only the lines in view. Paragraphs don't affect one another, so an edit relays out only its own. The work is a few small sequential passes whose results the page needs right away, so it stays on the CPU. The GPU earns its place where the same work repeats for every sample: drawing, and recomputing bars at a new zoom across hours of audio.

This is the layout wavearea is moving to. Some parts are still open: the scale to use on a phone, where a line holds four words; how to show two speakers talking over each other; and how the text should behave while a phrase is regenerated in the speaker's own voice. The workshop behind these figures lives in the wavearea repository.

SOURCES
  1. NASA, Houston We Have a Podcast, episode 434: Artemis III Training. Audio and transcript, public domain. The figures use its first two minutes.
  2. LibriVox, The Gift of the Magi by O. Henry, read by Betsie Bush. Public domain.
  3. D. E. Knuth and M. F. Plass, Breaking paragraphs into lines, Software: Practice and Experience 11(11), 1981.
  4. A. Graves, S. Fernández, F. Gomez and J. Schmidhuber, Connectionist temporal classification, ICML 2006.
  5. A. Baevski, H. Zhou, A. Mohamed and M. Auli, wav2vec 2.0, NeurIPS 2020. Model: facebook/wav2vec2-base-960h.
  6. LilyPond, Common notation for vocal music: melismas, extenders, hyphens.
  7. W3C, CSS Ruby Annotation Layout Module Level 1.
  8. Wavefont 3.8.2, gl-waveform 5.
  9. Type: Manrope by Mikhail Sharanda; in the figures, Newsreader by Production Type and Departure Mono by Helena Zhang.