Signal and Sensation

Two Numbers Make a Vowel

A vowel is two bumps in a spectrum. A voice can only send them at multiples of its own pitch, and a high voice cannot send enough.

Open fullscreen →

What it is

Eight vowels, each defined entirely by two numbers. Click one on the left to hear it. It’s a buzz through two resonators, with no recording, no formant tracking and no vocal folds.

The right panel shows the vowel as a shape in blue, and what a voice at the current pitch actually transmits in amber. Drag the slider to raise the pitch.

The red circle marks where the formants come back when you measure the transmitted signal. Raise the pitch and watch it walk away from the vowel you asked for.

How it works

Two two-pole resonators at F1 and F2, in parallel, fed by an impulse train. That’s the whole synthesiser. The resonator is three coefficients, normalised so its peak gain is exactly 1, which the tests check.

Formant recovery is a peak search over a log frequency grid using Goertzel, done in two separate ranges so the second formant can’t be mistaken for another look at the first.

What surprised me

My first version failed its own test and taught me the day’s actual subject.

I’d synthesised each vowel at 120 Hz and asked the peak finder to recover the formants. It couldn’t. For “heed”, F1 is 270 Hz, and a 120 Hz voice has harmonics at 240 and 360. There is no energy at 270 Hz at all. The peak finder returned 240, because 240 was the nearest thing that existed.

The formants aren’t in the signal. They’re in the envelope, and a voiced source samples that envelope only at multiples of its own pitch. That’s measurable:

source mean formant error vowels identified harmonics below 3 kHz
noise (continuous) 36 cents 8 of 8 n/a
voice at 90 Hz 59 cents 8 of 8 33
voice at 130 Hz 85 cents 8 of 8 23
voice at 200 Hz 195 cents 6 of 8 15
voice at 400 Hz 442 cents 2 of 8 7
voice at 600 Hz 373 cents 2 of 8 5

With noise excitation, which samples the envelope everywhere, all eight vowels come back within 36 cents. At 400 Hz, two do.

So intelligibility isn’t a fixed property of a voice, it’s a sampling rate, and the thing being sampled is a spectrum at intervals of the fundamental. A bass gets 33 samples of his vowel shape below 3 kHz. A soprano at the top of her range gets five, and five points can’t locate two peaks.

I’d always heard that sopranos’ vowels are unintelligible in their upper register and filed it as a fact about singing technique. It isn’t. It’s Nyquist applied along the frequency axis instead of the time axis, and the numbers fall out of a resonator and a peak finder in an afternoon.

What I would do next

Add vibrato. A few percent of pitch wobble slides the whole comb back and forth across the envelope, so a listener integrating over time should recover far more of it than any single instant contains. That would also explain why vibrato isn’t merely decorative.