Signal and Sensation

Hidden In Plain Sound

A loud tone does not make its quiet neighbour quieter. It makes it not exist. This is why MP3 works.

Open fullscreen →

What it is

A steady 1 kHz tone, and a second tone you drag around. The green region is where the second tone is predicted to be inaudible. Drag the probe into it and it vanishes, then turn the masker off and it’s plainly, obviously there, exactly as loud as it was a second ago.

That gap between physically present and perceptually present is the entire budget perceptual audio coding spends.

How it works

Loud sound saturates the region of the cochlea it excites, and a quieter tone landing in that region produces no additional response. It isn’t attenuated. It’s absent.

The model here is the standard two-slope spread of masking on the Bark scale, which is frequency measured in critical bands rather than Hz. Masking spreads at about 27 dB per Bark downward in frequency and only about 12 dB per Bark upward, so the green region is a lopsided triangle. A masker reaches much further up than down.

Concretely, with a 1 kHz masker at 80 dB, a probe needs to exceed 56 dB at 1250 Hz to be heard but only 37 dB at 800 Hz. So a 45 dB probe is inaudible above the masker and audible below it, which you can check by dragging.

What surprised me

I wrote a test asserting the Bark scale compresses high frequencies, comparing the octave 500→1000 Hz against 5000→10000 Hz. It failed. Those octaves span 3.77 and 3.89 Barks. Nearly identical, and the top one marginally wider.

The compression is real but it isn’t about octaves. Measured properly, a fixed 100 Hz span covers 0.96 Bark at 200 Hz and 0.069 Bark at 8 kHz, a factor of fourteen.

Through the mid range every octave is about four Barks: measured at 3.77, 4.59, 4.16 and 3.89 for octaves starting at 500 Hz, 1 kHz, 2 kHz and 5 kHz, so within a quarter of each other. The scale is close to logarithmic there. It’s only in fixed Hz that the squeeze appears.

That distinction matters for the thing the day is about. A codec allocates bits per critical band, so at 8 kHz one band covers 1.7 kHz of spectrum and everything inside it competes for the same handful of bits. Not because high frequencies are less important, but because the ear can’t separate them, and the octave-based intuition I started with hides exactly that.

What I would do next

Temporal masking, where a loud sound hides a quiet one arriving up to 200 ms after it and, stranger, up to 20 ms before.