Signal and Sensation

Which Is Louder

Two clips with identical peak levels, and one is obviously louder. Your answers reveal which of three measures you are actually following, and it is not the one on the meter.

Open fullscreen →

What it is

Headphones. Two clips play one after the other and you say which was louder. That’s the whole game.

What the page is doing is holding one of three measures deliberately equal and letting the other two disagree. Sometimes the two clips have exactly the same peak level. Sometimes the same average power. Sometimes the same loudness. You’re never told which.

Your answers then say which measure you were following, because only one of them can explain a run of choices. The bars on the right fill in as you play, and fifty per cent is what following nothing at all looks like.

How it works

Three numbers, all claiming to be how loud something is.

Peak is the tallest sample. It’s what a clipping indicator watches and it has almost nothing to do with how loud anything sounds.

RMS is the average power. Closer, and still wrong, because it treats 40 Hz and 3 kHz as equals and your ear doesn’t.

Loudness is the broadcast standard’s version: filter the signal to roughly the shape of a head and an outer ear, then take the mean square of that. K-weighting, from BS.1770.

That filter is two biquads, a high shelf lifting everything above 1.7 kHz by 4 dB and a high pass throwing away everything below 38 Hz. I designed them from the usual parameters first, which was a mistake worth keeping: the design comes out half a decibel high through the shelf transition around 1.6 kHz, which is squarely in the region that decides how loud things sound. So the standard’s own coefficients are used at 48 kHz, the designed ones are a documented fallback for other rates, and there’s a test asserting they differ by between 0.2 and 0.6 dB so nobody discovers that by accident.

Everything on the page is measured at 48 kHz for that reason.

One detail fell out that I like. K-weighting lifts 1 kHz by 0.698 dB and the standard’s channel offset is -0.691 dB. They cancel to seven thousandths of a decibel, which is exactly why a 1 kHz tone reads its own RMS level in LUFS. That only comes out right if the filters, the integration and the offset are all correct, so it’s the best single check in the file.

The pairs are made by taking material, running one copy through a crude limiter, then scaling both so the chosen measure matches. Every relationship is measured off the finished arrays rather than assumed, because a limiter and a normaliser interact and half a decibel of intended equality is easy to lose without noticing.

At equal peak, the limited copy of a drum pattern comes out 7.8 dB louder in RMS and 7.5 dB louder in loudness. Its crest factor drops from 21.4 dB to 13.4. That’s the loudness war in two numbers.

What surprised me

The scoring was punishing a measure for the pairs it had nothing to say about. I ran a simulated listener that follows loudness perfectly and it scored 63 per cent on loudness. It should score 100.

The reason is that holding RMS equal usually holds loudness equal too, since they differ only by a spectral weighting and a limiter barely changes the spectrum. Those pairs have no loudness opinion at all, and I was counting them as loudness misses. A perfect loudness-follower was being dragged towards the middle by trials where loudness had abstained.

Fixed by recording, per pair, which measures had a view at all and dividing only by those. The same listener now scores 100 per cent of 20 pairs on loudness, 50 per cent of 32 on peak, and that 50 is exactly the chance it should be.

RMS scores 90 per cent, and that’s the honest problem with this page. A loudness-following listener agrees with RMS nine times out of ten, because the two measures usually point the same way. Only bass-heavy material separates them, since that’s where the weighting actually bites.

So the page can tell you clearly that you’re not following peak. It can tell you that you are following something like power. It can’t cleanly separate power from loudness without many more trials on the one kind of material where they disagree, and the bar chart doesn’t say so.

Matching drums by RMS pushed them 3 dB over full scale. A drum pattern has a crest factor over 20 dB, so normalising it to -18 dB RMS puts the peaks at +3 dBFS and the output clips.

My test for that only checked one material, noise, which has a crest factor of 4.8 dB and is never in danger. The fix is a trim applied equally to both clips so every difference between them survives untouched, and the test now walks all four materials against all three held measures, which is what it should have done first.

That’s six days running where the first thing to break was the instrument.

What I would do next

Add material designed to split RMS from loudness rather than hoping it turns up. A pair where one clip is a low rumble and the other a mid-range buzz, matched on RMS, has loudness saying one thing and power saying nothing at all. Half a dozen of those in the rotation would turn the 90 per cent into a real separation instead of a coincidence.

The other thing missing is the gating. The real standard doesn’t take the mean square of the whole file, it measures in 400 ms blocks and throws away the quiet ones, which is what stops a track with long silences reading as quieter than it is. Everything here is short and continuous so gating changes nothing, and that will stop being true the moment the material gets any more realistic.