Signal and Sensation

How Many Voices Can You Follow

Two voices in different places are far easier to separate than two in the same place. The help has a size in decibels, and at this pitch almost none of it comes from your head being in the way.

Open fullscreen →

What it is

Headphones. A single note plays in the middle so you know what to listen for, then the same tune arrives buried in other tunes, and you say whether it went up or down.

Get two right and the target gets quieter. Get one wrong and it gets louder. That’s a staircase, and where it settles is how quiet the target can be before you lose it.

The page runs two of those side by side. In one, every voice comes from the same place. In the other they’re spread across 120 degrees around your head. The difference between the two thresholds is spatial release from masking, and it’s why you can hold a conversation in a pub and why a mono recording of the same room is unlistenable.

The buttons change how many voices are against you.

How it works

The spatialiser is day 58’s, brought across by copy rather than imported, so this day stands on its own the way the other fifty-nine do. Woodworth’s time difference and a head shadow shaped in ka, with the tests here re-asserting the properties this day leans on rather than trusting they travelled.

Two decisions shape the whole measurement.

The maskers keep constant total power as voices are added. Four voices each get half the amplitude of one, so going from one competitor to four doesn’t quietly turn into a level change. Whatever difficulty is left when energy is held constant is informational rather than energetic, which is the more interesting half.

The target sits at 420 Hz, deliberately. At 420 Hz the head casts a shadow of about 1.2 dB, which is nearly nothing. That matters because there are two separate reasons separating voices helps, and this pitch rules one of them out.

The first reason isn’t clever at all. Spread the sources and one ear ends up with a better view of the target than the other, purely because the head is in the way of different things. That’s the better ear advantage and it’s arithmetic, so the page computes it: the best target-to-masker ratio available at either ear, against the ratio you get with everything piled up. Maskers add as power, because independent sources do.

At 420 Hz with four voices across 120 degrees, that comes to 0.7 dB. So if the measured release is much larger, the head can’t be the explanation, and what’s left is two ears doing something neither could do alone.

That’s the dashed line on the chart. Anything between it and the green line is binaural.

Now the part that surprised me, which is about the measuring rather than the hearing.

What surprised me

A short session can’t measure this, and I nearly shipped one that pretended to. I started with six reversals per staircase, which is a normal number and takes about twenty trials. Then I ran the estimator against a simulated listener four thousand times to see what it actually recovers:

reversals averaged over bias standard deviation trials
6 last 4 -1.55 dB 3.27 dB 21
8 last 6 -1.49 dB 2.97 dB 28
12 last 8 -1.34 dB 2.31 dB 41
16 last 12 -1.25 dB 1.95 dB 53
24 last 16 -1.06 dB 1.39 dB 79

3.27 dB of noise on one threshold means 4.6 dB on the difference of two, and the effect I am trying to measure is around 5 to 8 dB. A twenty trial session would have returned a number with an error bar bigger than the thing it was measuring, printed to one decimal place, and looked entirely convincing.

The default is now twelve reversals, which costs 41 trials a cell and gets the difference down to about 3.3 dB. Still not tight. Filling the whole chart honestly is around 330 trials and almost nobody will do that, so a single cell with a couple of decibels of slop is what this page realistically gives you.

The bias is the good news. The staircase settles about 1.3 dB below the true threshold, every time, in the same direction. That’s a real bias and it would matter if either number were the answer. Neither is. The answer is the difference, both halves carry the same bias, and it cancels exactly.

So the quantity this page reports is better than either of the measurements it’s made from, which isn’t something I expected to be able to say.

The distractors were sharing notes with the target. A three note tune spanning six semitones an octave below the target finishes exactly where the target starts. That shared note is a free anchor with nothing to do with masking, and it was there in one span out of four, in one direction out of two. The generator now checks every candidate against every target note and moves register until it’s clear, and the test walks all four spans in both directions because the collision only appeared in one corner.

What I would do next

Run it at 3 kHz. The head shadow at 420 Hz is 1.2 dB and at 3 kHz it’s 11, so the better ear advantage stops being a rounding error and becomes most of the effect. The same page with a different target pitch should show the dashed line climbing to meet the green one, and the binaural gap closing.

That’s a real prediction and the code to test it is one constant. I left it at 420 Hz because ruling the head out entirely makes a cleaner first measurement, but the pair of pitches is the actual experiment.

The other thing I’d want is the same effect with the voices behind you. Front-back mirror pairs produce identical cues, as day 58 measured to fifteen decimal places, so a masker directly behind should mask exactly as well as one directly in front. If it doesn’t, that difference is spectral information from the outer ear, which neither of these two days has.