Signal and Sensation

How Fine Is Your Ear

Two tones, almost the same. Twenty questions later the page knows your pitch threshold to within about ten percent.

Open fullscreen →

What it is

You hear two tones. Say which was higher. Get it right twice and they move closer together, get it wrong and they move apart.

After twenty or thirty answers the dashed line is your pitch discrimination threshold. For most people somewhere between 3 and 15 cents, which is a fraction of a percent.

This is the first day where the reader is the instrument.

How it works

Testing at a fixed difference wastes nearly every trial. Too easy and you learn nothing, too hard and you learn nothing either. An adaptive staircase follows you instead, so almost all the questions land near your threshold where they’re actually informative.

The rule here is two-down-one-up: two correct answers make it harder, one wrong makes it easier. That converges on the difference you get right about 70.7% of the time, and the step size shrinks at every turning point so it settles rather than oscillating. The readout is the mean of the last six reversals, because the early ones are discarded while the staircase is still travelling rather than hovering.

The order of the two tones is randomised every trial, so you can’t learn “it’s always the second one”, and each tone gets a 20 ms fade. Without the fade the click at the edges is a louder cue than the pitch difference and you end up measuring click discrimination.

What surprised me

I couldn’t test this on myself, so I built a synthetic listener with a known threshold and ran the staircase against that. It immediately showed a bias I’d never have noticed otherwise.

Given a listener whose threshold is 2%, the staircase reports 1.54%. Given 0.6%, it reports 0.46%. Given 0.2%, it reports 0.15%. Every time, a factor of 0.76, across three orders of magnitude, with about 13% variation between runs.

That isn’t an error, it’s a definition mismatch, and it’s entirely predictable. My simulated listener’s “threshold” is the difference they get right 82% of the time. A two-down-one-up staircase converges on the 70.7% point, which is a smaller difference. The ratio between those two points on a psychometric curve is fixed, so the bias is a constant. Theory puts it near 0.66 and the measured 0.76 sits a little above, the rest being the finite step size and the averaging.

The useful part is that the bias is constant. An instrument with a stable, known offset is perfectly good. An instrument whose offset varies with what it’s measuring is not, and finding out which one you have means testing it against something whose answer you already know.

What I would do next

Run two staircases interleaved so a listener can’t tell which trial belongs to which, then check the two agree.