Imagine identifying a face as happy or sad and male or female at the same time. GRT says each face lands somewhere in a two-dimensional "perceptual space" — a cloud, not a point, because perception is noisy. You respond by asking which side of each line the cloud landed on.
Two things can go wrong with independence. The clouds can move: if a face's happiness changes depending on whether it's male, the dimensions aren't separable. The clouds can tilt: if noisy-happy and noisy-male travel together on a given trial, the dimensions aren't independent. Drag the sliders and watch both happen.
Each ellipse is one stimulus: a 1-SD contour of where it lands in perceptual space on a given trial. The dashed cross is the decision bound — you respond by asking which quadrant you fell into.
Read the marginals. Each dimension's two curves of the same colour are the same level of that dimension, paired with the two different levels of the other one (solid vs dotted). If they sit on top of each other, that dimension is separable. If they pull apart, it isn't.
Exact response probabilities implied by the space above — this is what you'd get with infinitely many trials.
This is the whole problem
The confusion matrix has 16 cells, but each row must sum to 1 — so it carries exactly 12 free numbers. And the space above has exactly 12 identified parameters. That is not a coincidence: it's why a single confusion matrix is enough to pin down the space, and why one extra parameter would make it impossible.
Solid = the space you built. Dashed rose = what was recovered from the noisy data alone. The gap between them is the error — and it is supposed to shrink as you raise the trial count.
What just happened?
We drew 100 trials per stimulus from your space, producing a noisy confusion matrix — a simulated experiment. Then we handed only that matrix to the model, which had never seen your sliders, and asked it to reconstruct the space.
Try this: set the trial count to 10 and run it a few times. The estimates jump around — and, crucially, the intervals get wide. That is the whole point. An honest method doesn't just get less accurate with less data; it tells you it has got less accurate.