Abstract
Deep neural networks are increasingly used as models of vision, yet they are evaluated against the average observer: average performance, average neural response, average representational geometry. This discards one of the richest constraints on theory: stable individual differences. Fitting the average is necessary but not sufficient. I argue that future models should be judged not only by their fit to mean human performance, but by whether populations of artificial observers reproduce the covariance structure and latent factors found across human observers. Two mature but disconnected literatures motivate this. Decades of factor-analytic psychophysics and electrophysiology show that interindividual variation is signal, not noise: the pattern of correlations among observers reveals the number and tuning of underlying mechanisms, and fractionates vision into many narrow factors rather than one general ability. Meanwhile, neural networks differing only in random initialization develop measurably different internal representations, so any single trained network is one sample from a distribution of possible observers. Joining these yields a computational psychometrics of vision: construct a designed population of artificial observers (varying in architecture, initialization, training diet, and front-end sampling), submit them to the same latent-variable analyses used on humans, and test whether the recovered factor structure matches. Recent work showing that seed-varied networks align with specific individuals is an existence proof. This article generalizes it into a framework, a staged roadmap, and a benchmark, together with the safeguards that keep such a benchmark from becoming an optimization target, for testing hypotheses about human visual variation with artificial observer populations.