Abstract
Large language model benchmarks often contain many more items than evaluated systems. We develop a transposed Bock–Aitkin EM algorithm for 1PL and 2PL random-item models, including correlated difficulty and log discrimination. Capabilities are structural parameters; item effects are integrated over a normal population. We derive posterior-averaged scoring equations and distinguish exact marginal-likelihood theory from the implemented Laplace–cubature moment iteration, which can stabilize without solving the likelihood score. A simulation comparison uses 100 replications in each of 16 conditions. Capability recovery improves with more items for the transposed procedures, but population-correlation bias persists in small model cohorts. All variational fits in this comparison reached their iteration cap, limiting cross-framework conclusions. In an application to 98,032 WILD items, numerically eligible correlated transposed 2PL and variational 2PL have similar predictive performance, with log loss and Brier score favoring different procedures. JML 2PL has the lowest test Brier score but does not meet its numerical eligibility rule. The simulation and application use different fitting implementations and baseline settings, so they do not jointly validate a single numerical procedure. These exploratory findings concern point estimation and random-cell prediction, not calibrated parameter uncertainty or generalization to new models.