Abstract
Grand canonical Monte Carlo (GCMC) simulations underpin machine-learning-accelerated screening of metal–organic frameworks (MOFs) for hydrogen storage, yet published workflows differ widely in simulation rigour. Minimum-image violation in GCMC screening of MOFs biases uptake but preserves rankings, and the error is predictable before any simulation is run. Here, the cost of the single-unit-cell shortcut (one crystallographic cell regardless of the interaction cutoff, violating the minimum-image convention) is quantified across a screening library for the first time. Uptakes of 85 CoRE MOF 2019 frameworks were computed with a converged per-framework supercell protocol and repeated in a single cell. Deliverablecapacity rankings are found to be remarkably stable (Spearman ρ = 0.996, bootstrap CI [0.989, 0.998]; top ten recovered exactly, no capacity forfeited), while absolute uptakes fall by up to 30%. The total replication factor m predicts the error poorly: frameworks needing only 4 × 1 × 1 replication lose as much uptake as 36× cases. The error tracks the worst single axis nmax, exposing rod secondary-building-unit frameworks (MOF-74 among them) as the class at risk; intermediate-supercell ladders confirm this directly, the same volume increase recovering 2.4% or 16.4% of the deficit according to which axis receives it. Stated without reference to the cutoff, a single cell is safe for ranking whenever every perpendicular width exceeds one cutoff radius, while absolute capacities require two. Recomputing only the frameworks with nmax ≥ 3, one in three of the set, brings the largest remaining error down to 1.3% while still saving 18% of the campaign