Abstract
Metagenomics has transformed our ability to characterize microbial communities without relying on cultivation, but translating complex microbiome datasets into reliable biological conclusions remains a major analytical challenge. Although AI agents have improved substantially at software engineering and general data analysis, it remains unclear whether they can reliably analyze and interpret real-world metagenomic data. We introduce MetagenomicsBench, a benchmark of 100 verifiable evaluations derived from published datasets spanning community structure, host-microbiome associations, microbial function, longitudinal dynamics, and microbiome interventions. Each evaluation provides experimental data and a deterministic grader that evaluates recovery of a key analytical or biological result. Across 10,200 trajectories from 34 model-harness configurations, the strongest configuration achieved a 60.0% pass rate. Performance varied across functional domains, data modalities, and biological systems, while greater cost, token usage, and tool use did not consistently correspond to higher accuracy. Trajectory analysis showed that agents often produced technically plausible analyses while making mistakes in problem interpretation, statistical reasoning, and biological interpretation. MetagenomicsBench provides a framework for measuring progress toward AI agents that can reliably make and validate the analytical and biological decisions required for metagenomic research.