Abstract
While genomic research has provided profound insights into human biology, it is estimated that a significant portion of chronic disease risk is driven by the exposome, defined as the cumulative measure of environmental influences throughout the lifespan. The exposome encompasses a vast array of non-genetic factors, including pollutants, dietary patterns, and socio-economic variables. Unraveling the impact of these factors requires a longitudinal approach, accounting for how exposures vary and accumulate over time. From a statistical perspective, this presents a significant challenge: researchers must analyze complex, highly correlated data while maintaining the ability to distinguish between population-wide trends and individual variability. The statistical cornerstone of such analysis is the Linear Mixed Model (LMM). The objective of this thesis is to extend the limits of these models and provide novel extensions of these methodologies for the complex datasets generated in exposome research.
The first part of this thesis focuses on understanding the theoretical limits of LMMs and developing solutions to overcome those limits by stabilizing the estimation process through regularization. In Chapter 3, we study the conditions under which an LMM is ``identifiable''—that is, ensuring that no two distinct sets of parameters generate the same model. Specifically, this research establishes new necessary and sufficient conditions for the identifiability of the random effect covariance matrix. We demonstrate that standard computational tools, such as the widely-used lme4 package in R, employ overly restrictive criteria. By relaxing these constraints, this work allows researchers to utilize more flexible model structures that were previously thought to be computationally or mathematically unreachable.
Chapter 2 addresses the challenges of including interaction terms. Indeed including interactions causes the complexity of the model to grow quadratically, often leading to statistical instability. We developed a Bayesian framework that enforces a structural hierarchy: an interaction is only considered relevant if its underlying main effects are also significant. This allows for the analysis of complex synergies without compromising the model’s reliability. Furthermore, Chapter 4 proposes a data-driven regularization of both the fixed and random effects of an LMM. While traditional methods often treat individual variation (random effects) as a secondary concern, we show that neglecting their precise estimation compromises uncertainty quantification and predictive power. To solve this without requiring manual tuning, we introduced an empirical Bayes procedure that estimates optimal regularization parameters directly from the data using a Laplace approximation to ensure computational feasibility.
The final stage of this thesis shifts toward the development of specialized tools designed for the structure of environmental data, where exposures often influence health in complex patterns. Chapter 5 introduces the ProfileLMM method, which combines the strengths of profile regression and LMMs. Because individuals are exposed to a large combination of environmental factors, we integrate a clustering technique to identify ``exposure profiles'' and link them directly to health outcomes over time. We applied this method to the textit{Lifelines} cohort, successfully disentangling the effects of complex air pollutant mixtures on blood pressure.
Finally, to ensure these advancements are accessible to the broader scientific community, Chapter 6 describes the R package resulting from this methodology. Released on the Comprehensive R Archive Network (CRAN), this software provides a streamlined, computationally efficient interface for performing profile regression combined with GLMM in large-scale datasets. The package generalizes the approach to support both binary and continuous outcomes and introduces novel applications, such as non-parametric piecewise fits for complex response variables. Collectively, these contributions provide a more robust mathematical foundation for turning complex environmental data into actionable public health insights.