Abstract
Industrial data-center systems contain complex component dependencies and long fault-propagation chains, making accurate root-cause localization difficult. Large language models (LLMs) provide a promising solution because of their strong ability to understand, organize, and reason over heterogeneous operational evidence. However, existing LLM-based methods mainly rely on statistical co-occurrence in pretraining corpora and may mistake correlation for causation in complex systems. In this paper, we propose CALM, a causally constrained LLM framework for fault diagnosis. The main idea is to learn the causal structure among fault units induced by system dependencies and use it to constrain the LLM reasoning process. Specifically, we train a graph neural network with independently learnable edge gates, impose sparsity directly on these gates, and validate the retained edges using post-training nonlinear Granger predictive gains. The resulting causal relationships constrain the LLM to trace root causes reliably. Experiments on a dataset collected from a power metering data center show that the proposed method achieves a root-cause localization accuracy of 82.6%, outperforming the best-performing baseline by 9.2 percentage points. The method also outputs explicit fault-propagation paths, indicating that causal-structure constraints can improve both the accuracy and interpretability of fault diagnosis.