Abstract
Large language models (LLMs) offer new opportunities to make statistical analysis more accessible to medical researchers, but their use introduces risks including hallucinated functions, computational errors, and inappropriate statistical procedures. This paper proposes an R-based tool-calling framework for LLM-assisted categorical data analysis, using Pearson’s chi-square test as a proof of concept. The framework separates natural-language interaction from statistical computation by allowing the LLM to identify the analytical requirement and invoke a predefined R statistical tool, rather than independently generating the statistical computation. A sandbox-based approach, in which the LLM generated and executed R code, was evaluated for comparison. In a 10-iteration simulation study, both approaches achieved 100\% agreement with direct R analysis in statistical test selection and returned identical p-values. The approaches were further evaluated using a real-world medical dataset, where complete agreement with direct R analysis was observed across six categorical analyses. These findings demonstrate the feasibility of integrating LLMs with established statistical software while retaining greater control over statistical computation. The main contribution of this work is a practical and reproducible framework that positions the LLM as an interface to validated statistical procedures rather than as the statistical computing engine itself. The approach provides a foundation for extending LLM tool calling to broader statistical workflows in medical research.