Can Large Language Models (LLMs) be Trusted for Power Analysis? An Empirical Evaluation

Authors

DOI:

https://doi.org/10.35566/jbds/jikim

Keywords:

Large language models, ChatGPT, Claude, Gemini, Llama, Statistical power analysis, Sample size calculation, Research methods

Abstract

Power analysis is critical for assuring rigor and validity of quantitative research yet remains underutilized due to technical challenges associated with specialized software. At the same time, large language models (LLMs) are being rapidly integrated into research practice, raising interest in their potential to assist statistical and research design tasks. However, despite their widespread adoption, the reliability of LLMs in supporting statistically rigorous procedures has not been systematically evaluated, posing risks for unexamined or overly optimistic use. To address this gap, we evaluated four widely used LLMs—ChatGPT (GPT-3.5, GPT-4, GPT-4o) and Llama 3.2—across two experiments. Experiment 1 examined whether LLMs could calculate required sample sizes for common statistical tests (two-sample t-test, one-way ANOVA, and χ² goodness-of-fit test) under different prompting strategies, including direct calculation versus R/Python code generation. Experiment 2 assessed models’ ability to identify missing input parameters necessary for power analysis, which is a task that requires methodological understanding. Results revealed that GPT-4 and GPT-4o performed well when generating R code for sample size estimation, but struggled with direct numerical calculation. Furthermore, while LLMs were able to detect missing information, their reliability varied by statistical context. Findings suggest that while LLMs may offer support in structuring and initiating power analysis, they cannot substitute for expert judgment. Overall, the study underscores the importance of critically evaluating LLM performance in statistically demanding tasks. Responsible integration of LLM requires critical oversight, cross-verification, and methodological evaluation.

References

Aberson, C. L. (2019). Applied power analysis for the behavioral sciences. Routledge.

Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., et al. (2023). Gpt-4 technical report. arXiv preprint arXiv:2303.08774. Retrieved from https://doi.org/10.48550/arXiv.2303.08774

Anderson, S. F., Kelley, K., & Maxwell, S. E. (2017). Sample-size planning for more accurate statistical power: A method adjusting sample effect sizes for publication bias and uncertainty. Psychological Science, 28(11), 1547-1562. Retrieved from https://doi.org/10.1177/0956797617723724

Arend, M. G., & Schäfer, T. (2019). Statistical power in two-level models: A tutorial based on monte carlo simulation. Psychological Methods, 24(1), 1–19. Retrieved from https://doi.org/10.1037/met0000195

Arnold, B. F., Hogan, D. R., Colford, J., & Hubbard, A. E. (2011). Simulation methods to estimate design power: An overview for applied research. BMC Medical Research Methodology, 11(1). Retrieved from https://doi.org/10.1186/1471-2288-11-94

Aubin Le Quéré, M., Schroeder, H., Randazzo, C., Gao, J., Epstein, Z., Perrault, S. T., et al. (2024). Llms as research tools: Applications and evaluations in hci data work. In Extended abstracts of the chi conference on human factors in computing systems (p. 1-7). Retrieved from https://doi.org/10.1145/3613905.3636301

Bakker, M., Hartgerink, C. H. J., Wicherts, J. M., & van der Maas, H. L. J. (2016). Researchers’ intuitions about power in psychological research. Psychological Science, 27(8), 1069-1077. Retrieved from https://psycnet.apa.org/doi/10.1177/0956797616647519

Benjamini, Y., & Hochberg, Y. (1995). Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the Royal Statistical Society. Series B (Methodological), 57(1), 289–300. Retrieved from https://doi.org/10.1111/j.2517-6161.1995.tb02031.x

Birhane, A., Kasirzadeh, A., Leslie, D., & Wachter, S. (2023). Science in the age of large language models. Nature Reviews Physics, 5(5), 277-280. Retrieved from https://doi.org/10.1038/s42254-023-00581-4

Biswas, S. (2023). Role of chatgpt in computer programming. Mesopotamian Journal of Computer Science, 2023, 8-16. Retrieved from https://doi.org/10.58496/MJCSC/2023/002

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., … Amodei, D. (2020). Language models are few-shot learners. arXiv preprint arXiv:2005.14165. Retrieved from https://doi.org/10.48550/arxiv.2005.14165

Button, K. S., Ioannidis, J. P., Mokrysz, C., Nosek, B. A., Flint, J., Robinson, E. S., & Munafò, M. R. (2013). Power failure: why small sample size undermines the reliability of neuroscience. Nature Reviews. Neuroscience, 14(5), 365–376. Retrieved from https://doi.org/10.1038/nrn3475

Champely, S., Ekstrom, C., Dalgaard, P., Gill, J., Weibelzahl, S., Anandkumar, A., … De Rosario, H. (2017). pwr: Basic functions for power analysis. https://cran.r-project.org/web/packages/pwr/. (Software)

Charfeddine, M., Kammoun, H. M., Hamdaoui, B., & Guizani, M. (2024). Chatgpt’s security risks and benefits: Offensive and defensive use-cases, mitigation measures, and future implications. IEEE Access, 12. Retrieved from https://doi.org/10.1109/ACCESS.2024.3367792

Chen, T. C., Kaminski, E., Koduri, L., Singer, A., Singer, J., Couldwell, M., … Wang, A. (2023). Chatgpt as a neuro-score calculator: Analysis of a large language model’s performance on various neurological exam grading scales. World Neurosurgery, 179, 342–347. Retrieved from https://doi.org/10.1016/j.wneu.2023.08.088

Chen, T. J. (2023). Chatgpt and other artificial intelligence applications speed up scientific writing. Journal of the Chinese Medical Association, 86(4), 351-353. Retrieved from https://doi.org/10.1097/JCMA.0000000000000900

Cheong, I., Xia, K., Feng, K. K., Chen, Q. Z., & Zhang, A. X. (2024). (a) i am not a lawyer, but…: Engaging legal experts towards responsible llm policies for legal advice. In The 2024 acm conference on fairness, accountability, and transparency (p. 2454-2469). Retrieved from https://doi.org/10.1145/3630106.3659048

Cohen, J. (1988). Statistical power analysis for the behavioral sciences. (2nd ed). L. Erlbaum Associates.

Colavizza, G. (2025). Large language models for social science research. https://www.summerschoolsineurope.eu/course/large-language-models-for-social-science-research/. ([Workshop]. Università della Svizzera italiana, Lugano, Switzerland)

Corp., I. (2021). Ibm spss statistics for mac, version 28.0. (Armonk, NY: IBM Corp.)

Correll, J., Mellinger, C., McClelland, G. H., & Judd, C. M. (2020). Avoid cohen’s ‘small’, ‘medium’, and ‘large’ for power analysis. Trends in Cognitive Sciences, 24(3), 200–207. Retrieved from https://doi.org/10.1016/j.tics.2019.12.009

Espejel, J. L., Ettifouri, E. H., Alassan, M. S. Y., Chouham, E. M., & Dahhane, W. (2023). Gpt-3.5, gpt-4, or bard? evaluating llms reasoning ability in zero-shot setting and performance boosting through prompts. arXiv preprint arXiv:2305.12477. Retrieved from https://doi.org/10.48550/arxiv.2305.12477

Evkaya, O., & de Carvalho, M. (2024). Decoding ai: The inside story of data analysis in chatgpt. arXiv preprint arXiv:2024.08480.. Retrieved from https://doi.org/10.48550/arxiv.2404.08480

Faul, F., Erdfelder, E., Lang, A.-G., & Buchner, A. (2007). Gpower 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behavior Research Methods, 39(2), 175–191. Retrieved from https://doi.org/10.3758/BF03193146

Finch, S., Cumming, G., & Thomason, N. (2001). Reporting of statistical inference in the journal of applied psychology: Little evidence of reform. Educational and Psychological Measurement, 61(2), 181–210. Retrieved from https://doi.org/10.1177/0013164401612001

for Social Science, O. D. I., & Innovations, E. (2025). Large language models in social science research. https://odissei-data.nl/event/workshop-llm/. ([Workshop]. Utrecht University, Utrecht, Netherlands)

Frank, M. C. (2023). Baby steps in evaluating the capacities of large language models. Nature Reviews Psychology, 2(8), 451-452. Retrieved from https://doi.org/10.1038/s44159-023-00211-x

Frieder, S., Pinchetti, L., Chevalier, A., Griffiths, R.-R., Salvatori, T., Lukasiewicz, T., … Berner, J. (2023). Mathematical capabilities of chatgpt. arXiv preprint arXiv:2301.13867. Retrieved from https://doi.org/10.48550/arXiv.2301.13867

Fritz, A., Scherndl, T., & Kühberger, A. (2013). A comprehensive review of reporting practices in psychological journals: Are effect sizes really enough? Theory & Psychology, 23(1), 98–122. Retrieved from https://doi.org/10.1177/0959354312436870

Hermann, C. E., Patel, J. M., Boyd, L., Growdon, W. B., Aviki, E., & Stasenko, M. (2023). Let’s chat about cervical cancer: Assessing the accuracy of chatgpt responses to cervical cancer questions. Gynecologic Oncology, 179, 164–168. Retrieved from https://doi.org/10.1016/j.ygyno.2023.11.008

Hu, X., Zhao, Z., Wei, S., Chai, Z., Ma, Q., Wang, G., et al. (2024). Infiagent-dabench: Evaluating agents on data analysis tasks. arXiv preprint arXiv:2401.05507. Retrieved from https://doi.org/10.48550/arXiv.2401.05507

Inc., S. I. (2023). Sas/stat© 15.3 user’s guide. (Cary, NC: SAS Institute Inc.)

in Europe, S. S. (2025). Large language models for social science research summer course. https://www.summerschoolsineurope.eu/course/large-language-models-for-social-science-research/. (Retrieved April 28, 2025)

Ioannidis, J. P. A. (2005). Why most published research findings are false. Research Integrity in the Biomedical Sciences, 2(8), 0696–0701. Retrieved from https://doi.org/10.1371/journal.pmed.0020124

Jahangiri, Y. (2023). Can chat generative pretraining transformer (chatgpt) be used for statistical analysis of research data? Journal of Vascular and Interventional Radiology, 34(12), 2242-2246. Retrieved from https://doi.org/10.1016/j.jvir.2023.09.010

Jalali, M. S., & Akhavan, A. (2024). Integrating ai language models in qualitative research: Replicating interview data analysis with chatgpt. System Dynamics Review, 40(3). Retrieved from https://doi.org/10.1002/sdr.1772

Jiang, B., Xie, Y., Hao, Z., Wang, X., Mallick, T., Su, W. J., et al. (2024). A peek into token bias: Large language models are not yet genuine reasoners. arXiv preprint arXiv:2406.11050. Retrieved from https://doi.org/10.48550/arXiv.2406.11050

Jiao, W., Wang, W., Huang, J. T., Wang, X., & Tu, Z. (2023). Is chatgpt a good translator? a preliminary study. arXiv preprint arXiv:2301.08745.. Retrieved from https://doi.org/10.48550/arXiv.2301.08745

Kashefi, A., & Mukerji, T. (2023). Chatgpt for programming numerical methods. arXiv preprint arXiv:2303.12093. Retrieved from https://doi.org/10.48550/arxiv.2303.12093

Khraisha, Q., Put, S., Kappenberg, J., Warraitch, A., & Hadfield, K. (2024). Can large language models replace humans in systematic reviews? evaluating gpt4’s efficacy in screening and extracting data from peerreviewed and grey literature in multiple languages. Research Synthesis Methods, 15(4). Retrieved from https://doi.org/10.1002/jrsm.1715

Kingsley, B. E., & Robertson, J. M. (2017). Exploring reticence in research methods: The experience of studying psychological research methods in higher education. Psychology Teaching Review, 23(2), 4–19. Retrieved from https://doi.org/10.53841/bpsptr.2017.23.2.4

Kitamura, F. C. (2023). Chatgpt is shaping the future of medical writing but still requires human judgment. Radiology, 307(2). Retrieved from https://doi.org/10.1148/radiol.230171

Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., & Iwasawa, Y. (2023). Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916.. Retrieved from https://doi.org/10.48550/arxiv.2205.11916

Lecler, A., Duron, L., & Soyer, P. (2023). Revolutionizing radiology with gpt-based models: Current applications, future possibilities and limitations of chatgpt. Diagnostic and Interventional Imaging, 104(6), 269–274. Retrieved from https://doi.org/10.1016/j.diii.2023.02.003

Le Mens, G. (2024). Using large language models for empirical research in social science. https://www.ibei.org/en/using-large-language-models-for-empirical-research-in-social-science_350753. ([Workshop]. Institut Barcelona Estudis Internacionals, Barcelona, Spain)

Liu, J., Xia, C. S., Wang, Y., & Zhang, L. (2024). Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems, 36. Retrieved from https://doi.org/10.48550/arXiv.2305.01210

Lixandru, I.-D. (2024). The use of artificial intelligence for qualitative data analysis: Chatgpt. Informatica Economica, 28(1), 57–67. Retrieved from https://doi.org/10.24818/issnl4531305/28.1.2024.05

Maxwell, S. E. (2004). The persistence of underpowered studies in psychological research: Causes, consequences, and remedies. Psychological Methods, 9(2), 147–163. Retrieved from https://doi.org/10.1037/1082-989X.9.2.147

Maxwell, S. E., Lau, M. Y., & Howard, G. S. (2015). Is psychology suffering from a replication crisis?: What does “failure to replicate” really mean? The American Psychologist, 70(6), 487–498. Retrieved from https://doi.org/10.1037/a0039400

Methnani, J., Latiri, I., Dergaa, I., Chamari, K., & Saad, H. B. (2023). Chatgpt for sample-size calculation in sports medicine and exercise sciences: A cautionary note. International Journal of Sports Physiology and Performance, 18(10), 1219–1223. Retrieved from https://doi.org/10.1123/ijspp.2023-0109

Mirzadeh, I., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S., & Farajtabar, M. (2024). Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models. arXiv preprint arXiv:2410.05229. Retrieved from https://doi.org/10.48550/arXiv.2410.05229

Moshagen, M., & Bader, M. (2023). sempower: General power analysis for structural equation models. Behavior Research Methods, 56(4), 2901–2922. Retrieved from https://doi.org/10.3758/s13428-023-02254-7

Myors, B., & Murphy, K. R. (2023). Statistical power analysis: A simple and general model for traditional and modern hypothesis tests. . (Fifth edition.). Routledge. Retrieved from https://doi.org/10.4324/9781003296225

Nigar, M. S., & Mohammed, Y. S. (2023). Use chatgpt to solve programming bugs. International Journal of Information Technology & Computer Engineering, 3(01), 17–22. Retrieved from https://doi.org/10.55529/ijitc.31.17.22

OpenAI. (2022). Chatgpt [large language model]. https://chat.openai.com/chat. (Published November 30, 2022)

Paul, J., Ueno, A., & Dennis, C. (2023). Chatgpt and consumers: Benefits, pitfalls and future research agenda. International Journal of Consumer Studies, 47(4), 1213–1225. Retrieved from https://doi.org/10.1111/ijcs.12928

Piccolo, S. R., Denny, P., Luxton-Reilly, A., Payne, S., & Ridge, P. G. (2023). Many bioinformatics programming tasks can be automated with chatgpt. arXiv preprint arXiv:2303.13528.. Retrieved from https://doi.org/10.48550/arxiv.2303.13528

Qin, X. (2023). Sample size and power calculations for causal mediation analysis: A tutorial and shiny app. Behavior Research Methods, 56(3), 1738–1769. Retrieved from https://doi.org/10.3758/s13428-023-02118-0

Rasheed, Z., Waseem, M., Ahmad, A., Kemell, K. K., Xiaofeng, W., Duc, A. N., & Abrahamsson, P. (2024). Can large language models serve as data analysts? a multi-agent assisted approach for qualitative data analysis. arXiv preprint arXiv:2402.01386. Retrieved from https://doi.org/10.48550/arXiv.2402.01386

Ray, P. P. (2023). Chatgpt: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. Internet of Things and Cyber-Physical Systems, 3, 121–154. Retrieved from https://doi.org/10.1016/j.iotcps.2023.04.003

Rosenfeld, A., & Lazebnik, T. (2024). Whose llm is it anyway? linguistic comparison and llm attribution for gpt-3.5, gpt-4 and bard. arXiv preprint arXiv:2402.14533.. Retrieved from https://doi.org/10.48550/arxiv.2402.14533

Sedlmeier, P., & Gigerenzer, G. (1989). Do studies of statistical power have an effect on the power of studies? Psychological Bulletin, 105(2), 309–316. Retrieved from https://doi.org/10.1037/0033-2909.105.2.309

StataCorp. (2023). Stata 18 stata power, precision, and sample-size reference manual. (College Station, TX: Stata Press)

Taylor, D. W., & Bosch, E. G. (1990). Cts: A clinical trials simulator. Statistics in Medicine, 9(7), 787–801. Retrieved from https://doi.org/10.1002/sim.4780090708

Vallat. (2018). Pingouin: Statistics in python. Journal of Open Source Software, 3(31), 1026. Retrieved from https://doi.org/10.21105/joss.01026

van Dis, E. A. M., Bollen, J., Zuidema, W., van Rooij, R., & Bockting, C. L. (2023). Chatgpt: Five priorities for research. Nature (London), 614(7947), 224–226. Retrieved from https://doi.org/10.1038/d41586-023-00288-7

Wang, H., Fu, T., Du, Y., Gao, W., Huang, K., Liu, Z., et al. (2023). Scientific discovery in the age of artificial intelligence. Nature, 620(7972), 47-60. Retrieved from https://doi.org/10.1038/s41586-023-06221-2

Watkins, R. (2024). Guidance for researchers and peer-reviewers on the ethical use of large language models (llms) in scientific research workflows. AI and Ethics, 4(4), 969-974. Retrieved from https://doi.org/10.1007/s43681-023-00294-5

Wilkinson, L., & on Statistical Inference., T. F. (1999). Statistical methods in psychology journals: Guidelines and explanations. American Psychologist, 54, 594-604. Retrieved from https://psycnet.apa.org/doi/10.1037/0003-066X.54.8.594

Xu, H., Sharaf, A., Chen, Y., Tan, W., Shen, L., Van Durme, B., et al. (2024). Contrastive preference optimization: Pushing the boundaries of llm performance in machine translation. arXiv preprint arXiv:2401.08417. Retrieved from https://doi.org/10.48550/arXiv.2401.08417

Zhang, T., Ladhak, F., Durmus, E., Liang, P., McKeown, K., & Hashimoto, T. B. (2024). Benchmarking large language models for news summarization. Transactions of the Association for Computational Linguistics, 12, 39-57. Retrieved from https://doi.org/10.1162/tacl_a_00632

Zhang, Z. (2014). Monte carlo based statistical power analysis for mediation models: Methods and software. Behavior Research Methods, 46(4), 1184–1198. Retrieved from https://doi.org/10.3758/s13428-013-0424-0

Zhao, H., Liu, Z., Wu, Z., Li, Y., Yang, T., Shu, P., et al. (2024). Revolutionizing finance with llms: An overview of applications and insights. arXiv preprint arXiv:2401.11641. Retrieved from https://doi.org/10.48550/arXiv.2401.11641

Zheng, Y., Koh, H. Y., Ju, J., Nguyen, A. T., May, L. T., Webb, G. I., & Pan, S. (2025). Large language models for scientific discovery in molecular property prediction. Nature Machine Intelligence, 1-11. Retrieved from https://doi.org/10.1038/s42256-025-00994-z

Zhu, J. J., Jiang, J., Yang, M., & Ren, Z. J. (2023). Chatgpt and environmental research. Environmental Science & Technology, 57(46), 17667–17670. Retrieved from https://doi.org/10.1021/acs.est.3c01818

Downloads

Published

2026-09-09

Issue

Section

Theory and Methods

How to Cite

Kim, H., Qi, J., Feng, Z., Zhang, X., Han, Y., He, J., & Ji, F. (2026). Can Large Language Models (LLMs) be Trusted for Power Analysis? An Empirical Evaluation. Journal of Behavioral Data Science, 6(2), 1-42. https://doi.org/10.35566/jbds/jikim