Jagged competencies: Measuring the reliability of generative AI in academic research

Large Language Models (LLMs) are increasingly viewed as a valuable tool for academic research. While LLMs have some benefits, a ‘crisis of replicability’ in management scholarship mitigates against unrestrained use. In this paper we investigate the reproducibility of LLM analyses. We analyze three L...

ver descrição completa

Detalhes bibliográficos
Autores: Thomas, Llewellyn, Romasanta, Angelo Kenneth, Pujol Priego, Laia
Tipo de documento: artigo
Data de publicação:2026
País:España
Recursos:Universitat Ramon Llull (URL)
Repositório:DAU Arxiu Digital de la Universitat Ramon Llull
OAI Identifier:oai:dau.url.edu:20.500.14342/6069
Acesso em linha:http://hdl.handle.net/20.500.14342/6069
https://doi.org/10.1016/j.jbusres.2025.115804
Access Level:Acceso aberto
Palavra-chave:Generative AI
LLM
Replication
Reproducibility
Consistency
Accuracy
Descrição
Resumo:Large Language Models (LLMs) are increasingly viewed as a valuable tool for academic research. While LLMs have some benefits, a ‘crisis of replicability’ in management scholarship mitigates against unrestrained use. In this paper we investigate the reproducibility of LLM analyses. We analyze three LLMs—ChatGPT, Claude and Mistral—over fifteen weeks, testing the consistency, accuracy and their interaction using the same prompts on the same data corpus. While our results demonstrate significant variations in reliability and consistency across the three LLMs, we also show that LLMs can exhibit deterministic and reliable behavior under specific, well-defined constraints. We argue that replicable LLM-based research will rely on understanding and validating the task-specific operational boundaries of the LLM. To ensure the responsible integration of LLMs into management research, we highlight a need for robust frameworks, transparency, ethical guidelines, and ongoing evaluation. We conclude with actionable guidance for management researchers.