با همکاری مشترک دانشگاه پیام نور و انجمن مدیریت دولتی ایران و انجمن مدیریت رفتار سازمانی

نوع مقاله : اکتشافی

نویسندگان

1 دانشیار، گروه مدیریت بازرگانی، دانشکده مدیریت و حسابداری، دانشگاه علامه طباطبایی، تهران، ایران.

2 دانشجوی دکتری، گروه رفتار سازمانی و منابع انسانی، دانشکده مدیریت و حسابداری، دانشگاه علامه طباطبایی، تهران، ایران.

چکیده

این پژوهش با هدف طراحی و اعتبارسنجی روش دلفی مبتنی بر هوش مصنوعی بر اساس بازی تقلید (تست تورینگ) انجام شد. سؤال اصلی تحقیق این بود که چگونه می‌توان از تست تورینگ به‌عنوان روشی نوین برای سنجش اعتبار نتایج دلفی مبتنی بر هوش مصنوعی استفاده کرد. این مطالعه از نوع کاربردی–توسعه‌ای است و جامعه پژوهش شامل ۳۰ نفر از متخصصان منابع انسانی، متخصصان هوش مصنوعی و ارزیابان مستقل بود. داده‌ها به‌صورت ثانویه گردآوری شد و خروجی سه مدل زبانی  Microsoft Copilot، ChatGPT و Gemini   به‌عنوان ورودی در تست تورینگ برای تحلیل میزان شباهت نتایج دلفی انسانی و هوش مصنوعی مورد استفاده قرار گرفت. این داده‌ها از مقالات قبلی نویسندگان که نیازهای توسعه‌ای مدیران منابع انسانی را با روش دلفی احصا کرده بودند، استخراج شده‌اند. این پژوهش به‌عنوان اولین مطالعه در زمینه استفاده از آزمون تورینگ برای ارزیابی یک روش پژوهشی، نشان داد که دلفی مبتنی بر هوش مصنوعی قادر است داده‌هایی تولید کند که از نظر شباهت به پاسخ‌های انسانی پایا و قابل اعتماد هستند. نتایج حاکی از آن است که مدل‌های زبانی بزرگ رفتار و سبک پاسخ‌دهی انسانی را به‌خوبی شبیه‌سازی کرده‌اند و تمایز بین پاسخ‌های انسانی و ماشینی برای شرکت‌کنندگان دشوار بوده است. این یافته‌ها اعتبار روش دلفی مبتنی بر هوش مصنوعی را تقویت کرده و نشان می‌دهد که این روش می‌تواند به‌عنوان ابزاری مؤثر در پژوهش، بدون کاهش کیفیت تحلیل انسانی، مورد استفاده قرار گیرد.

کلیدواژه‌ها

موضوعات

عنوان مقاله [English]

Validation of the Artificial Intelligence–Driven Delphi Method in Management Research Based on the Turing Test

نویسندگان [English]

  • hamed dehghanan 1
  • Zahra Pouramini 2

1 Associate Professor, Department of Management and Accounting, Allameh Tabataba`i University, Tehran, Iran.

2 Ph.D. Student, Department of Management and Accounting, Allameh Tabataba`i University, Tehran, Iran.

چکیده [English]

Introduction                                   
In the past decade, the expansion of artificial intelligence (AI) and large language models (LLMs) has brought a fundamental transformation to qualitative research methods and expert-based decision-making processes. The Delphi method, traditionally grounded in the collective judgments of human experts, has entered a new phase of evolution with the advent of AI. The present study aimed to design and validate an AI-based Delphi method while leveraging the imitation game, or Turing test, as a tool to assess the credibility of its outputs. The central question guiding this research was:
"How can the Turing test be utilized as a novel approach to evaluate the validity and similarity of AI-based Delphi results compared to traditional human-based Delphi?"
This question is particularly significant because researchers engaging with modern language models face a fundamental challenge: whether AI-generated responses can be considered as reliable as human judgments in terms of logic, reasoning, and coherence. This study seeks to provide a well-grounded answer and establish a methodological framework for the combined use of Delphi and AI in applied research contexts.
 
Mothodology
This research is of an applied-developmental nature and was conducted with the objective of validating the AI-based Delphi method through the Turing test. The study population consisted of 30 purposively selected participants, including human resources experts, AI specialists, and independent evaluators, ensuring the data could be examined from human, technical, and impartial judgment perspectives.
In prior studies, the AI-based Delphi methodology was comprehensively designed and explained, and in a separate study, it was applied to prioritize the developmental needs of human resources managers in the context of coaching. In the current study, the method is only briefly introduced, and the data derived from previous research serve as input for the Turing test.
In the AI-based Delphi process, large language models (ChatGPT, Copilot, and Gemini) acted as virtual experts, responding to research questions and presenting their reasoning explicitly through structured prompts. Subsequently, the outputs of these models were validated using the Turing test. Evaluators engaged in brief dialogues without knowing the source of the responses (human or AI) and were required to determine the nature of the respondent. Responses that could not be confidently distinguished were considered “successful” in the Turing test.
Thus, this study represents a continuation of prior research, examining the similarity, coherence, and credibility of AI-based Delphi outputs in a real-world research setting using actual data and the Turing test.
 
Findings
Findings indicated that 56% of participants were female and 44% male, with a mean age of 36 years (range: 27–52) and approximately 72% holding a master’s degree or higher. This diversity of expertise and experience provided a robust basis for analyzing AI model behavior in comparison with human responses. The data employed were secondary sources, comprising results from three rounds of traditional and AI-based Delphi in previous studies by the authors, focused on identifying developmental priorities for human resources managers. These datasets included competency ratings, comparisons between traditional and AI-based approaches, and reliability and validity matrices, with over 85% of the data confirmed as reliable and valid.
During the study, key questions were collected from both human and AI sources at each Delphi stage, and evaluators were tasked with distinguishing between human and machine responses, in line with the Turing test framework. Qualitative analysis revealed that AI-generated responses were often so similar to human answers that evaluators struggled to distinguish them, with an overall Turing test success rate of only 23%, indicating that over two-thirds of responses could not be correctly classified. This finding highlights the strong ability of the models to simulate human behavior.
Content analysis showed that evaluators primarily relied on cues such as grammatical and stylistic errors, vocabulary diversity, interactive responsiveness, display of emotion and empathy, variation in tone, and references to personal experience—features more prevalent in human responses. In contrast, AI responses were characterized by high linguistic precision, formal structure, uniform tone, accurate and error-free answers, consistent response speed, and absence of personal narratives or experiential examples. Behavioral and linguistic patterns further revealed that participants expected human responses to be more dynamic, personal, and occasionally non-linear, whereas machine responses were predominantly formal, precise, and structured.
 
Discussion and Conclusion
These findings not only confirm the capability of large language models to emulate human-like behavior and response style but also underscore the significance of linguistic, behavioral, and emotional cues in distinguishing between humans and AI. Furthermore, integrating human and AI Delphi data and analyzing evaluator judgments demonstrated that AI models could produce highly coherent responses, logically consistent themes, and alignment with management literature, establishing their effectiveness as tools for analyzing developmental needs of human resources managers.
Overall, the results indicate that the AI-based Delphi method, combined with the Turing test, can effectively evaluate the validity and similarity of machine-generated responses to human ones, enabling the production of reliable and credible data in real-world research settings.
As the first study to employ the Turing test for evaluating a research methodology, this study demonstrates that AI-based Delphi can generate data that are reliable and highly similar to human responses. The findings show that large language models successfully emulate human behavior and response style, making it challenging for participants to distinguish human from machine responses. This evidence strengthens the credibility of the AI-based Delphi method, suggesting it can serve as an effective research tool without compromising the quality of human-like analysis.
Based on the findings, practical recommendations include optimizing the design of Turing test scenarios and criteria to more accurately assess the capabilities of language models, integrating the Delphi method with the Turing test for practical validation of outputs, fully documenting the process and managing potential biases, attending to ethical considerations and data confidentiality, and developing and standardizing prompt-engineering skills for effective interaction with language models. Additionally, potential limitations—such as the influence of participants’ attitudes and experiences, ethical and confidentiality constraints, sensitivity of results to prompt design and model characteristics, limitations in sample diversity and size, and the inability to fully control the knowledge and biases embedded in language models—should be considered in the design and analysis of future studies.
These findings provide practical guidance for researchers employing AI-based Delphi methods, enabling the combined use of human and AI collective intelligence with higher accuracy and reliability.

کلیدواژه‌ها [English]

  • Artificial Intelligence
  • Delphi Method
  • Turing Test
  • Personal Development
Aithal, P. S., & Aithal, S. (2023). Use of AI-based GPTs in experimental, empirical, and exploratory research methods. International Journal of Case Studies in Business, IT, and Education (IJCSBE), 7(3), 411-425. http://dx.doi.org/10.2139/ssrn.4673846.
Alvarado R. (2022). Should we replace radiologists with deep learning?. Bioethics, 36(2), 121-133. https://doi.org/10.1111/bioe.12959.
Bertolotti, F., & Mari, L. (2025). An LLM-based Delphi study to predict GenAI evolution. arXiv preprint arXiv:2502.21092. https://doi.org/10.48550/arXiv.2502.21092
Borowiec ML, Dikow RB, Frandsen PB, McKeeken A, Valentini G, and White AE. (2022). Deep learning as a tool for ecology and evolution. Methods in Ecology and Evolution, 13(8), 1640-1660. https://doi.org/10.1111/2041-210X.13901
Bothra A, Cao Y, Černý J, and Arora G. (2023). The Epidemiology of Infectious Diseases Meets AI: A Match Made in Heaven. Pathogens, 12(2), 317. https://doi.org/10.3390/pathogens12020317.
Bughin, J., Manyika, J., & Woetzel, J. (2017). A Future That Works: Automation, Employment, and Productivity. McKinsey Global Institute.
Chang, T. A., & Bergen, B. K. (2024). Language model behavior: A comprehensive survey. Computational Linguistics, 50(1), 293-350. https://doi.org/10.1162/coli_a_00492.
Dehghanan, H., Pouramini, Z., Yazdanshenas, M., & Raeisi Vanani, I. (2025). AI-based Delphi methodology: An innovative approach to triangulation in management research. Journal of Public Management, in press. (In Persian)
Fink-Hafner, T. Dagen, M. Doušak, M. Novak, and M. (2019). Delphi method: strengths and weaknesses. Advances in Methodology and Statistics, 16(2),1–19,
Hudoud, A. (2025). Integrating Artificial Intelligence into Research Methodology: Examining Potential Bias and Mitigation Strategies. The Arab Journal For Quality Assurance in Higher Education, 18(64). https://doi.org/10.20428/ajqahe.v18i64.2680
Jannai, D., Meron, A., Lenz, B., Levine, Y., & Shoham, Y. (2023). Human or not? A gamified approach to the Turing test. arXiv preprint arXiv:2305.20010.
Kelly, J. (2023). Goldman Sachs predicts 300 million jobs will Be lost or degraded by artificial intelligence. Forbes. https://tinyurl.com/3xb437rb.
Knoth, N., Tolzin, A., Janson, A., & Leimeister, J. M. (2024). AI literacy and its implications for prompt engineering strategies. Computers and Education: Artificial Intelligence, 6, 100225. https://doi.org/10.1016/j.caeai.2024.100225.
Mai, V., Neef, C., & Richert, A. (2022). Clicking vs. writing-The impact of a chatbot’s interaction method on the working alliance in AI-based coaching. Coaching Theorie & Praxis, 8(1), 15-31. https://doi.org/10.1365/s40896-021-00063-3.
Mayer-Schönberger, & Cukier, K. (2013). Big Data: A Revolution That Will Transform How We Live, Work, and Think. Houghton Mifflin Harcourt.
Mueller, R. M., Thoring, K., Klöckner, H. W., & Larsen, K. (2024). Crafting Future Scenarios with the Help of AI: Potentials of a Hybrid Delphi Expert Panel.
Nasa, R. Jain, and D. Juneja.(2022).  Delphi methodology in healthcare research: how to decide its appropriateness. World journal of methodology, 11(4),116. https://doi.org/10.5662/wjm.v11.i4.116.
Pathak, J., Wikner, A., Fussell, R., Chandra, S., Hunt, B. R., Girvan, M., & Ott, E. (2018). Hybrid forecasting of chaotic processes: Using machine learning in conjunction with a knowledge-based model. Chaos: An interdisciplinary journal of nonlinear science28(4). https://doi.org/10.1063/1.5028373.
Pouramini, Z., Dehghanan, H., Yazdanshenas, M., & Raeisi Vanani, I. (2025). Prioritizing HR managers’ personal development programs in the field of coaching using the AI-based Delphi method. Iranian Journal of Management Research. in press.[Persian].
Robert M. French. (2000). The Turing Test: The first 50 years. Trends in Cognitive Sciences, 4(3),115–122.
Russell, S. J., & Norvig, P. (2021). Artificial intelligence: A modern approach (4th ed.). Pearson.
Shalev-Shwartz, S & Ben-David, S. (2014). Understanding Machine Learning: From Theory to Algorithms. Cambridge University Press.
Shang.(2023). Use of delphi in health sciences research: a narrative review. Medicine, 102(7), e32829.https://doi.org/10.1097/MD.0000000000032829.
Sherry Turkle. 2011. Life on the Screen. Simon and Schuster.
Speed, C., & Metwally, A. A. (2025). The Human-AI Hybrid Delphi Model: A Structured Framework for Context-Rich, Expert Consensus in Complex Domains. arXiv preprint arXiv:2508.09349.
Ujhelyi, A., Almosdi, F., & Fodor, A. (2022). Would you pass the turing test? Influencing factors of the turing decision. Psihologijske teme31(1), 185-202. https://doi.org/10.31820/pt.31.1.9.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., ... & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems35, 24824-24837.
Witten, I. H., Frank, E., & Hall, M. A. (2011). Data Mining: Practical Machine Learning Tools and Techniques. Morgan Kaufmann.