TY - GEN
T1 - Do LLMs Make Mistakes Like Students? Exploring Natural Alignments Between Language Models and Human Error Patterns
AU - Liu, Naiming
AU - Sonkar, Shashank
AU - Baraniuk, Richard
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2025.
PY - 2025
Y1 - 2025
N2 - Large Language Models (LLMs) have demonstrated remarkable capabilities in various educational tasks, yet their alignment with human learning patterns, particularly in predicting which incorrect options students are most likely to select in multiple-choice questions (MCQs), remains underexplored. Our work investigates the relationship between LLM generation likelihoods and student response distributions in MCQs with a specific focus on distractor selections. We collect a comprehensive dataset of MCQs with real-world student response distributions to explore two fundamental research questions: (1). RQ1 - Do the distractors that students more frequently select correspond to those that LLMs assign higher generation likelihood to? (2). RQ2 - When an LLM selects a incorrect choice, does it choose the same distractor that most students pick? Our experiments reveals moderate correlations between LLM-assigned probabilities and student selection patterns for distractors in MCQs. Additionally, when LLMs make mistakes, they are more likley to select the same incorrect answers that commonly mislead students, which is a pattern consistent across both small and large language models. Our work provides empirical evidence that despite LLMs’ strong performance on generating educational content, there remains a gap between LLM’s underlying reasoning process and human cognitive processes in identifying confusing distractors. Our findings also have significant implications for educational assessment development. The smaller language models could be efficiently utilized for automated distractor generation as they demonstrate similar patterns in identifying confusing answer choices as larger language models. This observed alignment between LLMs and student misconception patterns opens new opportunities for generating high-quality distractors that complement traditional human-designed distractors.
AB - Large Language Models (LLMs) have demonstrated remarkable capabilities in various educational tasks, yet their alignment with human learning patterns, particularly in predicting which incorrect options students are most likely to select in multiple-choice questions (MCQs), remains underexplored. Our work investigates the relationship between LLM generation likelihoods and student response distributions in MCQs with a specific focus on distractor selections. We collect a comprehensive dataset of MCQs with real-world student response distributions to explore two fundamental research questions: (1). RQ1 - Do the distractors that students more frequently select correspond to those that LLMs assign higher generation likelihood to? (2). RQ2 - When an LLM selects a incorrect choice, does it choose the same distractor that most students pick? Our experiments reveals moderate correlations between LLM-assigned probabilities and student selection patterns for distractors in MCQs. Additionally, when LLMs make mistakes, they are more likley to select the same incorrect answers that commonly mislead students, which is a pattern consistent across both small and large language models. Our work provides empirical evidence that despite LLMs’ strong performance on generating educational content, there remains a gap between LLM’s underlying reasoning process and human cognitive processes in identifying confusing distractors. Our findings also have significant implications for educational assessment development. The smaller language models could be efficiently utilized for automated distractor generation as they demonstrate similar patterns in identifying confusing answer choices as larger language models. This observed alignment between LLMs and student misconception patterns opens new opportunities for generating high-quality distractors that complement traditional human-designed distractors.
KW - AI-human Alignment
KW - Cognitive Student Modeling
KW - Large Language Models
KW - Student Misconceptions
UR - https://www.scopus.com/pages/publications/105012022356
UR - https://www.scopus.com/inward/citedby.url?scp=105012022356&partnerID=8YFLogxK
U2 - 10.1007/978-3-031-98459-4_26
DO - 10.1007/978-3-031-98459-4_26
M3 - Conference contribution
AN - SCOPUS:105012022356
SN - 9783031984587
T3 - Lecture Notes in Computer Science
SP - 364
EP - 377
BT - Artificial Intelligence in Education - 26th International Conference, AIED 2025, Proceedings
A2 - Cristea, Alexandra I.
A2 - Walker, Erin
A2 - Lu, Yu
A2 - Santos, Olga C.
A2 - Isotani, Seiji
PB - Springer Science and Business Media Deutschland GmbH
T2 - 26th International Conference on Artificial Intelligence in Education, AIED 2025
Y2 - 22 July 2025 through 26 July 2025
ER -