Evaluating Large Language Models in Turkish Short Answer Scoring: Validity, Reliability, and Fairness Perspectives
Küçük Resim Yok
Tarih
2026
Yazarlar
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
Sakarya University
Erişim Hakkı
info:eu-repo/semantics/closedAccess
Özet
This study examines the performance of large language models (LLMs) in Turkish short-answer assessments within the measurement and evaluation theory framework. The GPT, Gemini, Gemma, and LLaMA models were evaluated under zero-shot and one-shot conditions with rubric support. The results show that LLMs have high self-consistency, but decision reliability can vary depending on prompt format and example sensitivity. Formulating rubrics with clear and concrete performance indicators increases model-human alignment and assessment fairness. Furthermore, error analyses revealed that while models often exhibit systematic low-scoring tendencies, they also display an ‘inverted-U’ error pattern, showing higher reliability at score extremes but struggling significantly with evaluating partial knowledge (intermediate scores). The results indicate that LLMs can support teachers in formative assessment when properly structured rubrics are used, but ethical oversight and pedagogical responsibility remain indispensable in final decisions. © 2026, Sakarya University. All rights reserved.
Açıklama
Anahtar Kelimeler
Artificial Intelligence In Education, Automated Scoring, Evaluation Rubric, Large Language Models (Llm), Validity And Fairness
Kaynak
Sakarya University Journal of Computer and Information Sciences
WoS Q Değeri
Scopus Q Değeri
Q3
Cilt
9
Sayı
3












