Evaluating Large Language Models in Turkish Short Answer Scoring: Validity, Reliability, and Fairness Perspectives

Küçük Resim Yok

Tarih

2026

Dergi Başlığı

Dergi ISSN

Cilt Başlığı

Yayıncı

Sakarya University

Erişim Hakkı

info:eu-repo/semantics/closedAccess

Özet

This study examines the performance of large language models (LLMs) in Turkish short-answer assessments within the measurement and evaluation theory framework. The GPT, Gemini, Gemma, and LLaMA models were evaluated under zero-shot and one-shot conditions with rubric support. The results show that LLMs have high self-consistency, but decision reliability can vary depending on prompt format and example sensitivity. Formulating rubrics with clear and concrete performance indicators increases model-human alignment and assessment fairness. Furthermore, error analyses revealed that while models often exhibit systematic low-scoring tendencies, they also display an ‘inverted-U’ error pattern, showing higher reliability at score extremes but struggling significantly with evaluating partial knowledge (intermediate scores). The results indicate that LLMs can support teachers in formative assessment when properly structured rubrics are used, but ethical oversight and pedagogical responsibility remain indispensable in final decisions. © 2026, Sakarya University. All rights reserved.

Açıklama

Anahtar Kelimeler

Artificial Intelligence In Education, Automated Scoring, Evaluation Rubric, Large Language Models (Llm), Validity And Fairness

Kaynak

Sakarya University Journal of Computer and Information Sciences

WoS Q Değeri

Scopus Q Değeri

Q3

Cilt

9

Sayı

3

Künye