Evaluating Large Language Models in Turkish Short Answer Scoring: Validity, Reliability, and Fairness Perspectives
| dc.contributor.author | Kara, Abdulkadir | |
| dc.contributor.author | Yıldırım, Serkan | |
| dc.date.accessioned | 2026-09-01T15:48:04Z | |
| dc.date.available | 2026-09-01T15:48:04Z | |
| dc.date.issued | 2026 | |
| dc.department | Bayburt Üniversitesi | |
| dc.description.abstract | This study examines the performance of large language models (LLMs) in Turkish short-answer assessments within the measurement and evaluation theory framework. The GPT, Gemini, Gemma, and LLaMA models were evaluated under zero-shot and one-shot conditions with rubric support. The results show that LLMs have high self-consistency, but decision reliability can vary depending on prompt format and example sensitivity. Formulating rubrics with clear and concrete performance indicators increases model-human alignment and assessment fairness. Furthermore, error analyses revealed that while models often exhibit systematic low-scoring tendencies, they also display an ‘inverted-U’ error pattern, showing higher reliability at score extremes but struggling significantly with evaluating partial knowledge (intermediate scores). The results indicate that LLMs can support teachers in formative assessment when properly structured rubrics are used, but ethical oversight and pedagogical responsibility remain indispensable in final decisions. © 2026, Sakarya University. All rights reserved. | |
| dc.identifier.doi | 10.35377/saucis…1835608 | |
| dc.identifier.endpage | 994 | |
| dc.identifier.issn | 2636-8129 | |
| dc.identifier.issue | 3 | |
| dc.identifier.scopus | 2-s2.0-105045933773 | |
| dc.identifier.scopusquality | Q3 | |
| dc.identifier.startpage | 980 | |
| dc.identifier.uri | https://doi.org/10.35377/saucis…1835608 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.12403/8370 | |
| dc.identifier.volume | 9 | |
| dc.indekslendigikaynak | Scopus | |
| dc.language.iso | en | |
| dc.publisher | Sakarya University | |
| dc.relation.ispartof | Sakarya University Journal of Computer and Information Sciences | |
| dc.relation.publicationcategory | Makale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/closedAccess | |
| dc.snmz | KA_Scopus_20260820 | |
| dc.subject | Artificial Intelligence In Education | |
| dc.subject | Automated Scoring | |
| dc.subject | Evaluation Rubric | |
| dc.subject | Large Language Models (Llm) | |
| dc.subject | Validity And Fairness | |
| dc.title | Evaluating Large Language Models in Turkish Short Answer Scoring: Validity, Reliability, and Fairness Perspectives | |
| dc.type | Article |












