Evaluating Large Language Models in Turkish Short Answer Scoring: Validity, Reliability, and Fairness Perspectives

dc.contributor.authorKara, Abdulkadir
dc.contributor.authorYıldırım, Serkan
dc.date.accessioned2026-09-01T15:48:04Z
dc.date.available2026-09-01T15:48:04Z
dc.date.issued2026
dc.departmentBayburt Üniversitesi
dc.description.abstractThis study examines the performance of large language models (LLMs) in Turkish short-answer assessments within the measurement and evaluation theory framework. The GPT, Gemini, Gemma, and LLaMA models were evaluated under zero-shot and one-shot conditions with rubric support. The results show that LLMs have high self-consistency, but decision reliability can vary depending on prompt format and example sensitivity. Formulating rubrics with clear and concrete performance indicators increases model-human alignment and assessment fairness. Furthermore, error analyses revealed that while models often exhibit systematic low-scoring tendencies, they also display an ‘inverted-U’ error pattern, showing higher reliability at score extremes but struggling significantly with evaluating partial knowledge (intermediate scores). The results indicate that LLMs can support teachers in formative assessment when properly structured rubrics are used, but ethical oversight and pedagogical responsibility remain indispensable in final decisions. © 2026, Sakarya University. All rights reserved.
dc.identifier.doi10.35377/saucis…1835608
dc.identifier.endpage994
dc.identifier.issn2636-8129
dc.identifier.issue3
dc.identifier.scopus2-s2.0-105045933773
dc.identifier.scopusqualityQ3
dc.identifier.startpage980
dc.identifier.urihttps://doi.org/10.35377/saucis…1835608
dc.identifier.urihttps://hdl.handle.net/20.500.12403/8370
dc.identifier.volume9
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherSakarya University
dc.relation.ispartofSakarya University Journal of Computer and Information Sciences
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_Scopus_20260820
dc.subjectArtificial Intelligence In Education
dc.subjectAutomated Scoring
dc.subjectEvaluation Rubric
dc.subjectLarge Language Models (Llm)
dc.subjectValidity And Fairness
dc.titleEvaluating Large Language Models in Turkish Short Answer Scoring: Validity, Reliability, and Fairness Perspectives
dc.typeArticle

Dosyalar