Autors: Vangelova A., Gancheva, V. S.
Title: LLMs in Automated Assessment: A Role-Based Taxonomy and Framework for Controlled Educational Integration
Keywords: automated assessment, large language models, open-ended questions, RAG, semantic analysis

Abstract: Large language models (LLMs) are reshaping the automated assessment of open-ended student responses. Compared with earlier rule-based, statistical, and feature-engineered approaches, they enable a deeper interpretation of meaning, context, and argumentation. This development can be understood as a fifth generation of automated scoring systems, but it also raises a new question: not only what LLMs can do, but also how they can be deployed in education in a controlled and reliable manner. This paper presents a role-based taxonomy that distinguishes between generative LLMs used as direct virtual graders, encoder transformers used as semantic tools, and intermediate text-to-text models used in more formalized assessment tasks. It also discusses the main limitations of standalone LLM graders, including hallucinations, probabilistic instability, limited interpretability, bias, and weak grounding in domain-specific content. To address these issues, the paper presents a developed framework implemented in an integrated assessment system built on role prompting, rubric-constrained grading, Retrieval-Augmented Generation (RAG), structured machine-readable outputs, workflow orchestration, and LMS integration. The framework is further extended to multimodal assessment through vision-based evaluation of visual artifacts such as UML state diagrams. The main contribution of the paper is not only a conceptual framework, but also its realization in a working integrated system for automated assessment in a more traceable, pedagogically grounded, and institutionally reliable way.

References

  1. Gao R. Merzdorf H.E. Anwar S. Hipwell M.C. Srinivasa A.R. Automatic Assessment of Text-Based Responses in Post-Secondary Education: A Systematic Review Comput. Educ. Artif. Intell. 2024 6 100206 10.1016/j.caeai.2024.100206
  2. Rodrigues L. Xavier C. Costa N.T. Gašević D. Mello R.F. Is GPT-4 Fair? An Empirical Analysis in Automatic Short Answer Grading Comput. Educ. Artif. Intell. 2025 8 100428 10.1016/j.caeai.2025.100428
  3. Tan L.Y. Hu S. Yeo D.J. Cheong K.H. A Comprehensive Review on Automated Grading Systems in STEM Using AI Techniques Mathematics 2025 13 2828 10.3390/math13172828
  4. Sembey R. Hoda R. Grundy J. Emerging Technologies in Higher Education Assessment and Feedback Practices: A Systematic Literature Review J. Syst. Softw. 2024 211 111988 10.1016/j.jss.2024.111988
  5. Sun J. Song T. Peng W. Song J. A Survey of Automated Essay Scoring: Challenges, Advances, and Future Neurocomputing 2025 650 130916 10.1016/j.neucom.2025.130916
  6. Tang X. Chen H. Lin D. Li K. Harnessing LLMs for Multi-Dimensional Writing Assessment: Reliability and Alignment with Human Judgments Heliyon 2024 10 e34262 10.1016/j.heliyon.2024.e34262 39113951
  7. Jauhiainen S.J. Garagorry Guerra A. Evaluating Students’ Open-Ended Written Responses with LLMs: Using the RAG Framework for GPT-3.5, GPT-4, Claude-3, and Mistral-Large Adv. Artif. Intell. Mach. Learn. 2024 4 3097 3113 10.54364/AAIML.2024.44177
  8. Emirtekin E. Large Language Model-Powered Automated Assessment: A Systematic Review Appl. Sci. 2025 15 5683 10.3390/app15105683
  9. Jacobsen L.J. Weber K.E. The Promises and Pitfalls of Large Language Models as Feedback Providers: A Study of Prompt Engineering and the Quality of AI-Driven Feedback AI 2025 6 35 10.3390/ai6020035
  10. Mendonça P.C. Quintal F. Mendonça F. Evaluating LLMs for Automated Scoring in Formative Assessments Appl. Sci. 2025 15 2787 10.3390/app15052787
  11. Nkoyo T.A.F.E. Ijezue C.F. Amjad A.I. Amjad M. Butt S. Castañeda-Garza G. Advances in Auto-Grading with Large Language Models: A Cross-Disciplinary Survey Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2025), Vienna, Austria, 31 July–1 August 2025 Association for Computational Linguistics Stroudsburg, PA, USA 2025 477 498 10.18653/v1/2025.bea-1.35
  12. García-Varela F. Nussbaum M. Mendoza M. Martínez-Troncoso C. Bekerman Z. ChatGPT as a Stable and Fair Tool for Automated Essay Scoring Educ. Sci. 2025 15 946 10.3390/educsci15080946
  13. Grévisse C. LLM-Based Automatic Short Answer Grading in Undergraduate Medical Education BMC Med. Educ. 2024 24 1060 10.1186/s12909-024-06026-5 39334087
  14. Seßler K. Fürstenberg M. Bühler B. Kasneci E. Can AI Grade Your Essays? A Comparative Analysis of Large Language Models and Teacher Ratings in Multidimensional Essay Scoring Proceedings of the 15th International Learning Analytics and Knowledge Conference Dublin, Ireland 24–28 March 2025 462 472 10.1145/3706468.3706527
  15. Paranjape B. Lundberg S. Singh S. Hajishirzi H. Zettlemoyer L. Ribeiro M.T. ART: Automatic Multi-Step Reasoning and Tool-Use for Large Language Models arXiv 2023 10.48550/arXiv.2303.09014 2303.09014
  16. Singh J. Magazine R. Pandya Y. Nambi A.U. Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning arXiv 2025 2505.01441
  17. Vangelova A. Gancheva V. AI-Based Automated Scoring Layer Using Large Language Models and Semantic Analysis Appl. Sci. 2026 16 3537 10.3390/app16073537
  18. Sychev O. Anikin A. Prokudin A. Automatic Grading and Hinting in Open-Ended Text Questions Cogn. Syst. Res. 2020 59 264 272 10.1016/j.cogsys.2019.09.025
  19. Shermis M.D. Burstein J. Apel Bursky S. Introduction to Automated Essay Evaluation Handbook of Automated Essay Evaluation: Current Applications and New Directions Shermis M.D. Burstein J. Routledge/Taylor & Francis Group New York, NY, USA 2013 1 15
  20. Dikli S. An Overview of Automated Scoring of Essays J. Technol. Learn. Assess. 2006 5 1 36 Available online: https://ejournals.bc.edu/index.php/jtla/article/view/1640 (accessed on 18 December 2025)
  21. Mizumoto A. Eguchi M. Exploring the Potential of Using an AI Language Model for Automated Essay Scoring Res. Methods Appl. Linguist. 2023 2 100050 10.1016/j.rmal.2023.100050
  22. Attali Y. Burstein J. Automated Essay Scoring with e-rater®V.2 J. Technol. Learn. Assess. 2006 4 3
  23. Birla N. Kumar Jain M. Panwar A. Automated Assessment of Subjective Assignments: A Hybrid Approach Expert Syst. Appl. 2022 203 117315 10.1016/j.eswa.2022.117315
  24. Pecuchova J. Benko Ľ. Drlik M. Automated Grading of Open-Ended Questions in Higher Education Using GenAI Models Int. J. Artif. Intell. Educ. 2025 35 3813 3846 10.1007/s40593-025-00517-2
  25. Cisneros-González J. Gordo-Herrera N. Barcia-Santos I. Sánchez-Soriano J. JorGPT: Instructor-Aided Grading of Programming Assignments with Large Language Models (LLMs) Future Internet 2025 17 265 10.3390/fi17060265
  26. Cipriano E. Ferrato A. Limongelli C. Schicchi D. Taibi D. Leveraging Large Language Models to Assist Teachers in Code Grading Artificial Intelligence in Education Cristea A.I. Walker E. Lu Y. Santos O.C. Isotani S. Springer Nature Cham, Switzerland 2025 Volume 15880 204 217 10.1007/978-3-031-98459-4_15
  27. Bernik A. Radošević D. Čep A. A Comparative Study of Large Language Models in Programming Education: Accuracy, Efficiency, and Feedback in Student Assignment Grading Appl. Sci. 2025 15 10055 10.3390/app151810055
  28. Papachristou I. Dimitroulakos G. Vassilakis C. Automated Test Generation and Marking Using LLMs Electronics 2025 14 2835 10.3390/electronics14142835
  29. Ndukwe I.G. Amadi C.E. Nkomo L.M. Daniel B.K. Automatic Grading System Using Sentence-BERT Network Artificial Intelligence in Education Bittencourt I.I. Cukurova M. Muldner K. Luckin R. Millán E. Springer International Publishing Cham, Switzerland 2020 Volume 12164 224 227 10.1007/978-3-030-52240-7_41
  30. Dada I.D. Akinwale A.T. Tunde-Adeleke T.-J. A Structured Dataset for Automated Grading: From Raw Data to Processed Dataset Data 2025 10 87 10.3390/data10060087
  31. Balakrishnan R.M. Pati P.B. Singh R.P. S S. Kumar P. Fine-Tuned T5 for Auto-Grading of Quadratic Equation Problems Procedia Comput. Sci. 2024 235 2178 2186 10.1016/j.procs.2024.04.206
  32. Li J. Gui L. Zhou Y. West D. Aloisi C. He Y. Distilling ChatGPT for Explainable Automated Student Answer Assessment arXiv 2023 10.48550/arXiv.2305.12962 2305.12962
  33. Pack A. Barrett A. Escalante J. Large Language Models and Automated Essay Scoring of English Language Learner Writing: Insights into Validity and Reliability Comput. Educ. Artif. Intell. 2024 6 100234 10.1016/j.caeai.2024.100234
  34. Lee G.-G. Latif E. Wu X. Liu N. Zhai X. Applying Large Language Models and Chain-of-Thought for Automatic Scoring Comput. Educ. Artif. Intell. 2024 6 100213 10.1016/j.caeai.2024.100213
  35. Organisciak P. Acar S. Dumas D. Berthiaume K. Beyond Semantic Distance: Automated Scoring of Divergent Thinking Greatly Improves with Large Language Models Think. Ski. Creat. 2023 49 101356 10.1016/j.tsc.2023.101356
  36. Chu S. Kim J. Wong B. Yi M. Rationale Behind Essay Scores: Enhancing S-LLM’s Multi-Trait Essay Scoring with Rationale Generated by LLMs arXiv 2025 10.48550/arXiv.2410.14202 2410.14202
  37. Seneviratne H.M.T.W. Manathunga S.S. Artificial Intelligence Assisted Automated Short Answer Question Scoring Tool Shows High Correlation with Human Examiner Markings BMC Med. Educ. 2025 25 1146 10.1186/s12909-025-07718-2 40764994
  38. Oğuz E. Can Generative AI Figure Out Figurative Language? The Influence of Idioms on Essay Scoring by ChatGPT, Gemini, and Deepseek Assess. Writ. 2025 66 100981 10.1016/j.asw.2025.100981
  39. Xu W. Kassim M.S.S. Hoo W.L. Yang W. Xu T. Explainable AI for Education: Enhancing Essay Scoring via Rubric-Aligned Chain-of-Thought Prompting Int. J. Mod. Phys. C 2026 37 2542013 10.1142/S0129183125420136
  40. Kinder A. Briese F.J. Jacobs M. Dern N. Glodny N. Jacobs S. Leßmann S. Effects of Adaptive Feedback Generated by a Large Language Model: A Case Study in Teacher Education Comput. Educ. Artif. Intell. 2025 8 100349 10.1016/j.caeai.2024.100349
  41. Seo H. Hwang T. Jung J. Kang H. Namgoong H. Lee Y. Jung S. Large Language Models as Evaluators in Education: Verification of Feedback Consistency and Accuracy Appl. Sci. 2025 15 671 10.3390/app15020671
  42. Landis J.R. Koch G.G. The Measurement of Observer Agreement for Categorical Data Biometrics 1977 33 159 174 10.2307/2529310
  43. Taghipour K. Ng H.T. A Neural Approach to Automated Essay Scoring Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Austin, TX, USA, 1–5 November 2016 Association for Computational Linguistics Stroudsburg, PA, USA 2016 1882 1891 10.18653/v1/D16-1193
  44. Ludwig S. Mayer C. Hansen C. Eilers K. Brandt S. Automated Essay Scoring Using Transformer Models Psych 2021 3 897 915 10.3390/psych3040056
  45. Villegas-Ch W. Gutierrez R. García-Ortiz J. Guevara V. Explainable Educational Assistant Integrated in Moodle: Automated Semantic Assessment and Adaptive Tutoring Based on NLP and XAI Discov. Artif. Intell. 2025 5 191 10.1007/s44163-025-00438-y
  46. Plyer L. Marcou G. Perves C. Bonachera F. Varnek A. Implementation of a Soft Grading System for Chemistry in a Moodle Plugin: Reaction Handling J. Cheminform. 2024 16 90 10.1186/s13321-024-00889-y 39090756
  47. Sychev O. Questions for Teaching Phrase Building with Automatic Feedback Softw. Impacts 2023 15 100461 10.1016/j.simpa.2022.100461
  48. Gradescope Guides. Using Gradescope LTI 1.3 with Moodle as an Instructor Available online: https://guides.gradescope.com/hc/en-us/articles/23587719150349-Using-Gradescope-LTI-1-3-with-Moodle-as-an-Instructor (accessed on 24 April 2026)
  49. Gradescope Guides. Grading Submissions with Rubrics Available online: https://guides.gradescope.com/hc/en-us/articles/22249389005709-Grading-submissions-with-rubrics (accessed on 24 April 2026)
  50. Turnitin Guides. Using the AI Writing Report Available online: https://guides.turnitin.com/hc/en-us/articles/22774058814093-Using-the-AI-Writing-Report (accessed on 24 April 2026)
  51. Barenji R.V. Salimi N. Khoshgoftar S. An LLM-Powered Assessment Retrieval-Augmented Generation (RAG) for Higher Education arXiv 2026 10.48550/arXiv.2601.06141 2601.06141

Issue

Applied Sciences (Switzerland), vol. 16, 2026, Switzerland, https://doi.org/10.3390/app16136617

Вид: статия в списание, публикация в издание с импакт фактор, публикация в реферирано издание, индексирана в Scopus