Using Large Language Models for Automated Grading of Student Writing
Introduction
Full credit goes to the original author, linked below. All blog posts were reposted either with permission of the author, or by anonymous submission by SAIRC members like yourself.
A research digest of Impey et al. (2024), a University of Arizona study on GPT-4 grading of short student science writing. Notable findings include the comparability of instructor-created and LLM-generated rubrics — and that acceptable grading performance required giving the model both a rubric and an example answer.