Skip to main navigation Skip to search Skip to main content

Validating human and automated scoring of essays against “True” scores

  • Yoav Cohen*
  • , Effi Levi
  • , Anat Ben-Simon
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

20 Scopus citations

Abstract

In the current study, two pools of 250 essays, all written as a response to the same prompt, were rated by two groups of raters (14 or 15 raters per group), thereby providing an approximation to the essay’s true score. An automated essay scoring (AES) system was trained on the datasets and then scored the essays using a cross-validation scheme. By eliminating one, two, or three raters at a time, and by calculating an estimate of the true scores using the remaining raters, an independent criterion against which to judge the validity of the human raters and that of the AES system, as well as the interrater reliability was produced. The results of the study indicated that the automated scores correlate with human scores to the same degree as human raters correlate with each other. However, the findings regarding the validity of the ratings support a claim that the reliability and validity of AES diverge: although the AES scoring is, naturally, more consistent than the human ratings, it is less valid.

Original languageEnglish
Pages (from-to)241-250
Number of pages10
JournalApplied Measurement in Education
Volume31
Issue number3
DOIs
StatePublished - 3 Jul 2018

Bibliographical note

Publisher Copyright:
© 2018 Taylor & Francis.

Fingerprint

Dive into the research topics of 'Validating human and automated scoring of essays against “True” scores'. Together they form a unique fingerprint.

Cite this