Is the AI marking the quality of the thinking, or the quality of the writing?
Fluent, confident writing can receive higher marks even when the reasoning itself is weak.
AI graders can be influenced by features that should not matter very much. Polished prose can make weak reasoning look stronger, rough phrasing can lower a mark, and longer answers often do better than shorter ones. If the system is supposed to assess the quality of students' thinking, this is one of the first things we need to test. It has also not been studied much with handwritten responses of the kind this course will use.
You'll build a set of responses in which reasoning quality and writing quality vary independently. Some will contain strong reasoning expressed in rough prose; others will contain weak reasoning expressed in polished prose. Expert raters who do not know the purpose of the manipulation will check every response before the grading study begins.
The same responses will also be copied out in both neat and messy handwriting, allowing you to test whether legibility affects the result. The AI and trained human markers will then grade the full set. You'll examine whether their marks follow the quality of the reasoning, the quality of the writing, the handwriting, or some combination of the three. You'll also test whether comparative judgement reduces any of these biases.
- Start by judging twenty polished-but-weak and rough-but-sound pairs yourself, and measure whether fluency affects your own judgements.
- Learn how to build and validate experimental materials using expert panels.
- Test whether comparative judgement and a tuned grader reduce the bias, using a performance standard specified in advance.
Your study sets the standard the grader has to meet