§ 2027 Honours in the Tangen Lab

Can AI mark students' thinking?

In 2028, a new UQ course will ask hundreds of students to write by hand, in class, about the reasoning behind their own work. An AI-based system will help mark those responses. Before it's used in a real course, we need to know how well it works, where it can go wrong, and what safeguards are needed. In 2027, four honours students will each run a study answering one part of that question. This page describes those four projects. If you're starting honours in 2027, one of them could be yours.

SUPERVISOR  /  Prof. Jason Tangen, School of Psychology
SEATS  /  FOUR, ONE PER PROJECT
EXPRESSIONS OF INTEREST  /  j.tangen@uq.edu.au
A seminar table set for four, with four stacks of handwritten exam scripts, black pens, and a single orange pen
The lab table, set for fourRecruiting for 2027
01 The setup

Every student has spent fifteen years being marked.

§ 01

Four questions you already have opinions about.

Program The Longhand Program
Feeds Thinking with AI, 2028
Shares One corpus · one ethics approval · one pipeline

The lab is building a marking system called Longhand. It reads handwritten exam answers, produces clean transcripts, and helps mark the reasoning in them, with people checking its work at every step. From 2028 it will be used in a course that asks students to write in class about the reasoning behind their own work: by hand, with nothing open, four times in the semester.

Whether a system like that should be used is really four questions. Is the AI marking the quality of the thinking, or the quality of the writing? How good are people at catching its mistakes? Could students learn to fool it? And how do students feel about receiving a mark from an AI? Those are the four honours projects for 2027. Each one is a complete honours thesis, and each belongs to one student.

The results will be used. Each study is preregistered, with its standards set in advance, and together the four provide the evidence for deciding how (and whether) the system is used in the 2028 course. Some of the connections are direct: the moderation study shapes what moderators see on screen, and the validity study sets the standard the grader has to meet before it marks anything that counts.

One question cuts across all four projects

Is comparative judgement more trustworthy than absolute marking?

Instead of asking the AI to give each response a score on its own, comparative judgement asks it to decide which of two responses is better. All four projects test some part of this question. In March, each student will make a formal prediction about how comparative judgement will perform, and those predictions will be tested against the full set of results at the end of the year.

02 The 2027 projects

Pick the question that grabs you.

Two handwritten pages under a magnifying loupe: one immaculate, one a scrawl with crossings-out
Project 01 · VALValidity
01 / Validity

Is the AI marking the quality of the thinking, or the quality of the writing?

Fluent, confident writing can receive higher marks even when the reasoning itself is weak.

AI graders can be influenced by features that should not matter very much. Polished prose can make weak reasoning look stronger, rough phrasing can lower a mark, and longer answers often do better than shorter ones. If the system is supposed to assess the quality of students' thinking, this is one of the first things we need to test. It has also not been studied much with handwritten responses of the kind this course will use.

You'll build a set of responses in which reasoning quality and writing quality vary independently. Some will contain strong reasoning expressed in rough prose; others will contain weak reasoning expressed in polished prose. Expert raters who do not know the purpose of the manipulation will check every response before the grading study begins.

The same responses will also be copied out in both neat and messy handwriting, allowing you to test whether legibility affects the result. The AI and trained human markers will then grade the full set. You'll examine whether their marks follow the quality of the reasoning, the quality of the writing, the handwriting, or some combination of the three. You'll also test whether comparative judgement reduces any of these biases.

  • Start by judging twenty polished-but-weak and rough-but-sound pairs yourself, and measure whether fluency affects your own judgements.
  • Learn how to build and validate experimental materials using expert panels.
  • Test whether comparative judgement and a tuned grader reduce the bias, using a performance standard specified in advance.

Your study sets the standard the grader has to meet

A tall stack of marked scripts with one orange page-marker tab sticking out
Project 02 · MODModeration
02 / Moderation

How good are people at catching the AI's mistakes?

The usual safeguard is to have a person check the AI's marks. We know surprisingly little about how well that works.

When people review a machine's output, they sometimes accept errors they would not have made themselves. We do not yet know how often that happens in marking, or what helps people catch those errors. That matters because human moderation is one of the main safeguards proposed for AI-assisted marking.

You'll study this using a signal detection experiment, an approach the lab has used for many years in research with fingerprint examiners. Moderators will review scripts that have already been marked by an AI. A known proportion of those marks will contain deliberate errors.

If a moderator spots one of those errors, that counts as a hit. If they overturn a mark that was actually correct, that counts as a false alarm. Both matter. Missing an error is a problem, but unnecessary corrections also cost time and reduce the efficiency that AI-assisted marking is meant to provide.

You'll test whether people are better at detecting errors when they see different kinds of information about the AI's judgement. For example, they might see only the mark, the AI's explanation for the mark, or information about whether several independent AI judgements agreed. You'll also compare experienced tutors with novices.

  • Do the moderation task yourself in the first week and calculate your own hit and false-alarm rates.
  • Compare different ways of presenting the AI's judgement: the mark alone, its reasoning, or its consistency across several readings.
  • Learn signal detection theory as you need it, using your own data to make the ideas concrete.

Your study helps determine what moderators should see on screen

A handwritten exam script with an orange-backed playing card tucked under one corner, and two dice nearby
Project 03 · GAMEGaming
03 / Gaming

Can students learn to fool the marker?

If the grading system has weaknesses, we want to find them before students do.

The responses in this course will be handwritten, completed in class, and written from memory. Students will not be able to use external tools during the assessment. But they could still learn strategies in advance that make a weak response look stronger than it really is. They might write more, use formal-sounding language, sound unusually confident, or finish a weak argument with a strong-sounding conclusion.

Previous research has identified several biases that could make these strategies effective. This project asks whether they actually work under realistic assessment conditions.

You'll test them in two stages. First, you'll make controlled edits to real responses so that the effect of each strategy can be measured cleanly. Then you'll run a handwritten red-team study. Participants will be coached to try to increase their marks without improving the quality of their reasoning, and they will write under exam-like conditions.

Finally, you'll test possible safeguards, including comparative judgement, rubrics that limit the value of generic content, and several independent grading runs. By the end, you'll be able to say which strategies work, how much they change marks, and which safeguards reduce their effect.

  • Write a convincing account of work you have not actually done and use it to examine how easy it is to create an impression of understanding.
  • Run a coached red-team study under exam-like conditions. Some strategies may help, some may do nothing, and some may make the response worse.
  • Code responses for specific, verifiable detail using methods adapted from deception research.

Your study tests whether the assessment can be made resistant to gaming

A balance scale weighing a folded handwritten script against a small orange machine-printed slip
Project 04 · FAIRFairness
04 / Fairness

How do students feel about receiving a mark from an AI?

The other three projects ask whether the marks themselves can be trusted. This project asks how students respond to them.

Existing research does not give a simple answer. In some studies, people judge exactly the same feedback more harshly when they are told it came from an AI. In others, students see AI marking as fairer because they believe it is less likely to favour particular people.

The response may depend on features that a course can actually control. These include how the source of the mark is described, whether students can see the transcript the system used, whether a human has checked the result, and whether students have a meaningful way to appeal.

In this project, participants will receive real marks and feedback on their own writing. You'll vary whether they are told the mark came from a human, an AI, or an AI whose judgement was checked by a human. You'll also vary how the result is presented, for example as an individual score or as a comparison with other responses.

You'll measure perceived fairness, trust, willingness to appeal, and whether students actually use the feedback they receive.

  • Receive the same feedback yourself under two different source labels before you design the study.
  • Build your measures from procedural-justice research, including accuracy, voice and correctability.
  • Find out which design choices change how students respond to a mark, and what it means if none of them make much difference.

Your study helps determine how marks should be presented to students

03 The 2027 year, by design

The year is planned backwards from thesis submission.

§ 03

The groundwork is done before you arrive.

Corpus About 700 handwritten answers, already transcribed and expert-rated
Ethics One application covering all four projects
Pipeline Tested from beginning to end on UQ-governed infrastructure

Honours projects can lose a great deal of time to practical delays: ethics approval arriving late, materials being built in a rush, or data collection continuing well into second semester. This program is designed to avoid those problems.

Most of the shared groundwork will be completed before the 2027 honours year begins. Summer Research scholars will build the common corpus, the four projects will be covered by one ethics application, and the grading pipeline will already be working before orientation.

That changes the shape of the year. Each study can be specified early, the materials come from a corpus that already exists, and preregistration and data collection can be completed in Semester 1. The second half of the year is then available for analysis and writing.

SUMMER
Before day one
The lab's Summer Research scholars will collect about 700 handwritten responses from roughly 100 first-year students. They will scan the responses, check each transcript against the handwriting, and organise expert ratings of response quality. The marking guides, anchor sets and grading pipeline will also be completed and tested. By the time honours begins, the common materials and infrastructure will already be in place.
FEB
You experience the system first
The first day starts with the assessment itself rather than a reading list. You'll handwrite a short response of the same kind the 2028 students will write, and put it through the grading pipeline planned for that course. You'll receive the transcript, the grading levels and the feedback comment on the same kind of results screen students would see. Then you'll switch roles and hand-mark five responses yourself. By the end of the session, you'll have experienced the main issues behind all four projects from both sides.
MAR
Shared reading, your own project, and predictions
The four students begin with five papers that introduce the main ideas behind the program. They cover examples such as people mistaking machine-generated poems for work by recognised poets, the role of evaluative judgement in university learning, the effect of labelling feedback as AI-generated, experts accepting an incorrect machine judgement, and an automated licensing-exam marker giving credit for answers that copied wording from the question. Each student then moves into the literature for their own project. Your branch begins with a small self-experiment, and the methods you need are introduced as you begin designing the study rather than taught separately in advance. During March, all four students also make formal predictions about comparative judgement. Those predictions are recorded before the results are known and then tested against the full set of findings later in the year.
APR–MAY
Preregister, then collect the data
The analysis plan is finalised before the first participant is run. Each project uses materials drawn from the shared corpus, and data collection is designed to be finished by the end of May.
JUN–JUL
Analyse the results together
Winter is devoted to analysis, working both with the lab and with the other three honours students. Because the main analyses were preregistered, you begin with a clear plan for what to test rather than deciding on the analysis after seeing the results. The first full draft of the introduction is due on 30 July.
AUG–OCT
Write the thesis
The final part of the year is mainly for writing, with data collection already completed several months earlier. The thesis is due on 6 October 2027. After submission, the findings from all four projects will feed directly into decisions about how the marking system should be used in the 2028 course.
The corpus01

700handwritten answers

Written by
About 100 first-year students, over summer
Prepared
Scanned, transcribed, and checked against the handwriting
On day one
Ready and waiting for you
The structure02

4students, 4 projects

Your own
One complete project, written up as your thesis
Shared
The corpus, the ethics approval, and the pipeline
Weekly
All four meet at one table to compare findings
The deadlines03

MAYdata collection ends

30 July
Introduction draft due
6 October
Thesis due
After that
Your results shape the 2028 course
04 The lab

Four projects, one table.

§ 04

Linked projects, weekly meetings, formal predictions.

Meets Weekly, all four projects at one table
Lab record ~50 honours theses since 2007
Also in the lab Four PhD students, active collaborations

The four projects are wired together by design. The validity study's measured failure modes seed the moderation study's errors. The validity materials supply the gaming study's edit bases. The moderation study's winning display becomes the "checked by a human" condition the fairness study describes to its participants. Sharing what you're finding at the weekly meeting is part of the method, and it means each of you finishes the year fluent in all four studies.

The predictions made in March (will comparative judgement beat absolute scoring on validity, moderation, gaming, and fairness?) are tested against the full results later in the year, seat by seat.

You'd be joining a lab that has supervised roughly fifty honours theses since 2007 and has spent two decades studying expert judgement: fingerprint examiners, forensic scientists, first responders. This year, the judgement being studied is a marker's. Meet the lab →

05 Take a seat

If one of these questions grabbed you, say so.

§ 05

Start the conversation now.

Email j.tangen@uq.edu.au
Subject "Honours 2027" + the project's name
Formal applications Via the School of Psychology, for 2027 entry

Formal applications for 2027 honours run through the School of Psychology's usual process, but conversations can start earlier. Email a few lines about which project interests you and why. If you'd like to see the machinery before deciding, come by: the corpus, the pipeline, and the marking screens are all real, and I'm happy to walk you through them.

There's no prerequisite reading and no AI expertise required; the program was designed for students arriving fresh. If you've ever looked at a mark on your own work and wondered whether anyone really read it, that's enough to start with.