Part oneThe man with the stopwatch
Amsterdam, 1946. A Dutch psychologist named Adriaan de Groot wants to understand what makes expert chess players so good. He’s also one of the strongest players in the Netherlands, so he knows the game well enough to study the question properly.
His question is straightforward: what are grandmasters doing differently?
The usual explanation was that grandmasters calculate further ahead and have unusually good memories. De Groot initially assumed the same thing.
So he ran the obvious test. He showed players a position from a real game. Twenty-odd pieces, mid-battle. Five seconds. Then he took it away and asked them to rebuild it from memory.
The grandmasters rebuilt it almost perfectly. Club players managed four or five pieces. That result seemed to support the memory explanation.
The important follow-up was run in 1973 by Bill Chase and Herbert Simon.
They kept everything the same. Same board, same pieces, same five seconds. One change. They took the pieces and scattered them at random. Positions that could never occur in a real game.
The grandmasters dropped to the level of the beginners.
The result wasn’t evidence of photographic memory. Experienced players had learned a very large number of familiar chess patterns—estimates are around fifty thousand. In a real position, a grandmaster doesn’t process every piece separately. They recognise several familiar configurations. In a random position, those configurations aren’t there, so the advantage disappears.
So the difference was specialised knowledge, not better memory in general. It had been built over decades, and much of it operated automatically. Grandmasters could describe parts of their reasoning, but they couldn’t give a complete account of how they recognised a position.
You did this at breakfast
Reading gives us a familiar example of the same process.
When you read the news this morning, you didn’t consciously identify each letter and assemble each word. Most of that processing happened automatically. A fluent reader can explain some reading strategies, but not every operation involved in recognising a sentence.
That fluency still took years to develop: phonics, spelling, comprehension, and a great deal of practice reading aloud. Over time, children stop having to concentrate on each component and can focus on meaning instead.
Once the process becomes fluent, it is easy to forget how much practice it required.
Cognitive scientists call these familiar units chunks. Research on chess, radiology, music, firefighting and fingerprint examination shows a similar pattern: repeated practice allows people to process several related elements as a single familiar unit. Experts then use those units quickly, often without being able to describe every step.
Why exams worked
This creates a practical problem for education.
If we cannot observe learning directly, how do we assess and certify it?
We usually infer learning from performance: the rebuilt chessboard, the essay, the exam response, or the answer to a follow-up question. We cannot observe the underlying knowledge directly, but we can examine what a person produces with it.
Exams and qualifications rely on that relationship. Board papers often do as well: the document is treated as evidence that the author has understood the issue and worked through the relevant evidence.
Until recently, producing a good essay generally required the writer to understand the material, organise an argument and make decisions about evidence.
Generative AI weakens that connection between the quality of the document and the understanding of the person submitting it.
A student can now produce a plausible essay very quickly without doing all of that work. On the page, it may be difficult to tell how much the student understood or contributed.
I want to focus on what that means for learning, rather than only for cheating or detection. The work involved in writing essays and solving problems helps students develop three kinds of judgement: distinguishing strong work from weak work, deciding when work is good enough, and judging their own understanding.
The map
I’ll take those in order. First, how people learn to distinguish one thing from another. Second, how they decide what counts as good enough. Third, how accurately they judge their own knowledge. We’ll pause for questions after each section, and then look at practical implications for schools and boards.
I’ll start with fingerprint examination, which is the area I’ve studied for the past twenty years, and with the difference between the television version and the real process.
Part twoThe examiner’s eye
The CSI myth
On television, a technician scans a fingerprint, the database produces a name, and the result appears as a confirmed match. It looks like a largely automated process.
In practice, a crime-scene print is often partial, smudged or distorted. Skin does not leave exactly the same impression every time. An examiner compares two images and decides whether the similarities are sufficient to support an identification. They can point to relevant features and explain much of the comparison, but some of their pattern recognition comes from experience and is difficult to put into words.
This is another example of learned pattern recognition. Our research shows that qualified fingerprint examiners are very good at distinguishing matching prints from similar non-matches. It also shows that their decisions can be affected by the same limitations that affect other forms of human judgement. The Mayfield case is a well-known example.
Madrid
On the eleventh of March, 2004, bombs go off on four commuter trains in Madrid. A hundred and ninety-three people are killed. Spanish police find a partial fingerprint on a bag of detonators. This is now the largest terrorism investigation in Europe, so the print goes out to agencies around the world, including the FBI.
The FBI runs it through their database and lands on an American: a lawyer in Oregon named Brandon Mayfield.
A senior FBI examiner compares the prints and declares a match. A second examiner verifies it. A third concurs. Mayfield is arrested. And when his defence team asks for an independent expert, the court appoints one, a well-regarded examiner with no ties to the FBI. He examines the prints, and he confirms the match too.
Four examiners reached the same conclusion.
Spanish police continued to disagree with the identification. A few weeks later, they identified the actual source of the print as an Algerian national named Ouhnane Daoud. The FBI released Mayfield, apologised publicly and later paid him two million dollars.
It would be easy to explain the error as carelessness or misconduct, but the evidence points to a more ordinary problem. Context can influence expert judgement: the seriousness of the case, knowledge that another examiner has reached a conclusion, or the way a candidate print is presented. Because those influences are not always conscious, examiners may remain confident even when their judgement has been affected.
The important point is that confidence does not necessarily tell us whether a judgement has been influenced.
What the eye actually is
When fingerprint expertise works well, the examiner distinguishes relevant detail from distortion and coincidental similarity. That ability develops through years of comparing prints and receiving feedback. A lecture can provide principles and terminology, but it cannot replace the comparisons needed to develop the skill.
In cognitive science, this is called discrimination: learning to distinguish categories that initially look similar.
Students develop discrimination when they compare examples, make judgements and receive feedback. In writing, that means learning to recognise whether an argument is clear, whether evidence is relevant and whether one explanation is stronger than another. AI-generated text gives students a new type of material to evaluate, and they may not yet have enough experience to evaluate it reliably.
A recent medical study illustrates the problem.
Fifty doctors
2024. A randomised trial, published in JAMA Network Open. Fifty experienced physicians, family medicine, internal medicine, emergency medicine, are given difficult diagnostic cases. Half work with their usual resources. Half get their usual resources plus GPT-4.
The researchers expected that access to GPT-4 might improve the doctors’ performance.
Doctors alone scored 74 per cent. Doctors with GPT-4: 76. Statistically, that’s a tie.
Then the researchers ran the model on its own, no doctor. Almost 90.
The doctors did not consistently use the model’s better answers to improve their own performance.
The doctors had extensive diagnostic training, but that training focused on patients, clinical histories and examinations. It did not necessarily prepare them to decide when a fluent model-generated explanation should override their own view. When the model disagreed with them, they did not always identify which answer was better.
The relevance to a board is that experience in budgets, contracts, curriculum or staffing does not automatically confer expertise in evaluating AI-generated analysis. As AI-written or AI-edited material enters board papers, the fact that it reads well should not be treated as evidence that its reasoning is sound.
What this asks of you
There are three practical responses.
In the classroom, put two essays on the same question side by side and ask students which is stronger and why. Have them explain their decision aloud, then give them feedback on the judgement. For boards and committees, periodically read a complete underlying document—a budget or safeguarding report, for example—rather than relying only on summaries. In both cases, the relevant skill is developed by working directly with the material.
That’s the first ability: discrimination. Before I move on, what questions or objections do you have?
Part threeThe control room
England, the 1970s
The second ability is easier to explain through research on automation.
In the 1970s, Lisanne Bainbridge studied control-room operators in settings such as chemical plants and power stations. Automated systems handled routine adjustments, while operators monitored the system and intervened when something unusual happened. The arrangement appeared sensible: automate routine work and leave exceptions to people.
In her 1983 paper Ironies of Automation, Bainbridge pointed out a problem with that arrangement. When operators no longer perform routine tasks, they get less practice. Yet the situations in which they must intervene are usually the most difficult ones. Automation can therefore reduce the experience people need for the work that remains.
Her summary is still useful, so I’ll read it directly:
“By taking away the easy parts of his task, automation can make the difficult parts of the human operator’s task more difficult.”
The same problem can arise when AI removes routine practice from education.
Her prediction is running an hour from here
Fingerprint examination provides a current example.
Queensland’s fingerprint bureau uses an automated process often called lights-out latents. An algorithm handles easier comparisons without an examiner, while more difficult cases are sent to people.
That changes the mix of cases available to trainees. Straightforward comparisons used to provide early practice. If those cases are automated, new examiners begin with a greater proportion of difficult material. As the system improves, the remaining human caseload may become more difficult again.
The efficiency gain may therefore create a training problem unless the organisation provides practice in another way.
The threshold for automation is usually chosen to meet operational goals such as speed and accuracy. Training effects also need to be part of that decision.
The easy part and the hard part
The same issue applies to writing. Producing a first draft is only part of the task. Students also need to decide whether the argument works, whether the evidence is sufficient and whether the piece is ready to submit. They develop that judgement by drafting, revising and receiving feedback many times.
If AI produces the draft, students may be asked to evaluate a polished text without having had enough experience creating and revising one themselves.
The second half of the examiner’s judgement
Fingerprint decisions also involve a second skill.
After examiners identify relevant similarities, they must decide whether the evidence is sufficient to report a match. Two competent examiners can see the same details and still reach different conclusions because they require different amounts of evidence.
The evidence does not determine that threshold on its own. A lower threshold increases the risk of a false identification; a higher threshold increases the risk of missing a genuine match. The choice therefore depends partly on how the organisation weighs those two errors.
In cognitive science, this threshold is called a criterion. Students make similar decisions when they ask whether they have enough evidence, whether a response is complete, or whether a piece of work is ready to submit. They improve those decisions through repeated practice and feedback.
AI-generated work often arrives polished and well structured. If students mainly receive finished-looking text, they get less practice deciding what needs revision and when a piece of work is actually complete.
At your table
Boards face a related issue. Capital works proposals, fee recommendations and vendor submissions usually arrive in a finished format. The quality of the presentation can influence whether the evidence feels sufficient, even though presentation quality and evidentiary quality are different things. AI makes polished presentation easier to produce.
One useful response is to specify the decision criteria before reading a consequential proposal. What evidence is required? What would rule the proposal out? Writing those conditions down makes it easier to assess the submission against an independent standard.
In the classroom, ask students to describe the criteria before they assess a draft, whether it is their own or AI-generated. Assess the quality of those criteria as well as the final work. For significant board decisions, record the criteria in advance and periodically review whether they were appropriate for the risks involved.
That’s the second ability: setting a criterion. I’ll pause again for questions.
Part fourThe best essays in the room
The experiment that should have been good news
The third section begins with a study in which AI improved the students’ essays.
In 2025, the British Journal of Educational Technology published a study of 117 university students. They worked in a lab, with every click and keystroke recorded. Everyone wrote the same two-hour essay from the same materials, then had an hour to revise it. That’s where the experiment happened. Some revised alone, some with ChatGPT, some with a human writing tutor, and some with an automated feedback checklist. One design detail matters: ChatGPT was allowed to advise, but not to write the essay for them.
The ChatGPT group produced the strongest essays. On that outcome alone, the tool looked beneficial.
Because Fan and colleagues recorded the revision process, they could also examine how the students worked. The ChatGPT group interacted repeatedly with the chatbot, but did much less planning, checking and evaluation of their own work. The essays improved while students engaged in less self-monitoring. This happened even though the chatbot was configured to advise rather than write the essay.
Fan and colleagues described this pattern as metacognitive laziness: relying on the tool to guide the process rather than monitoring the work independently.
A small MIT preprint reported a related result. It involved 54 students, so I would treat it as preliminary. Students wrote essays with or without AI while wearing EEG caps. The AI-assisted group showed lower measures of brain engagement, and 83 per cent could not quote a sentence from their essay shortly after finishing. The result suggests that producing a document with AI does not necessarily involve close engagement with its content.
The surprise test
The next question is whether these process differences affect later learning.
In a 2025 study in Rio de Janeiro, 120 business students were randomly assigned to research a topic either with ChatGPT or with books, articles and ordinary search. They had two weeks to prepare a short presentation for their peers. Forty-five days later, they returned for a test they had not been told about in advance.
The group that studied unaided: 68.5 per cent. The ChatGPT group: 57.5.
Both groups had reported similar confidence during the original task. Their confidence therefore did not reflect the later difference in retention. A separate study examined whether students could judge the quality of their work at the time they produced it.
Urban and colleagues tested 145 university students in Prague in 2024. Students acted as consultants to a toy company and proposed three ways to make a stuffed bunny outsell Lego. Half used ChatGPT and half did not. Two experts, who did not know which condition each student was in, scored the solutions. Students also rated their own work.
The ChatGPT group produced stronger and more original solutions. They also reported that the task felt easier, required less effort and left them more confident.
However, within the ChatGPT group, students’ self-ratings were essentially unrelated to the expert scores. The correlation was close to zero. Students who found ChatGPT more useful also tended to overestimate their performance by more.
In this study, feeling confident while using the tool was not a reliable indicator of solution quality.
This third ability is metacognition: accurately judging what you understand, what you are uncertain about and when you need to check. If AI improves the submitted work without improving the student’s understanding, grades and feedback may become harder for the student to interpret.
The invented reasons
An older experiment shows why people’s explanations of their own decisions need to be treated carefully.
In a 2005 study in Lund, Petter Johansson and colleagues showed participants pairs of photographs and asked which face they found more attractive. The researcher then handed them the selected photograph and asked them to explain the choice.
On some trials, a concealed card-switch meant that the researcher actually handed over the photograph the participant had rejected.
Most participants did not notice the swap. Only about one in eight detected it immediately, and roughly three quarters of the swaps went undetected even under a broader measure. Participants then gave detailed reasons for choosing the photograph they had actually rejected, sometimes referring to features that were not present in their original choice.
The participants were not deliberately misleading the researcher. They generated plausible explanations without recognising that the photograph had changed. This shows that a confident explanation does not guarantee accurate access to the process that produced a decision.
That matters when we ask people to describe how critically they worked with AI.
The room’s version
The next study involved working adults rather than students.
In 2025, researchers from Microsoft and Carnegie Mellon asked 319 knowledge workers to describe three AI-assisted tasks from their jobs: creating something, finding information or getting advice. Across 936 tasks, participants reported where they had used critical thinking and where the effort had shifted. They spent less effort producing material and more effort checking AI output. Greater trust in the tool was associated with less reported critical thinking, while confidence in their own ability was associated with more.
This is relevant to boards because much of a board’s work involves reviewing material produced by other people. Effective review depends on current domain knowledge and on understanding how the underlying work is done. If AI changes that underlying process, boards need more than a general instruction to check the output.
What this asks of you
In the classroom, ask students to identify which parts of a substantial piece they are confident about and where they are uncertain. Compare predicted marks with actual marks over time. Ask students to explain their own document without referring to it. These activities provide direct evidence about how accurately they are monitoring their understanding.
At board level, make it acceptable to say, “I haven’t understood this.” For major decisions, ask members to state their level of confidence and what that confidence is based on. This helps distinguish confidence based on relevant evidence from confidence based on a clear and polished presentation.
That’s the third ability: metacognition. I’ll pause for questions, then bring the three abilities together and discuss possible responses.
Part fiveThe lift
Their word
The three sections describe related forms of judgement.
Discrimination is the ability to distinguish relevant differences. Criterion is the threshold used to decide what counts as sufficient. Metacognition is the ability to judge your own understanding. All three depend on practice and feedback, and all three can be affected when AI changes who performs the underlying work.
PMSA often brings these forms of judgement together under the word discernment.
In this context, discernment is not a single general capacity. It includes learned abilities to notice differences, apply appropriate standards and monitor one’s own knowledge. That gives schools something more specific to protect and assess.
The Turkish experiment
A study in a Turkish high school shows how the design of an AI tool can change its effect on learning.
The study took place in a large high school in Turkey during the 2023 school year. Bastani and colleagues at the University of Pennsylvania later published it in PNAS. Nearly a thousand students, in Years 9 to 11, took part in four ordinary ninety-minute maths lessons. The teacher reviewed a topic, the class practised problems, and then everyone sat a short closed-book exam alone. Both practice and exam contributed to real grades.
Whole classrooms were randomised into three groups. One group used textbooks and notes. Another used a GPT-4 chatbot—essentially ChatGPT. The third used a version the researchers had built with the school’s maths teachers: the same model, but redesigned to behave like a tutor. It would not simply give the answer. It offered hints, asked questions and moved one step at a time, using the teachers’ worked solutions and common mistakes behind the scenes.
During practice, the chatbot classrooms scored 48 per cent higher than the textbook classrooms. The tutor group scored 127 per cent higher. Both AI conditions improved performance while help was available.
On the closed-book exam, the standard-chatbot classrooms scored 17 per cent below the textbook classrooms. The tutor group performed at about the same level as the textbook group. The standard chatbot improved assisted practice performance but was followed by lower independent performance; the tutor improved practice performance without the same penalty.
The interaction records help explain the difference. Students often asked the standard chatbot for answers and copied what it produced. When asked directly for answers, the chatbot made logical errors on 42 per cent of the problems. The tutor was designed to provide hints and questions rather than complete answers, so students had to continue working through the problem.
Students’ perceptions were not reliable indicators of these effects. The chatbot group did not report learning less, and the tutor group overestimated its exam performance.
The underlying model was the same. What changed was how students were allowed to use it. That design choice is something a school can govern.
What good use looks like
A useful design principle is to use AI to prompt further thinking rather than replace the task students need to practise.
For example, a student could finish an essay and then ask the model to identify a weak argument. The student would still need to decide whether the criticism is valid and revise or defend the passage. After reading a difficult chapter, a student could use the model for retrieval questions and note where they cannot answer without help. In both cases, the student continues to perform the writing, recall and judgement being learned.
Boards can use the same principle. A model might identify claims in a vendor proposal that need closer examination, suggest information missing from a staff paper, or generate possible failure scenarios for a proposed decision. Board members and staff would then verify those suggestions against the evidence. The tool is being used to expand the review, not to make the decision.
There are two governance implications. First, students, staff and board members need examples and training in appropriate use; good practice is not self-evident. Second, the capabilities and limitations of these tools change frequently. The institution needs clear responsibility for reviewing its guidance and evidence on a regular schedule. That responsibility extends beyond technical maintenance.
Two children
I’ll finish with the staircase image used in the slides.
Two children reach the top floor. One used the stairs and one used the lift. At the top, both can see the same view.
If we look only at the final position, we cannot tell which route they took.
But only one has practised climbing. If the next task requires that capacity, the route matters.
An essay, diagnosis or board paper is an output. We also need to ask what knowledge and judgement were developed while producing it. In the terms I’ve used today, that includes discrimination, criterion and metacognition. PMSA describes these together as discernment. AI can bypass some of the practice that develops them, or it can be designed to support that practice.
The outcome depends on the choices schools make about when and how the technology is used.
That is what I mean by saying that the climbing was the point. Thank you.
Sources
- de Groot, A. D. (1965). Thought and Choice in Chess. Mouton. (Original studies, 1938–1946.)
- Chase, W. G., & Simon, H. A. (1973). Perception in chess. Cognitive Psychology. doi.org/10.1016/0010-0285(73)90004-2
- Ericsson, K. A., Chase, W. G., & Faloon, S. (1980). Acquisition of a memory skill. Science. doi.org/10.1126/science.7375930
- Castles, A., Rastle, K., & Nation, K. (2018). Ending the reading wars: Reading acquisition from novice to expert. Psychological Science in the Public Interest. doi.org/10.1177/1529100618772271
- Green, D. M., & Swets, J. A. (1966). Signal Detection Theory and Psychophysics. Wiley.
- U.S. Department of Justice, Office of the Inspector General (2006). A Review of the FBI’s Handling of the Brandon Mayfield Case.
- Goh, E., et al. (2024). Large language model influence on diagnostic reasoning: A randomized clinical trial. JAMA Network Open. doi.org/10.1001/jamanetworkopen.2024.40969
- Bainbridge, L. (1983). Ironies of automation. Automatica. doi.org/10.1016/0005-1098(83)90046-8
- Fan, Y., et al. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology. doi.org/10.1111/bjet.13544
- Kosmyna, N., et al. (2025). Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task. MIT Media Lab preprint. arxiv.org/abs/2506.08872
- Barcaui, A. (2025). ChatGPT as a cognitive crutch: Evidence from a randomized controlled trial on knowledge retention. Social Sciences & Humanities Open. doi.org/10.1016/j.ssaho.2025.102287
- Urban, M., et al. (2024). ChatGPT improves creative problem-solving performance in university students: An experimental study. Computers & Education. doi.org/10.1016/j.compedu.2024.105031
- Johansson, P., Hall, L., Sikström, S., & Olsson, A. (2005). Failure to detect mismatches between intention and outcome in a simple decision task. Science. doi.org/10.1126/science.1111709
- Lee, H.-P., et al. (2025). The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. CHI Conference on Human Factors in Computing Systems. doi.org/10.1145/3706598.3713778
- Bastani, H., et al. (2025). Generative AI can harm learning. PNAS. doi.org/10.1073/pnas.2422633122