The talk · PMSA Leadership Conference · 17 July 2026

Learning, Judgement and Decision‑Making in the Age of AI

The keynote in full — all sixty‑six slides, with the words that went with them.


This page is a record of the talk, to revisit or share. The pre‑reading note that preceded it is at tangenlab.com/pmsa/prereading.

Title slide: PMSA Leadership Conference 2026 — Learning, Judgement and Decision-Making in the Age of AI — what to protect, what to build, what to measure. Jason Tangen, The University of Queensland. A white chess king stands on a wooden board beside laboratory glassware.
Slide 1 · Title

Part oneThe man with the stopwatch

Slide: Amsterdam, 1946 — Adriaan de Groot wanted to know what grandmasters actually have. Everyone assumed calculation and photographic memory. So did he. An antique wooden chess clock sits beside a chess set.
Slide 2 · Amsterdam, 1946

Amsterdam, 1946. A Dutch psychologist named Adriaan de Groot wants to understand what makes expert chess players so good. He’s also one of the strongest players in the Netherlands, so he knows the game well enough to study the question properly.

His question is straightforward: what are grandmasters doing differently?

The usual explanation was that grandmasters calculate further ahead and have unusually good memories. De Groot initially assumed the same thing.

So he ran the obvious test. He showed players a position from a real game. Twenty-odd pieces, mid-battle. Five seconds. Then he took it away and asked them to rebuild it from memory.

The grandmasters rebuilt it almost perfectly. Club players managed four or five pieces. That result seemed to support the memory explanation.

Slide: The experiment — A real position. Five seconds. Rebuild it from memory. Grandmasters were nearly perfect. Club players managed four or five pieces. A mid-game chessboard with a brass stopwatch resting on its corner.
Slide 3 · The experiment

The important follow-up was run in 1973 by Bill Chase and Herbert Simon.

They kept everything the same. Same board, same pieces, same five seconds. One change. They took the pieces and scattered them at random. Positions that could never occur in a real game.

The grandmasters dropped to the level of the beginners.

Slide: The follow-up — Same pieces, placed on random squares. The advantage disappeared. The grandmasters performed like the beginners. Two chessboards side by side, labelled game position and random position. Chase and Simon, Cognitive Psychology, 1973.
Slide 4 · The follow-up

The result wasn’t evidence of photographic memory. Experienced players had learned a very large number of familiar chess patterns—estimates are around fifty thousand. In a real position, a grandmaster doesn’t process every piece separately. They recognise several familiar configurations. In a random position, those configurations aren’t there, so the advantage disappears.

Slide: What they actually had — They never had better memories. Fifty thousand familiar patterns, built over decades, and no idea they were using them. A close-up of a carved wooden knight.
Slide 5 · What they actually had

So the difference was specialised knowledge, not better memory in general. It had been built over decades, and much of it operated automatically. Grandmasters could describe parts of their reasoning, but they couldn’t give a complete account of how they recognised a position.

You did this at breakfast

Slide: You did this at breakfast — Ask a fluent adult how they read. They can't tell you. It took years to build, and then it disappeared from view. A folded newspaper and a cup of coffee on a breakfast table.
Slide 6 · You did this at breakfast

Reading gives us a familiar example of the same process.

When you read the news this morning, you didn’t consciously identify each letter and assemble each word. Most of that processing happened automatically. A fluent reader can explain some reading strategies, but not every operation involved in recognising a sentence.

That fluency still took years to develop: phonics, spelling, comprehension, and a great deal of practice reading aloud. Over time, children stop having to concentrate on each component and can focus on meaning instead.

Once the process becomes fluent, it is easy to forget how much practice it required.

Cognitive scientists call these familiar units chunks. Research on chess, radiology, music, firefighting and fingerprint examination shows a similar pattern: repeated practice allows people to process several related elements as a single familiar unit. Experts then use those units quickly, often without being able to describe every step.

Diagram: Chunking — practice compresses the work into chunks; once a chunk forms, it disappears from its owner's view. An ascending staircase labelled letters, syllables, words, phrases, meaning, over an axis marked years of practice. After Chase and Simon, 1973; Castles, Rastle and Nation, 2018.
Slide 7 · Chunking
Slide: Chunking in practice — Seven digits became 79. He turned random numbers into running times, ages and dates. Paper strips of digits and a stopwatch on a desk. Ericsson, Chase and Faloon, Science, 1980.
Slide 8 · Chunking in practice

Why exams worked

This creates a practical problem for education.

If we cannot observe learning directly, how do we assess and certify it?

We usually infer learning from performance: the rebuilt chessboard, the essay, the exam response, or the answer to a follow-up question. We cannot observe the underlying knowledge directly, but we can examine what a person produces with it.

Slide: The assessment problem — We cannot see learning directly. We assess what a person can produce with it. A chessboard, an open notebook, a typed manuscript and a voice recorder on a table.
Slide 9 · The assessment problem

Exams and qualifications rely on that relationship. Board papers often do as well: the document is treated as evidence that the author has understood the issue and worked through the relevant evidence.

Until recently, producing a good essay generally required the writer to understand the material, organise an argument and make decisions about evidence.

Slide: The exam hall — Until two years ago, there was only one way to get a good essay. You had to do the thinking. So the essay was proof of the thinking. An empty exam hall with rows of wooden desks.
Slide 10 · The exam hall

Generative AI weakens that connection between the quality of the document and the understanding of the person submitting it.

A student can now produce a plausible essay very quickly without doing all of that work. On the page, it may be difficult to tell how much the student understood or contributed.

Slide: Since ChatGPT — Now there are two ways to produce that essay. And on the page, they look identical. Two identical-looking stacks of typed pages on a lab table.
Slide 11 · Since ChatGPT

I want to focus on what that means for learning, rather than only for cheating or detection. The work involved in writing essays and solving problems helps students develop three kinds of judgement: distinguishing strong work from weak work, deciding when work is good enough, and judging their own understanding.

The map

I’ll take those in order. First, how people learn to distinguish one thing from another. Second, how they decide what counts as good enough. Third, how accurately they judge their own knowledge. We’ll pause for questions after each section, and then look at practical implications for schools and boards.

Slide: Where we're going this morning — Telling things apart: why experts can't always explain their own judgement. Good enough: who decides when a piece of work is finished. Knowing that you know: why confidence and understanding can come apart.
Slide 12 · This morning

I’ll start with fingerprint examination, which is the area I’ve studied for the past twenty years, and with the difference between the television version and the real process.

Part twoThe examiner’s eye

Section divider: I — The examiner's eye — telling things apart. A magnifying glass held over a card of enlarged fingerprint ridge patterns.
Slide 13 · Part I · The examiner’s eye

The CSI myth

On television, a technician scans a fingerprint, the database produces a name, and the result appears as a confirmed match. It looks like a largely automated process.

Slide: On television — A fingerprint match is a beep and a green light. In real life, it is a human being making a judgement call. A brass loupe beside a card of inked fingerprints.
Slide 14 · On television

In practice, a crime-scene print is often partial, smudged or distorted. Skin does not leave exactly the same impression every time. An examiner compares two images and decides whether the similarities are sufficient to support an identification. They can point to relevant features and explain much of the comparison, but some of their pattern recognition comes from experience and is difficult to put into words.

This is another example of learned pattern recognition. Our research shows that qualified fingerprint examiners are very good at distinguishing matching prints from similar non-matches. It also shows that their decisions can be affected by the same limitations that affect other forms of human judgement. The Mayfield case is a well-known example.

Madrid

On the eleventh of March, 2004, bombs go off on four commuter trains in Madrid. A hundred and ninety-three people are killed. Spanish police find a partial fingerprint on a bag of detonators. This is now the largest terrorism investigation in Europe, so the print goes out to agencies around the world, including the FBI.

Slide: Madrid, 11 March 2004 — A partial print on a bag of detonators. The FBI matched it to a lawyer in Oregon named Brandon Mayfield. An evidence bag holding a card with a smudged partial fingerprint.
Slide 15 · Madrid, 11 March 2004

The FBI runs it through their database and lands on an American: a lawyer in Oregon named Brandon Mayfield.

A senior FBI examiner compares the prints and declares a match. A second examiner verifies it. A third concurs. Mayfield is arrested. And when his defence team asks for an independent expert, the court appoints one, a well-regarded examiner with no ties to the FBI. He examines the prints, and he confirms the match too.

Four examiners reached the same conclusion.

Spanish police continued to disagree with the identification. A few weeks later, they identified the actual source of the print as an Algerian national named Ouhnane Daoud. The FBI released Mayfield, apologised publicly and later paid him two million dollars.

Slide: Four experts agreed — Three FBI examiners confirmed it. So did the court's independent expert. The print belonged to someone else. The FBI apologised and paid two million dollars. A tall stack of case files topped with wire-rimmed glasses.
Slide 16 · Four experts agreed

It would be easy to explain the error as carelessness or misconduct, but the evidence points to a more ordinary problem. Context can influence expert judgement: the seriousness of the case, knowledge that another examiner has reached a conclusion, or the way a candidate print is presented. Because those influences are not always conscious, examiners may remain confident even when their judgement has been affected.

The important point is that confidence does not necessarily tell us whether a judgement has been influenced.

Slide: What went wrong — They weren't careless. They were seeing with chunks. Expert judgement fails silently, with complete confidence. A loupe and a tally counter reading zero on a metal tabletop.
Slide 17 · What went wrong

What the eye actually is

When fingerprint expertise works well, the examiner distinguishes relevant detail from distortion and coincidental similarity. That ability develops through years of comparing prints and receiving feedback. A lecture can provide principles and terminology, but it cannot replace the comparisons needed to develop the skill.

In cognitive science, this is called discrimination: learning to distinguish categories that initially look similar.

Diagram: The first ability — Telling things apart. Scientists call it discrimination. Built from thousands of pairs, over years, with feedback. Two overlapping bell curves labelled noise and signal, with the overlap marked as where telling apart is hard. After Green and Swets, 1966.
Slide 18 · The first ability

Students develop discrimination when they compare examples, make judgements and receive feedback. In writing, that means learning to recognise whether an argument is clear, whether evidence is relevant and whether one explanation is stronger than another. AI-generated text gives students a new type of material to evaluate, and they may not yet have enough experience to evaluate it reliably.

Slide: How judgement develops — Students learn by comparing, deciding and getting feedback. AI gives them more text to evaluate, but not necessarily the experience to evaluate it well. A desk with a stack of marked-up essays and a plaque listing clear argument, relevant evidence, stronger explanation.
Slide 19 · How judgement develops

A recent medical study illustrates the problem.

Fifty doctors

2024. A randomised trial, published in JAMA Network Open. Fifty experienced physicians, family medicine, internal medicine, emergency medicine, are given difficult diagnostic cases. Half work with their usual resources. Half get their usual resources plus GPT-4.

Diagram: Stanford, Boston, Virginia, late 2023 — Fifty doctors, six real cases, one hour. The cases had never been published anywhere, so the model had not seen them. Flowchart: 50 physicians split into usual resources and usual resources plus GPT-4, six diagnostic cases in one hour, reasoning scored by blinded experts. Goh et al., JAMA Network Open, 2024.
Slide 20 · The design

The researchers expected that access to GPT-4 might improve the doctors’ performance.

Doctors alone scored 74 per cent. Doctors with GPT-4: 76. Statistically, that’s a tie.

Then the researchers ran the model on its own, no doctor. Almost 90.

Chart: The scores, per case — The doctors couldn't tell when the machine was right. Physicians alone 74 per cent, physicians with GPT-4 76 per cent, GPT-4 alone 90 per cent. Goh et al., JAMA Network Open, 2024.
Slide 21 · The scores

The doctors did not consistently use the model’s better answers to improve their own performance.

The doctors had extensive diagnostic training, but that training focused on patients, clinical histories and examinations. It did not necessarily prepare them to decide when a fluent model-generated explanation should override their own view. When the model disagreed with them, they did not always identify which answer was better.

Slide: Why it matters here — Skill in one domain doesn't carry to another. Thirty years of judging patients said nothing about judging model output. A stethoscope coiled beside a closed laptop.
Slide 22 · Why it matters here

The relevance to a board is that experience in budgets, contracts, curriculum or staffing does not automatically confer expertise in evaluating AI-generated analysis. As AI-written or AI-edited material enters board papers, the fact that it reads well should not be treated as evidence that its reasoning is sound.

What this asks of you

There are three practical responses.

In the classroom, put two essays on the same question side by side and ask students which is stronger and why. Have them explain their decision aloud, then give them feedback on the judgement. For boards and committees, periodically read a complete underlying document—a budget or safeguarding report, for example—rather than relying only on summaries. In both cases, the relevant skill is developed by working directly with the material.

Slide: What to do about it — Three tasks that train the eye. For students: put two essays side by side and ask which is stronger, and why. For teachers: have students defend their verdict aloud, then give them feedback. For the board: once a term, read one safeguarding report or budget in full.
Slide 23 · What to do about it

That’s the first ability: discrimination. Before I move on, what questions or objections do you have?

Slide: Over to you — Questions welcome. Two espresso cups beside an antique brass microscope.
Slide 24 · Over to you

Part threeThe control room

Section divider: II — The control room — good enough. A close-up of a vintage circular gauge with a brass needle.
Slide 25 · Part II · The control room

England, the 1970s

The second ability is easier to explain through research on automation.

In the 1970s, Lisanne Bainbridge studied control-room operators in settings such as chemical plants and power stations. Automated systems handled routine adjustments, while operators monitored the system and intervened when something unusual happened. The arrangement appeared sensible: automate routine work and leave exceptions to people.

Slide: England, the 1970s — Lisanne Bainbridge studied the operators automation left behind. The machines did the routine work. The people supervised. A vintage control room with a wall of analogue dials and an empty chair.
Slide 26 · England, the 1970s

In her 1983 paper Ironies of Automation, Bainbridge pointed out a problem with that arrangement. When operators no longer perform routine tasks, they get less practice. Yet the situations in which they must intervene are usually the most difficult ones. Automation can therefore reduce the experience people need for the work that remains.

Her summary is still useful, so I’ll read it directly:

“By taking away the easy parts of his task, automation can make the difficult parts of the human operator’s task more difficult.”

Slide: A brass pressure gauge under lamplight, with the quote — By taking away the easy parts of his task, automation can make the difficult parts of the human operator's task more difficult. Lisanne Bainbridge, Ironies of Automation, 1983.
Slide 27 · Bainbridge, 1983

The same problem can arise when AI removes routine practice from education.

Her prediction is running an hour from here

Fingerprint examination provides a current example.

Queensland’s fingerprint bureau uses an automated process often called lights-out latents. An algorithm handles easier comparisons without an examiner, while more difficult cases are sent to people.

That changes the mix of cases available to trainees. Straightforward comparisons used to provide early practice. If those cases are automated, new examiners begin with a greater proportion of difficult material. As the system improves, the remaining human caseload may become more difficult again.

Slide: Meanwhile, in Queensland — The easy prints go to the machine now. Those prints were how new examiners learned the craft. As the AI improves, the humans get only the harder leftovers. A stack of fingerprint cards feeding into a sorting machine.
Slide 28 · Meanwhile, in Queensland

The efficiency gain may therefore create a training problem unless the organisation provides practice in another way.

The threshold for automation is usually chosen to meet operational goals such as speed and accuracy. Training effects also need to be part of that decision.

The easy part and the hard part

The same issue applies to writing. Producing a first draft is only part of the task. Students also need to decide whether the argument works, whether the evidence is sufficient and whether the piece is ready to submit. They develop that judgement by drafting, revising and receiving feedback many times.

If AI produces the draft, students may be asked to evaluate a polished text without having had enough experience creating and revising one themselves.

Slide: Forty years early — The easy part of an essay is producing the draft. The hard part is knowing whether it's good enough, and that part just got harder. A handwritten page and fountain pen on an antique desk.
Slide 29 · Forty years early

The second half of the examiner’s judgement

Fingerprint decisions also involve a second skill.

After examiners identify relevant similarities, they must decide whether the evidence is sufficient to report a match. Two competent examiners can see the same details and still reach different conclusions because they require different amounts of evidence.

The evidence does not determine that threshold on its own. A lower threshold increases the risk of a false identification; a higher threshold increases the risk of missing a genuine match. The choice therefore depends partly on how the organisation weighs those two errors.

Slide: Back in the lab — Same evidence. Different thresholds. Different decisions. A lower threshold risks a false identification. A higher one risks missing a genuine match. Two examiners at magnifying lamps, overlaid with two evidence scales marked lower criterion, report match, and higher criterion, not enough to report.
Slide 30 · Back in the lab

In cognitive science, this threshold is called a criterion. Students make similar decisions when they ask whether they have enough evidence, whether a response is complete, or whether a piece of work is ready to submit. They improve those decisions through repeated practice and feedback.

Diagram: The second ability — Where to set the bar is a choice. Cognitive scientists call the bar a criterion. Two overlapping bell curves bisected by a vertical line labelled the bar, with arrows marked lenient, more wrong matches, and strict, more cases let go. After Green and Swets, 1966.
Slide 31 · The second ability

AI-generated work often arrives polished and well structured. If students mainly receive finished-looking text, they get less practice deciding what needs revision and when a piece of work is actually complete.

Slide: The same decision — Students also have to decide when work is ready. Practice and feedback teach them what needs revision. Finished-looking AI text can reduce that practice. A gauge marked draft 1, draft 2, draft 3, looks ready, ready to submit, beside a stack of annotated drafts.
Slide 32 · The same decision

At your table

Boards face a related issue. Capital works proposals, fee recommendations and vendor submissions usually arrive in a finished format. The quality of the presentation can influence whether the evidence feels sufficient, even though presentation quality and evidentiary quality are different things. AI makes polished presentation easier to produce.

One useful response is to specify the decision criteria before reading a consequential proposal. What evidence is required? What would rule the proposal out? Writing those conditions down makes it easier to assess the submission against an independent standard.

Slide: The next proposal in your board pack — Write down what a yes would require, before you read it. Then the proposal has to clear your bar instead of setting it. A blank stack of paper and a fountain pen on a dark table.
Slide 33 · The next proposal in your board pack

In the classroom, ask students to describe the criteria before they assess a draft, whether it is their own or AI-generated. Assess the quality of those criteria as well as the final work. For significant board decisions, record the criteria in advance and periodically review whether they were appropriate for the risks involved.

Slide: What to do about it — Making the bar visible. In the classroom: students state the bar before they see any draft. In assessment: mark the bar-setting, not just the work. At the board: once a year, ask whether past bars matched the stakes.
Slide 34 · What to do about it

That’s the second ability: setting a criterion. I’ll pause again for questions.

Slide: Over to you — Questions welcome. A glass carafe and two glasses of water beside an open notebook.
Slide 35 · Over to you

Part fourThe best essays in the room

Section divider: III — Can students judge their own understanding? — metacognition. An index card with three prompts: what do I understand, what am I unsure about, what should I check.
Slide 36 · Part III · Metacognition

The experiment that should have been good news

The third section begins with a study in which AI improved the students’ essays.

In 2025, the British Journal of Educational Technology published a study of 117 university students. They worked in a lab, with every click and keystroke recorded. Everyone wrote the same two-hour essay from the same materials, then had an hour to revise it. That’s where the experiment happened. Some revised alone, some with ChatGPT, some with a human writing tutor, and some with an automated feedback checklist. One design detail matters: ChatGPT was allowed to advise, but not to write the essay for them.

Diagram: Every click and keystroke recorded — One task, four revision conditions. Flowchart: 117 students, same two-hour task, split into no extra support, ChatGPT advice, human expert, and checklist tools, then one-hour revision with activity logged, then original and revised essays scored. Fan et al., British Journal of Educational Technology, 2025.
Slide 37 · The design

The ChatGPT group produced the strongest essays. On that outcome alone, the tool looked beneficial.

Slide: The result — The ChatGPT group wrote the best essays. Stop reading there, and you'd mandate the tool by Friday. Four stacks of paper of varying heights on a marble counter.
Slide 38 · The result

Because Fan and colleagues recorded the revision process, they could also examine how the students worked. The ChatGPT group interacted repeatedly with the chatbot, but did much less planning, checking and evaluation of their own work. The essays improved while students engaged in less self-monitoring. This happened even though the chatbot was configured to advise rather than write the essay.

Diagram: What the recordings showed — ChatGPT became the centre of the revision process. In the ChatGPT group, revision revolved around the chatbot; in the human-expert group, revision stayed connected to reading, checking the task and evaluating the work. Simplified from the authors' process maps. Fan et al., 2025.
Slide 39 · What the recordings showed

Fan and colleagues described this pattern as metacognitive laziness: relying on the tool to guide the process rather than monitoring the work independently.

A small MIT preprint reported a related result. It involved 54 students, so I would treat it as preliminary. Students wrote essays with or without AI while wearing EEG caps. The AI-assisted group showed lower measures of brain engagement, and 83 per cent could not quote a sentence from their essay shortly after finishing. The result suggests that producing a document with AI does not necessarily involve close engagement with its content.

Slide: Minutes after finishing — 83 per cent couldn't quote one sentence from their own essay. A preprint with 54 students; an illustration, not proof. An EEG electrode cap on a stand. Kosmyna et al., MIT Media Lab, 2025.
Slide 40 · Minutes after finishing

The surprise test

The next question is whether these process differences affect later learning.

Slide: From performance to learning — A better essay is one outcome. Learning is another. The next question is what students can do later, without the same support. A student writing by hand at a desk while a laptop shows a chat interface.
Slide 41 · From performance to learning

In a 2025 study in Rio de Janeiro, 120 business students were randomly assigned to research a topic either with ChatGPT or with books, articles and ordinary search. They had two weeks to prepare a short presentation for their peers. Forty-five days later, they returned for a test they had not been told about in advance.

Diagram: Rio de Janeiro, 2024 — Two weeks to learn a topic and teach it to your peers. Half could use ChatGPT, half used books, articles and ordinary search. Nobody knew a test was coming. Flowchart: 120 students, ChatGPT allowed or books and search only, research a topic and present it to peers, surprise test 45 days later. Barcaui, Social Sciences and Humanities Open, 2025.
Slide 42 · Rio de Janeiro

The group that studied unaided: 68.5 per cent. The ChatGPT group: 57.5.

Slide: The scores — Both groups had felt equally confident. Studied unaided: 68.5 per cent. Studied with ChatGPT: 57.5 per cent. Barcaui, 2025; randomised, 120 participants.
Slide 43 · The scores

Both groups had reported similar confidence during the original task. Their confidence therefore did not reflect the later difference in retention. A separate study examined whether students could judge the quality of their work at the time they produced it.

Urban and colleagues tested 145 university students in Prague in 2024. Students acted as consultants to a toy company and proposed three ways to make a stuffed bunny outsell Lego. Half used ChatGPT and half did not. Two experts, who did not know which condition each student was in, scored the solutions. Students also rated their own work.

Diagram: Prague, 2024 — Three ideas to make a stuffed bunny outsell Lego. Half worked with ChatGPT, half without. Blind experts scored every solution, then every student rated their own work. Flowchart: 145 students, with ChatGPT or without, three ideas to improve a toy, experts score it and students rate their own. Urban et al., Computers and Education, 2024.
Slide 44 · Prague, 2024

The ChatGPT group produced stronger and more original solutions. They also reported that the task felt easier, required less effort and left them more confident.

However, within the ChatGPT group, students’ self-ratings were essentially unrelated to the expert scores. The correlation was close to zero. Students who found ChatGPT more useful also tended to overestimate their performance by more.

In this study, feeling confident while using the tool was not a reliable indicator of solution quality.

Slide: The third ability — ChatGPT improved the work. Students still struggled to judge it. Their own ratings did not line up with the experts' ratings. Two circles labelled student's rating and expert's rating, joined by a broken line marked did not line up. Urban et al., Computers and Education, 2024.
Slide 45 · The third ability

This third ability is metacognition: accurately judging what you understand, what you are uncertain about and when you need to check. If AI improves the submitted work without improving the student’s understanding, grades and feedback may become harder for the student to interpret.

The invented reasons

An older experiment shows why people’s explanations of their own decisions need to be treated carefully.

Diagram: From judgement to explanation — Can people accurately explain how they reached a decision? That matters when we ask students to describe how they worked with AI. Two circles labelled the decision and the explanation, joined by a dotted line and a question mark.
Slide 46 · From judgement to explanation

In a 2005 study in Lund, Petter Johansson and colleagues showed participants pairs of photographs and asked which face they found more attractive. The researcher then handed them the selected photograph and asked them to explain the choice.

On some trials, a concealed card-switch meant that the researcher actually handed over the photograph the participant had rejected.

Diagram: Lund, Sweden, 2005 — Pick the more attractive face, then say why. 120 people, fifteen pairs of faces, three of them secretly swapped by sleight of hand. Flowchart: which face is more attractive, point to one, sleight of hand swaps the cards, now explain your choice. Johansson et al., Science, 2005.
Slide 47 · Lund, Sweden, 2005

Most participants did not notice the swap. Only about one in eight detected it immediately, and roughly three quarters of the swaps went undetected even under a broader measure. Participants then gave detailed reasons for choosing the photograph they had actually rejected, sometimes referring to features that were not present in their original choice.

Slide: What happened next — People gave confident reasons for choices they never made. About three swaps in four went completely unnoticed. People praised details of the photograph they had actually rejected. Johansson et al., Science, 2005.
Slide 48 · What happened next

The participants were not deliberately misleading the researcher. They generated plausible explanations without recognising that the photograph had changed. This shows that a confident explanation does not guarantee accurate access to the process that produced a decision.

Slide: The part to remember — Giving reasons is a skill that never switches off. The reasons feel real from the inside, even when they can't be right. A vintage reel-to-reel tape recorder by a window.
Slide 49 · The part to remember

That matters when we ask people to describe how critically they worked with AI.

The room’s version

The next study involved working adults rather than students.

Diagram: Microsoft and Carnegie Mellon, 2025 — This study is about people like us. 319 knowledge workers each described three real AI tasks from their own jobs: creating, finding out, or getting advice. For each task: did you think critically, and where did the effort go? Lee et al., CHI, 2025; 936 tasks.
Slide 50 · Microsoft and Carnegie Mellon

In 2025, researchers from Microsoft and Carnegie Mellon asked 319 knowledge workers to describe three AI-assisted tasks from their jobs: creating something, finding information or getting advice. Across 936 tasks, participants reported where they had used critical thinking and where the effort had shifted. They spent less effort producing material and more effort checking AI output. Greater trust in the tool was associated with less reported critical thinking, while confidence in their own ability was associated with more.

Slide: The correlation — More trust in the tool was associated with less reported critical thinking. Work shifted from producing to checking the machine's output. That matters because checking is much of what a board does. Lee et al., CHI, 2025.
Slide 51 · The correlation

This is relevant to boards because much of a board’s work involves reviewing material produced by other people. Effective review depends on current domain knowledge and on understanding how the underlying work is done. If AI changes that underlying process, boards need more than a general instruction to check the output.

What this asks of you

In the classroom, ask students to identify which parts of a substantial piece they are confident about and where they are uncertain. Compare predicted marks with actual marks over time. Ask students to explain their own document without referring to it. These activities provide direct evidence about how accurately they are monitoring their understanding.

Slide: What to do about it — Reading your own gauge. In the classroom: students say what they're sure of and where they're guessing. In assessment: predicted mark beside actual mark, until they converge. At the board: make it normal to say you haven't understood a paper.
Slide 52 · What to do about it

At board level, make it acceptable to say, “I haven’t understood this.” For major decisions, ask members to state their level of confidence and what that confidence is based on. This helps distinguish confidence based on relevant evidence from confidence based on a clear and polished presentation.

Slide: Board-level practice — Make it acceptable to say, I haven't understood this. For major decisions, ask how confident people are, and what that confidence is based on. Relevant evidence, or a clear and polished presentation? A board member raising a hand toward a dashboard on a screen.
Slide 53 · Board-level practice

That’s the third ability: metacognition. I’ll pause for questions, then bring the three abilities together and discuss possible responses.

Slide: Over to you — Questions welcome. A white teapot with steam rising and two teacups on a wooden table.
Slide 54 · Over to you

Part fiveThe lift

Their word

The three sections describe related forms of judgement.

Discrimination is the ability to distinguish relevant differences. Criterion is the threshold used to decide what counts as sufficient. Metacognition is the ability to judge your own understanding. All three depend on practice and feedback, and all three can be affected when AI changes who performs the underlying work.

Slide: The three abilities — Telling apart. Setting the bar. Knowing that you know. Discrimination, criterion, and metacognition. A magnifying glass, a balance scale and a gauge on a shelf.
Slide 55 · The three abilities

PMSA often brings these forms of judgement together under the word discernment.

Slide: Your word for it — Discernment. Your schools have used this word for a hundred and fifty years. Cognitive science shows what it is made of.
Slide 56 · Your word for it

In this context, discernment is not a single general capacity. It includes learned abilities to notice differences, apply appropriate standards and monitor one’s own knowledge. That gives schools something more specific to protect and assess.

The Turkish experiment

A study in a Turkish high school shows how the design of an AI tool can change its effect on learning.

The study took place in a large high school in Turkey during the 2023 school year. Bastani and colleagues at the University of Pennsylvania later published it in PNAS. Nearly a thousand students, in Years 9 to 11, took part in four ordinary ninety-minute maths lessons. The teacher reviewed a topic, the class practised problems, and then everyone sat a short closed-book exam alone. Both practice and exam contributed to real grades.

Whole classrooms were randomised into three groups. One group used textbooks and notes. Another used a GPT-4 chatbot—essentially ChatGPT. The third used a version the researchers had built with the school’s maths teachers: the same model, but redesigned to behave like a tutor. It would not simply give the answer. It offered hints, asked questions and moved one step at a time, using the teachers’ worked solutions and common mistakes behind the scenes.

Diagram: A high school in Turkey, 2023 — Four ordinary maths lessons, three kinds of help. Whole classrooms randomised; practice with the assigned help, then a closed-book exam alone; both counted toward real grades. Flowchart: nearly 1,000 students by classroom, split into textbook and notes, GPT-4 chatbot, and GPT-4 tutor with guardrails. Bastani et al., PNAS, 2025.
Slide 57 · A high school in Turkey

During practice, the chatbot classrooms scored 48 per cent higher than the textbook classrooms. The tutor group scored 127 per cent higher. Both AI conditions improved performance while help was available.

Chart: The practice sessions — During practice, both AI groups pulled far ahead. Textbook and notes: baseline. GPT-4 chatbot: plus 48 per cent. GPT-4 tutor: plus 127 per cent. Bastani et al., PNAS, 2025.
Slide 58 · The practice sessions

On the closed-book exam, the standard-chatbot classrooms scored 17 per cent below the textbook classrooms. The tutor group performed at about the same level as the textbook group. The standard chatbot improved assisted practice performance but was followed by lower independent performance; the tutor improved practice performance without the same penalty.

Slide: Then the exam, alone — The chatbot classrooms landed below students who never had AI. Minus 17 per cent for the GPT-4 chatbot; zero per cent, no penalty, for the GPT-4 tutor, relative to the textbook classrooms. Bastani et al., PNAS, 2025.
Slide 59 · Then the exam, alone

The interaction records help explain the difference. Students often asked the standard chatbot for answers and copied what it produced. When asked directly for answers, the chatbot made logical errors on 42 per cent of the problems. The tutor was designed to provide hints and questions rather than complete answers, so students had to continue working through the problem.

Slide: Why the difference — Students asked the chatbot for answers, and copied them. Asked straight for an answer, it was wrong about half the time. The tutor only gave hints, so there was nothing to copy. An open notebook of handwritten maths with a calculator resting on it. Bastani et al., PNAS, 2025.
Slide 60 · Why the difference

Students’ perceptions were not reliable indicators of these effects. The chatbot group did not report learning less, and the tutor group overestimated its exam performance.

The underlying model was the same. What changed was how students were allowed to use it. That design choice is something a school can govern.

Slide: Same model, opposite results — Same model. Different design. Opposite outcomes. How students were allowed to use it changed what they learned. A bronze compass resting on a rolled blueprint.
Slide 61 · Same model, opposite results

What good use looks like

A useful design principle is to use AI to prompt further thinking rather than replace the task students need to practise.

For example, a student could finish an essay and then ask the model to identify a weak argument. The student would still need to decide whether the criticism is valid and revise or defend the passage. After reading a difficult chapter, a student could use the model for retrieval questions and note where they cannot answer without help. In both cases, the student continues to perform the writing, recall and judgement being learned.

Boards can use the same principle. A model might identify claims in a vendor proposal that need closer examination, suggest information missing from a staff paper, or generate possible failure scenarios for a proposed decision. Board members and staff would then verify those suggestions against the evidence. The tool is being used to expand the review, not to make the decision.

Slide: The rule that worked in Turkey — Let it question the work. Don't let it do the work. For an essay: ask it for the weakest argument, then defend or rewrite. For a hard text: have it quiz you until you can't answer. For this board: ask it for the three weakest claims in the next vendor pitch.
Slide 62 · The rule that worked in Turkey

There are two governance implications. First, students, staff and board members need examples and training in appropriate use; good practice is not self-evident. Second, the capabilities and limitations of these tools change frequently. The institution needs clear responsibility for reviewing its guidance and evidence on a regular schedule. That responsibility extends beyond technical maintenance.

Slide: Two jobs for governance — Using AI well is a skill, and a moving target. It has to be taught, and someone has to watch it change. An antique barometer mounted on a white wall.
Slide 63 · Two jobs for governance

Two children

I’ll finish with the staircase image used in the slides.

Two children reach the top floor. One used the stairs and one used the lift. At the top, both can see the same view.

If we look only at the final position, we cannot tell which route they took.

But only one has practised climbing. If the next task requires that capacity, the route matters.

Slide: Two children — From the top floor, you cannot tell them apart. Only one of them practised the climb. A stairwell with concrete steps and a brass handrail beside a lift door.
Slide 64 · Two children

An essay, diagnosis or board paper is an output. We also need to ask what knowledge and judgement were developed while producing it. In the terms I’ve used today, that includes discrimination, criterion and metacognition. PMSA describes these together as discernment. AI can bypass some of the practice that develops them, or it can be designed to support that practice.

The outcome depends on the choices schools make about when and how the technology is used.

Slide: The climbing was the point. A man in a long coat climbing a concrete staircase toward the light.
Slide 65 · The climbing was the point

That is what I mean by saying that the climbing was the point. Thank you.

Slide: Thank you — What to protect, what to build, what to measure. Jason Tangen, tangenlab.com/pmsa. A study table of notebooks and instruments, with a figure standing at the window.
Slide 66 · Thank you

Sources