The numbers are stark. Top scores on an AI-supervised remote examination jumped fivefold compared to historical norms, a statistical anomaly so severe that the examining body had no choice but to invalidate the results entirely. Fifty-eight thousand students now face a resit. The cause, according to reporting by Ars Technica, was a failure of the AI proctoring system to adequately detect or deter cheating at scale. Remote invigilation technology, deployed without sufficient human checks, turned a high-stakes assessment into an open-book free-for-all.
This is not an argument against AI in education. It is an argument against magical thinking about what AI can currently do on its own. AI proctoring tools, such as those offered by Proctorio, Examity, and similar platforms, work by flagging suspicious behaviour for human review. The keyword is flagging. The human review part is not optional. According to research published by the AI Now Institute, automated proctoring systems have documented failure rates across a range of demographic and environmental variables, and perform worst when deployed without a trained human in the loop to interpret their outputs.
Scotland's education bodies have been thoughtful here, at least so far. The Scottish Qualifications Authority has moved cautiously on AI-assisted assessment, and the Scottish Government's AI in Education guidance, updated in 2024, explicitly calls for human oversight to remain central to any automated process involving learner outcomes. That caution looks prescient this week. The lesson from this international case is not that the technology failed in isolation. It is that the institution deploying it assumed the technology was the entire system, when it was only ever meant to be part of one.
For schools, colleges, and training providers across Edinburgh and the rest of Scotland, the practical implication is this: AI tools in assessment contexts are genuinely useful for reducing admin load, flagging patterns, and scaling invigilation capacity. The University of Edinburgh's Bayes Centre and similar research environments have explored AI-assisted marking and feedback with real, positive results. But those results come from programmes where human educators remain in control of final decisions. The moment an institution treats AI output as the final word rather than a useful signal, it has crossed a line it may not be able to walk back from without significant reputational damage.
There is a broader point here for any organisation using AI to automate a high-stakes process, whether that is an exam, a hiring decision, a clinical triage, or a financial assessment. The tool is only as good as the governance wrapped around it. Speed and scale are AI's genuine gifts. Judgement, context, and accountability still belong to people. The 58,000 students sitting a resit they should not have needed are paying the price for a system that forgot the second half of that equation.
