HomeCommunity ResourcesBeyond the Illusion of Competence: What a 93.5% Exam Failure Taught Me About Teaching With AI October 1, 2026 Community Resources, News When the 46 students in my translation course were allowed unrestricted access to AI tools in their final examination, 93.5% of them failed, and every failing paper contained the same machine-generated error. After twenty-seven years of teaching, that result forced me to rebuild my pedagogy from the ground up. This is an honest account of what I changed, why, and what I will report at OEB 2026. Last year I did something that many colleagues considered reckless. In the final examination of my third-year literary translation course at the University of Belgrade, students were granted unrestricted access to AI tools: ChatGPT, DeepL, Google Translate, Gemini, anything they wished. The task was an excerpt from Antonio Muñoz Molina’s novel No te veré morir, calibrated to their level of Spanish. The one condition was a rule they knew from the entire course, because it is the rule of the profession: zero tolerance for material errors. In professional translation, a single critical error is grounds for rejecting a delivered text, no matter how elegant the rest of it sounds. The result: 93.5% of the cohort, 43 of 46 students, failed. And the three who passed had merely avoided the worst reading; none of them rendered the critical passage with full accuracy. The qualitative analysis was more disturbing than the number. Every failing student had submitted the identical material error, at the identical spot: the narrator’s rueful admission that he had spent the evening hablando yo solo, doing all the talking, was rendered as a confession of talking to himself, alone. ChatGPT, DeepL, Google Translate and Gemini each produced that same distortion, and the students carried it into their papers with complete confidence and without hesitation. The illusion of competence Learning scientists have long described the illusion of competence: the feeling of mastery that fluent rereading or recognition produces without real understanding underneath. What I witnessed was its AI-era form: a state in which the fluency of machine output is mistaken for understanding, and AI is treated as an infallible authority rather than what it actually is, a draft that requires human judgment. Here is what makes the case so instructive: the construction itself is not obscure. Wake any of these students at midnight and ask what hablando yo solo means, and they would tell you. The surrounding passage, in which the narrator addresses his interlocutor throughout, makes the meaning unmistakable; no dictionary and no cultural encyclopaedia are needed, only an attentive reading. The failure was not linguistic. My students are capable young people who did what the interface invited them to do: they read a fluent, grammatical, stylistically pleasant answer and concluded that it must also be correct. The machine’s confidence became their confidence. Verification was one click away, and nobody clicked. This is why I resist framing the situation as a cheating problem or a plagiarism problem. Those framings assume students know the right answer and take a shortcut around it. What I observed was different and, I would argue, more serious: students who could no longer see the error at all, because the tool had already answered on their behalf. The most revealing moment came after the exam. Confronted with the error, students turned to ChatGPT to check whether the literal translation could still be defended, and received the confirmation they were looking for. Faced with a choice between their teacher’s correction and the tool that had produced the mistake, they asked the tool. AI was no longer a drafting aid; it had become an authority entitled to overrule expert human judgment. The reset First, the students. In our examination system, as in much of continental Europe, a failed exam is simply retaken at a later session, and no one’s studies end over it. We went through the error together, my students and I, and when they sat the examination again, they passed. That shared work is where the rebuilding began. A failure of that magnitude is not a grading incident. It is empirical evidence that the course itself, however well it had worked for over two decades, no longer taught what students most need. So I rebuilt it, and the question guiding the redesign was simple: if intelligent systems generate plausible answers faster than students can evaluate them, what exactly must I teach, and what exactly should an exam measure? Four changes define the redesigned course. Post-editing moved to the centre of the syllabus. For years, working with machine output sat at the margins of translator training, treated as a lesser skill. That hierarchy is now untenable. Critically interrogating a fluent machine draft is not a supplement to translation competence; for this generation, it is where translation competence is most visible. Comparative critique became a core classroom exercise. Students now post-edit and analyse the outputs of several large language models side by side, on the same source text, in class. When four tools agree and are all wrong, something clicks that no lecture can produce. The shared error stops being invisible and becomes an object of study. Cognitive fundamentals returned, deliberately. Pragmatic awareness, sustained attention, discourse-level reasoning: precisely the capacities that erode when thinking is outsourced. We train them explicitly, as a counterweight, not out of nostalgia but because they are the preconditions for evaluating machine output at all. Assessment became process-oriented. The exam no longer evaluates only the delivered product. It evaluates where the student intervened in the machine draft, what they verified, and what evidence their decisions rest on. A student who catches and corrects a subtle machine error demonstrates something an unedited perfect output never can. We reward the judgment, not the fluency. Why this is everyone’s problem My case study comes from literary translation, but nothing in the argument depends on it. Every discipline now faces students who can obtain a plausible answer to almost any task in seconds. Every discipline therefore faces the same pedagogical question: how do we teach, and how do we credibly assess, the human judgment that decides whether a plausible answer is a correct one? This is what draws me to this year’s OEB theme, Sovereign Learning: Trust, Agency and the Education Reset. Sovereignty in learning, as I understand it after this experience, is not independence from AI tools. My students will use them for the rest of their professional lives, and they should. Sovereignty is the reconstructed capacity to stand over the output, to doubt it productively, and to take responsibility for the final decision. It is human judgment rebuilt as something teachable, practisable, and testable. And it makes the teacher not less necessary but more so, because judgment is not transmitted by systems that always answer; it is cultivated by people who know when to withhold an answer and ask a better question. What I will report in Berlin The full case study, with the methodology, the comparative error analysis and the study’s limitations, is published in the EDULEARN26 conference proceedings (doi.org/10.21125/edulearn.2026.2013). But the story does not end there. The redesigned cohort sat their final examination in June 2026, again with full access to AI tools, again under zero tolerance for material errors. At my OEB 2026 session, Beyond the Illusion of Competence: Teaching and Assessing Critical Judgment in the AI-Enabled Classroom, I will present the outcomes of that exam transparently, including what did not work: the trade-offs I have not resolved, the exercises that looked good on paper and failed in the classroom, and the costs of process-oriented assessment that no one warns you about. I went into this reset without the comfort of certainty, twenty-seven years in and starting over. What I can promise is an honest account from inside the experiment, and a set of concrete practices you can take back to your own classroom, whatever your discipline. I hope you will join the conversation in Berlin. Written for OEB 2026 by Dr Jasmina Nikolić. Join Jasmina at OEB 2026 Leave a Reply Cancel ReplyYour email address will not be published.CommentName* Email* Website Save my name, email, and website in this browser for the next time I comment.