Oral Exams in Combatting AI Use in Universities

Updated: Aug 26

When Students Can Generate Answers, Universities Need to Ask Different Questions
For centuries, universities have operated on a relatively simple assumption: if a student submits an answer, the answer can be treated as evidence of what the student knows.
Artificial intelligence has weakened that assumption.
A student can now produce a coherent essay, summarize a difficult reading, construct an argument, or answer a conceptual question without necessarily possessing the knowledge that the finished product appears to demonstrate. The problem is therefore larger than cheating. Generative AI has created a verification problem for higher education.
Universities are no longer asking only whether an answer is correct. Increasingly, they must ask whether the student can actually account for the answer. This helps explain the unexpected return of one of academia’s oldest assessment methods: the oral exam.
AI Did Not Destroy the Exam. It Separated the Answer From the Student
Written assessment traditionally combines two things that universities care about: the quality of the answer and the competence of the person producing it.
Generative AI can separate them.
A student with limited understanding can potentially submit work displaying vocabulary, structure, argumentation, and conceptual sophistication far beyond what they could independently produce. The resulting document may still be excellent. What becomes uncertain is what exactly that excellence measures.
This distinction matters because universities do not award degrees to essays. They award degrees to people.
The institutional problem, therefore, is not simply that ChatGPT can generate text. It is that increasingly sophisticated generated text makes authorship a weaker proxy for competence.
This is particularly significant for take-home assessments. Catherine Hartmann of the University of Wyoming has argued for reconsidering oral examinations in undergraduate humanities education precisely in response to generative AI’s disruption of conventional written assessment. Her approach reflects a broader revival of interest in viva voce and interactive oral assessment as universities search for ways to establish whether students can explain what their submitted work appears to demonstrate. (1)
The important shift is subtle. The university moves from asking:
What answer can you produce?
to asking:
What can you do with your answer once I start questioning it?
That is a much more difficult task to outsource.
The Value of the Follow-Up Question
The strongest feature of an oral examination is not actually that students have to speak. It is that the examiner can respond. A written examination is largely static. The university asks a question; the student provides an answer; the answer is evaluated afterward. An oral examination creates a feedback loop. Why? What do you mean by that? How does this concept apply to a different case? What evidence would change your conclusion? You have contradicted your previous argument. Which position do you defend? Each additional question increases the amount of understanding the student must demonstrate.
This is why oral assessment can provide evidence that conventional written work increasingly struggles to provide. The examiner can observe whether students understand relationships between concepts, recognize weaknesses in their reasoning, transfer knowledge to unfamiliar situations, and defend conclusions when challenged. Contemporary guidance on oral assessment similarly emphasizes its capacity to probe understanding and critical thinking dynamically rather than merely evaluate a predetermined response. (2)
AI can produce an answer. Understanding becomes visible when the answer is disturbed.
From Product to Performance
This distinction points toward a broader transformation in assessment. Traditional written assessment evaluates a product of knowledge. Oral assessment can evaluate a performance of knowledge. The difference resembles the distinction between possessing a script and knowing how to perform the role.
Imagine a student submitting an excellent analysis of political framing. The essay correctly discusses how different descriptions of the same policy can activate different interpretations. A conventional assessment might stop there. An oral examiner can continue:
If framing is so powerful, why do some frames fail?
Can competing political actors successfully use the same frame? Would framing theory predict the same effect among voters with strong partisan identities? Now memorization —or AI-generated prose— is no longer sufficient. The student must manipulate the concept rather than merely reproduce its definition.
This makes oral examination interesting for reasons extending beyond academic integrity. It tests something universities frequently claim to value but do not always measure effectively: whether students can think with knowledge rather than simply present it. Research and institutional guidance increasingly describe oral assessment in similar terms, emphasizing its potential to support deeper learning, critical thinking, communication skills, and authentic demonstration of competence. (3)
Oral Exams Create Their Own Measurement Problem
There is, however, a danger in treating oral examinations as the obvious solution to generative AI. They solve one measurement problem partly by creating another.
A student who understands a subject exceptionally well may be nervous, introverted, slower in spontaneous speech, or simply better at constructing arguments through writing. Another student may possess weaker conceptual understanding but be charismatic, verbally fluent, and comfortable improvising.
An oral examination can therefore accidentally measure performance confidence alongside disciplinary knowledge. The distinction becomes especially important when students are being assessed in a second language, have disabilities affecting oral communication, or experience significant assessment anxiety.
Recent work on oral assessment consequently stresses that equity cannot be treated as an afterthought. Hu and Hashim argue that scalable oral assessment requires separating disciplinary reasoning from language production where appropriate, alongside standardization and safeguards designed to improve validity and reliability. UCL guidance similarly warns that poorly designed oral assessments can increase anxiety and introduce unintended barriers or bias. (4)
The paradox is clear. Universities may adopt oral examinations because they want to measure students more authentically, only to introduce new variables that make the measurement less reliable.
The Scalability Problem
There is also a much simpler obstacle: time. A professor can administer the same written examination to 200 students simultaneously. A twenty-minute individual oral examination for those same students requires roughly 67 hours of examination time before preparation, scheduling, breaks, grading, or administration are considered.
That changes the economics of assessment considerably. Oral examinations have long faced precisely these concerns. Research on viva voce assessment has identified examiner workload, inter-rater reliability, and student anxiety as persistent limitations. (5)
This makes the idea of replacing written examinations entirely with oral ones difficult to defend. But complete replacement may also be the wrong objective. One emerging alternative is the mini-viva or targeted oral verification. Students might complete an essay, project, or AI-permitted assignment and subsequently participate in a short oral discussion designed to establish whether they understand and can defend what they submitted.
Another proposal is sampled viva assessment, in which only a proportion of students are selected for oral verification. This preserves some deterrent and verification value without requiring every assessment to become a lengthy individual examination. (6) The oral exam then becomes less of a replacement for written assessment than a second layer of authentication.
Perhaps Universities Should Stop Trying to Detect AI
This leads to a more consequential question. Much of the institutional response to generative AI has focused on determining whether students used AI. That may eventually prove to be the wrong question.
As AI becomes embedded in writing software, search engines, research tools, operating systems, and professional workflows, distinguishing completely “AI-generated” work from completely “human-generated” work may become increasingly artificial.
Assessment design can instead focus on something universities can actually observe:
Can the student demonstrate ownership of the reasoning?
A student might use AI to brainstorm, organize information, test counterarguments, or improve language. Whether those practices are acceptable depends on the rules and learning objectives of the course. But eventually the student can still be asked:
Explain this. Defend this. Apply this somewhere else. Tell me where this argument could fail.
That approach changes the function of assessment. Instead of trying to construct an environment in which AI does not exist, universities design assessments around competencies that students must ultimately demonstrate themselves. Recent research is beginning to move in precisely this direction. Work on assessment in AI-rich environments increasingly argues for combining AI-supported learning with mechanisms through which students must subsequently demonstrate or defend their competence. (7)
The Return of the Human Question
There is something almost paradoxical about the direction higher education may now be moving. The more sophisticated machines become at producing answers, the more valuable it becomes to watch a human being explain one. Oral examinations will not solve every problem created by generative AI. They are expensive, difficult to standardize, potentially stressful, and unsuitable as universal replacements for written assessment. But their resurgence reveals something more important about what AI has changed.
For a long time, universities could reasonably treat the production of intellectual work as evidence of intellectual competence. Generative AI has weakened that connection. The challenge for universities is therefore not to make every assessment AI-proof. It is to design assessments in which understanding remains visible even when producing an answer becomes easy. In that environment, the most valuable question an examiner can ask may also be the simplest:
Why?
References:
1. Catherine Hartmann (11 Sep 2025): Oral Exams for a Generative AI World:
Managing Concerns and Logistics for Undergraduate Humanities Instruction, College Teaching, DOI: 10.1080/87567555.2025.2558563
2. UNSW Sydney. Teaching - Oral Assessment. https://www.teaching.unsw.edu.au/assessment-methods/oral-assessment
3. UNSW Sydney. Teaching - Oral Assessment. https://www.teaching.unsw.edu.au/assessment-methods/oral-assessment
4. Hu, H., & Hashim, H. (2026). Oral Assessment in the AI Era—Equity, Validity, and Scale. Educational Researcher, 55(5), 339-340.
5. Alcorn SR, Cheesman MJ. Technology-assisted viva voce exams: A novel approach aimed at addressing student anxiety and assessor burden in oral assessment. Curr Pharm Teach Learn. 2022 May;14(5):664-670. doi: 10.1016/j.cptl.2022.04.009. Epub 2022 May 17. PMID: 35715108.
6. University College London. Times Higher Education article explores AI and the future of university assessment. https://www.ucl.ac.uk/uclic/news/2025/mar/times-higher-education-article-explores-ai-and-future-university-assessment
7. Elkhodr M and Gide E (2026) Assurance by design: embedding the SAGE Defend step in AI-integrated higher education assessment. Front. Educ. 11:1872630. doi: 10.3389/feduc.2026.1872630

Comments