Remote oral assessment has quietly become high stakes. Universities run thesis defences over video. Certification bodies award credentials on a spoken answer. Employers hire from interviews held entirely in a browser.

Voice exam fraud is the result, and it takes more than one form. In each case the institution trusts two things at once. That the speaker is the person enrolled, and that the words are their own. Generative AI weakened the second one fast, faster than most assessment policies were rewritten. A candidate no longer needs to know the material. They need a second screen and the ability to read naturally.

Scope note: this article describes these methods so institutions can defend against them. It is written for assessment designers, academic integrity teams, and security staff choosing a proctoring or voice verification system.

01 / The gap

Why conventional proctoring misses most of this

Online proctoring, also sold as remote proctoring, was built for written online exams on a screen. It watches for visual cues: a second face, a glance away, a change of window, a second monitor. Those cues worked well against the cheating of the time.

A spoken exam breaks that model twice. First, the answer itself is what matters, and proctoring never looks at it. Second, glancing away is normal when someone is thinking. The strongest visual cue loses its meaning, because a candidate reading from a device just out of frame looks like a candidate remembering.

The seven methods below are ordered roughly by how often institutions report them. Each is paired with the signal that actually catches it.

02 / The methods

Seven ways a remote voice exam gets faked

1. Proxy impersonation

What it looks like
Someone else sits the exam from the start. A paid expert takes a certification viva, or a stronger student defends a weaker one's work.
Why proctoring misses it
Nothing looks wrong. With no invigilator in the room, one person sits alone and answers well for the full session.
What catches it
Speaker verification against an enrolment sample. The live voice is matched to a recording the real candidate made under supervision. A mismatch settles it, however good the answers are.

2. Reading AI-generated answers aloud

What it looks like
The real candidate sits the exam and speaks in their own voice. A second device out of shot runs a chatbot such as ChatGPT, and they read its answer back.
Why proctoring misses it
Every visual signal is clean. Voice matching passes too, because it really is the enrolled candidate speaking.
What catches it
Content analysis of the answer. Read-aloud text has tells: even pacing, tidy sentences, textbook words, no false starts. The answers are complete but thin on the candidate's own work.

3. A hidden earpiece with a live AI assistant

What it looks like
Software listens to the question through the candidate's own machine and writes an answer. It reads that answer into an earbud, and the candidate repeats it.
Why proctoring misses it
There is no second screen to glance at, so the gaze stays natural. There is no second person either, so a room scan proves nothing.
What catches it
Content analysis, plus the pause before the answer. The spoken answer carries the same tells as method 2. A consistent gap while the assistant writes is often the clearer signal.

4. Off-mic coaching

What it looks like
Someone sits out of shot and feeds answers, whispered or written on a card. The candidate repeats them.
Why proctoring misses it
Room scans happen once, at the start. A coach who arrives later, or was never in shot, stays invisible.
What catches it
Both signals, but only partly. Voice matching catches the coach only if they speak loudly enough to register. Detection is weaker here than for the other six, and an assessment board should be told that.

5. Mid-session handover

What it looks like
The real candidate starts the exam and passes the identity check. A substitute takes over for the hard section.
Why proctoring misses it
The check ran once, and nothing retests it. A few seconds with the camera blocked covers the swap.
What catches it
Speaker verification on every answer, not once per session. Each answer is matched against the enrolment sample on its own. A handover shows up on the substitute's first answer.

6. Pre-recorded or played-back audio

What it looks like
Answers are recorded in advance and played into the call. Software stands in for the microphone.
Why proctoring misses it
The audio is real speech from a real person. If the candidate recorded it, voice matching passes too.
What catches it
Acoustics and timing. Played-back audio misses the room sound of the live channel and starts too promptly. Varying the question wording also makes this much harder, and costs nothing.

7. Voice cloning

What it looks like
A cloned copy of the candidate's voice speaks answers written by someone else. This is an audio deepfake, and it now needs only a short sample such as a public talk.
Why proctoring misses it
Visual proctoring is irrelevant here. A good clone is built to defeat voice matching, the very check most institutions rely on.
What catches it
Synthetic speech detection, separate from voice matching. Matching asks whether two voices are the same, not whether either is real. Any system relying on voice identity alone is vulnerable here.
03 / The pattern

Two questions cover all seven

Read the seven together and they come down to two questions. Both are asked of every answer, not once per session.

MethodIs the enrolled candidate speaking?Did the candidate write the answer?
1. Proxy impersonationNo, and this alone catches itNot the deciding signal
2. AI answers read aloudYes, passes cleanlyNo, and this alone catches it
3. Hidden earpiece assistantYes, passes cleanlyNo, and the pause adds a second signal
4. Off-mic coachingSometimes, if the coach is audibleNo, with weaker confidence
5. Mid-session handoverNo, from the swap onwardNot the deciding signal
6. Played-back audioPasses if self-recordedCaught by timing and acoustics
7. Voice cloningPasses by designNeeds synthetic speech detection

Neither question is enough alone. A system that only checks identity is blind to methods 2 and 3, the fastest growing. A system that only reads answers is blind to methods 1 and 5, the oldest. The two checks fail in opposite directions, so you need both.

This is the design VivaGuard is built around. It confirms the enrolled candidate is the one answering, and flags read-aloud and AI-generated answers. It runs on the institution's own servers, so no audio, transcript, or result leaves them.

04 / Evaluation

What to ask a vendor

  • Is identity checked once, or on every answer? A check at the start cannot catch a swap that happens later.
  • Is answer content analysed at all? If answers are only transcribed and stored, methods 2 and 3 pass every time.
  • Is synthetic speech detection separate from voice matching? Ask directly. The two are often sold as one thing, and method 7 defeats matching alone.
  • Where does the audio go? Ask which subprocessors receive recordings, in which countries, and for how long.
  • What does the system output when it is unsure? A flag for human review is fine. An automated verdict on a thesis defence is not.
  • What is the false-positive rate on accented speech? Ask who it was tested on. Voice matching and text classifiers both show accuracy gaps across accents.

That last point matters most. A missed cheat means a credential awarded wrongly. A false positive means a student accused of something they did not do. Those two costs are not equal. Set the system up accordingly, with human review before any consequence.

05 / FAQ

Frequently asked questions

What is a viva exam?

A viva, short for viva voce, is an oral exam. The candidate answers aloud instead of in writing. Universities use vivas to defend a thesis or verify coursework. Professional bodies use them for certification, and employers use the same format for structured technical interviews. Because the answer is spoken, a viva tests whether a candidate can explain their own work under questioning.

Can online proctoring detect AI-generated answers?

Conventional proctoring watches the room and the screen through a webcam. It spots a glance away, a second face, or a tab change. It does not read the content of a spoken answer. A candidate reading a chatbot answer from a device out of shot triggers no alert, because nothing visible has gone wrong. Catching that needs analysis of the answer's own language, looking at its phrasing and structure.

What is speaker verification, and how is it different from voice recognition?

Speaker verification, sometimes sold as voice biometrics, asks who is talking. It is the identity verification step, and it compares live speech against a sample the candidate enrolled earlier. Speech recognition asks what was said, turning speech into text. An exam system needs both: verification to confirm the enrolled candidate is answering, and content analysis to judge who wrote the answer.

Why do institutions want self-hosted exam proctoring?

Exam audio is personal data, and often biometric data. A voice sample used to verify identity counts as biometric under the EU GDPR and India's DPDP Act. Sending recordings to a third-party cloud creates a cross-border transfer, a retention duty, and a subprocessor to disclose. Universities, certification bodies, and regulated employers increasingly ask for systems that run on their own infrastructure, so no audio, transcript, or result leaves their servers. This is usually called an on-premise or air-gapped deployment, and it is the simplest answer to a data residency requirement.

Does recording a candidate's voice require consent?

In most places, yes. Voice enrolment for identity checks is usually treated as biometric processing. That carries a higher consent standard than ordinary personal data. Tell candidates before enrolment what is recorded, what it is compared against, how long it is kept, and who can see it. Offer a documented alternative for anyone who declines. Confirm the exact requirement with your own counsel, as the rules differ by jurisdiction.

06 / Conclusion

Design the assessment, not just the software

No system removes the need to design the assessment well. Varying the question wording blunts methods 6 and 7. Asking follow-ups about the candidate's own work blunts methods 2 and 3, because a generated answer is fluent in general and thin on detail. Requiring them to reason from something only they have seen weakens all seven.

Detection and assessment design work together. Good software makes the easy methods hard. A well-designed viva makes the hard ones pointless.

VivaGuard confirms the enrolled candidate is the one answering and flags read-aloud or AI-generated answers. It runs on your own servers, so no audio or transcript leaves them. See how it works →

← Back to Blog