BlogEducators
How to tell if a student used AI: what the signs are worth, what detectors miss, and what to do instead
Every teacher now reads an essay and wonders. The honest answer is that no sign and no detector can settle it on its own: the signs have innocent explanations and the detectors flag the wrong students. Here is what each signal is actually worth, what the detector studies say, how to run a fair process when you suspect it, and how to take the question off the table on the devices your school manages.
The Nullgen team
12 min read
The essay is clean, the argument is tidy, and something about it does not sound like the student who wrote last month’s paper. Every teacher has had that moment since late 2022, and most have typed the same question into a search box: how do you tell if a student used AI?
The honest answer is that you usually cannot be sure, and the tools that promise certainty do not deliver it. That is not a reason to give up. It is a reason to know what each signal is worth, to use a process that is fair to the student in front of you, and to stop relying on after-the-fact judgment where you do not have to. This guide covers all three, with the research behind it.
The honest answer first
AI-written text has no fingerprint. A language model produces ordinary sentences, and a student who edits them produces sentences that are more ordinary still. Reading closely can raise a suspicion; it cannot confirm one. The same is true of the software that scores essays for AI: it produces a probability, and a probability is not a finding.
What you can establish is process. Drafts, version history, notes, an in-class writing sample, and a five-minute conversation about the argument tell you whether the student can account for the work. That evidence holds up with a parent, a principal, or an appeals panel. A detector percentage does not, and several universities have said so in writing.
Six signs worth noticing, and why each one can be innocent
Experienced teachers notice the same things, and they are worth noticing. Each one also has an explanation that has nothing to do with AI, which is why no single sign should start a disciplinary process.
The voice changed
Vocabulary, sentence length, and confidence jump compared with in-class work. Innocent version: a tutor, a parent, a writing center, or a student who finally cared about the topic.
Sources that do not exist
Plausible authors, real journals, and page numbers that lead nowhere. Models invent references fluently. Innocent version: sloppy note-taking, or a citation copied from a secondary source.
Fluent but empty
Every paragraph is correct and none of it refers to the class discussion, the assigned reading’s specific passages, or the student’s own examples. Innocent version: a student who did not do the reading and is writing around it.
Leftover prompt artifacts
A stray “Certainly!”, a line that addresses the user, Markdown headings in a Word document, or an essay that answers a slightly different question than the one assigned. This is the closest thing to a tell. Innocent version: a template the student was taught, or an autocorrect accident.
Perfect scaffolding, shallow analysis
Introduction, three balanced points, a conclusion that restates the introduction, and no risk taken anywhere. Innocent version: that is exactly the structure many students were taught.
No trail
A document pasted in whole with one revision, no notes, no outline, and a student who cannot say how the argument came together. Innocent version: a student who drafts by hand or in another app.
Notice the pattern: the signs that are easiest to spot are also the easiest for a student to remove, and the students most likely to trip them are often the ones writing in a second language or following a formula they were taught. That is the same problem the detectors have, at scale.
What AI detectors actually deliver
The first widely used detector was OpenAI’s own. It correctly identified 26 percent of AI-written text and flagged 9 percent of human writing as AI, and OpenAI withdrew it after six months for low accuracy. Commercial tools have improved on paper, but the independent record has not. A Stanford study ran seven popular detectors over essays by non-native English speakers and found they flagged most of them as AI-written, while the same detectors were nearly perfect on essays by US eighth graders.
61.3%
Average share of TOEFL essays by non-native English writers that seven detectors flagged as AI-generated. The same tools were near-perfect on essays by US eighth graders.
Liang et al., Patterns, 2023
9–15%
Share of published research abstracts, submitted exactly as written, that commercial detectors still flagged as AI-generated in a 2026 study. Non-STEM fields were flagged more often.
Karr et al., 2026
750
Papers a year that Vanderbilt estimated Turnitin’s claimed 1 percent false-positive rate would wrongly flag across 75,000 submissions, before it turned the detector off.
Vanderbilt University, 2023
The pattern repeats wherever someone tests independently. A 2023 study by a European team ran fourteen tools and found none reached 80 percent accuracy, and all of them got worse when the AI text was lightly paraphrased. Turnitin itself puts its document-level false-positive rate under 1 percent only for papers with 20 percent or more AI writing, and its sentence-level rate around 4 percent. And a 2026 study found the incentives backwards: abstracts that had been lightly polished with AI were flagged 38 to 80 percent of the time, while fewer than 4 percent of AI rewrites passed through a “humanizer” were caught.
The target moves every few months. Detectors are trained on the output of yesterday’s models. Each new model release, and each new “humanizer” service, resets the arms race in the student’s favor.
A paraphrase pass defeats them. Asking the model to rewrite in a plainer voice, or running the text through a second tool, is enough to drop detection to chance in most published tests.
Small error rates are large numbers. A 1 percent false-positive rate sounds small until it meets a school that grades fifty thousand assignments a year. Each of those is a student accused of something they did not do.
The errors are not random. Non-native writers, students who write formally, and students who follow a taught template are flagged more often, which turns a detection problem into an equity problem.
Watermarking will not rescue this either. Google now watermarks Gemini’s text output, but only Google can check that watermark, it covers only Google’s own models, and a student who rewrites the paragraph removes it. There is no technical path to a reliable verdict after the fact, which is why the rest of this guide is about process and prevention.
What to do when you suspect it: a fair process
A suspicion deserves a response, but the response has to survive being wrong. The steps below are the ones that hold up with students, parents, and appeal panels, and none of them depends on a detector score.
- 1
Compare with a known sample
Pull in-class writing, earlier assignments, or a quick paragraph written in front of you. A real gap in voice and ability is evidence; a detector percentage is not. If you have no baseline, start collecting one for every student now.
- 2
Check the sources
Look up every citation. Invented references and quotations that do not appear in the cited work are concrete, documentable problems whether or not AI produced them.
- 3
Ask for the process
Version history in Google Docs or Word, notes, outlines, and drafts. A genuine essay has a trail. Ask for it neutrally, as something you ask every student for, so the request itself is not an accusation.
- 4
Talk about the work, not the tool
Ask the student to explain the thesis, defend the weakest paragraph, or rewrite one section with you watching. A student who wrote it can. Lead with curiosity; most students will tell you what happened if the first question is not “did you cheat?”
- 5
Decide on what the student could not account for
Record what you found: the gap from the baseline, the missing trail, the sources that do not exist, the paragraph they could not explain. That record is the finding. If the evidence is thin, the answer is a conversation and a redo, not a referral.
- 6
Keep consequences proportionate and predictable
A first instance is a teaching moment and a rewrite under supervision. A pattern is a referral. Publish the ladder in advance so students know the stakes and no case turns on one teacher’s mood.
One rule to hold onto: never put a student’s name next to a detector percentage in an email or a referral. Once that number is in the file it tends to become the case, and it is the one piece of evidence that will not survive a challenge.
Make the question come up less often
The most reliable way to stop wondering whether a student used AI is to design assessments where the question has a built-in answer. Our guide to preventing AI cheating covers assessment design in depth; the short version is below.
Collect an in-class writing sample early in the term so every later assignment has a baseline.
Grade the process as well as the product: outline, draft, revision, and a short reflection on what changed and why.
Add a two-minute conversation or presentation to any high-stakes written assignment. It is faster than a detector and far more accurate.
State, assignment by assignment, what AI may and may not do. Students cannot follow a rule that lives only in the handbook.
On school devices, remove the question entirely
Everything above is about personal devices and work done at home. On the devices your school manages, the question does not have to come up. Nullgen’s free browser extension blocks the prompt box itself on ChatGPT, Claude, Gemini, Perplexity, and the 100+ other AI tools in its catalog, on Chromebooks, lab PCs, and classroom laptops. Research tabs, the learning platform, and the library catalog keep working; only typing or pasting into an AI prompt stops.
That changes the teacher’s job. Work done in the lab or on the Chromebook cart is the student’s own because the shortcut was never available, so there is nothing to detect and nobody to accuse. When a lesson calls for AI, a student requests access from the device and the teacher approves it for the period, the day, or the unit.
Force-install through Google Workspace, Jamf, or Intune, so Chromebooks and lab images enroll on their own and students cannot remove it.
Allow and deny policies per student, per class, or for the whole school, with every approval on a timer.
An Approver role for teachers: they decide requests without administering students or devices.
Managed iPhones and iPads join the same dashboard through the Nullgen app, which blocks known AI apps and websites on the device.
Two boundaries worth stating. Nullgen blocks known AI tools, and the list keeps growing; it does not read essays or detect AI writing. And it covers devices the school manages, not a student’s phone at home. Parents who want the same block on family devices can use the Family plan.
Key takeaways
No sign and no detector can prove a student used AI. What you can establish is whether the student can account for the work.
The classic signs are worth noticing, but each has an innocent explanation, and the students most likely to trip them are often writing in a second language.
Independent tests put detectors below 80 percent accuracy, with false positives concentrated on non-native writers and evasion as easy as a paraphrase.
When you suspect it: baseline, sources, process trail, a conversation about the work, then a proportionate response. Keep detector scores out of the file.
On school-managed devices, block the prompt instead. Work done there is the student’s own, and teachers approve AI when the lesson calls for it.
Frequently asked questions
Can an AI detector score prove a student cheated?
No. Detectors output a probability, independent tests put their accuracy below 80 percent, and their errors fall hardest on non-native English writers. Several universities, including Vanderbilt, have switched detectors off for exactly that reason. Use a score, if at all, as a reason to look more closely.
Can I ask a student to prove they wrote it?
You can ask for the process: drafts, version history, notes, and a conversation about the argument. Frame it as something you ask of everyone, and judge what the student can explain rather than demanding proof of a negative.
What about Grammarly, Google Docs suggestions, and other writing aids?
Decide and say in advance. Many schools allow spelling and grammar help but not generative rewriting. Students cannot follow a line you have not drawn, and detectors cannot tell the two apart either.
Is it safe to paste student work into a free detector?
Be careful. Student work is personal data, and a free tool may keep it or train on it. Check your district’s approved-vendor rules before uploading anything a student wrote.
Does Nullgen detect AI writing?
No. Nullgen prevents the prompt on devices the school manages: the AI tool’s input box stops working on Chromebooks and lab PCs, and known AI apps and websites are blocked on managed iPhones and iPads. It never reads essays or sees what a student types.
What about homework done at home on a personal device?
That is where assessment design and clear rules do the work, and where a conversation with parents helps. Parents can put the same block on family devices with the Family plan, and your school policy can make the expectations explicit.
Further reading
- OpenAI, “New AI classifier for indicating AI-written text,” January 2023 (withdrawn July 2023)
- Liang, Yuksekgonul, Mao, Wu, and Zou, “GPT detectors are biased against non-native English writers,” Patterns, 2023
- Weber-Wulff et al., “Testing of detection tools for AI-generated text,” International Journal for Educational Integrity, 2023
- Vanderbilt University, “Guidance on AI detection and why we’re disabling Turnitin’s AI detector,” August 2023
- Turnitin, “Understanding AI writing detection: false positive rates”
- Karr, Khvatskii, Hua, and Chawla, “Why AI Detection Fails for Academic Integrity,” arXiv, 2026
Stop guessing on school devices
Install the free extension on one lab image or classroom set and the question disappears there. When you are ready, the Education plan rolls it out through Google Workspace, Jamf, or Intune, with approvals in teachers’ hands.
