Teacher and student in a one-on-one conversation across a desk
Back to Blog
Editorial
July 8, 20267 min read

Student Says They Didn't Use AI But the Detector Flagged Them. Now What?

It's one of the hardest situations any teacher faces with AI detection. Here is the step-by-step playbook for handling it fairly, what to document, and when to escalate.

First: slow down. The detector score creates a sense of urgency, especially when the percentage is high. But an academic-integrity decision made in the moment after seeing a flag is much harder to defend than the same decision made after a 24-hour pause, a conversation with the student, and a review of the available evidence.

First thing: slow down

The score is one input. The student's denial is another. Neither alone is a verdict. Document the score, have the conversation, look for corroborating evidence before any action. The process below is the playbook.

The score is one input. The student's denial is another. Neither alone is a verdict. The process below is what we recommend after years of covering AI detection in classrooms, drawing on published research on detector reliability and the standing guidance from major academic-integrity offices.

Teacher and student in a one-on-one conversation across a desk
An honest conversation with the student is the lowest-cost, highest-information next step after a flag.

The Step-by-Step Playbook

The process scales to whatever your institutional policy allows. Most of it works in a high school classroom as well as a university seminar.

  1. Read the submission yourself, slowly. Before consulting the score, read the work as a piece of writing. Does the voice, vocabulary, and argument match what you know about this student? If you have prior samples, compare. The detector cannot see this context.
  2. Note the score and the date. AI detector scores are not always stable across runs. Take a screenshot or download the report. If the case progresses, you will need the exact score the original decision was based on.
  3. Have the conversation in person, calmly. Open with what you are seeing rather than an accusation: "The detector flagged a portion of this submission. Walk me through how you wrote it." A student who actually wrote the work can usually explain their process, their sources, and which sections gave them trouble. A student who relied heavily on AI typically cannot.
  4. Listen to the explanation. Some explanations are immediately disqualifying (sources that don't exist, claims about effort that don't match a document version history). Some are reasonable and call the detector score into question (an ESL student whose writing is naturally more uniform, a student who used permitted spell-check or grammar tools, a student whose draft history shows iterative editing).
  5. Look at version history if available. Google Docs, Microsoft 365, and most modern writing platforms keep a revision history that shows how the document was built. A document with extensive iterative editing is hard to fake. A document pasted in as a single block is its own signal.
  6. Check the citations. ChatGPT and similar tools regularly fabricate citations. If the submission claims sources, verify a sample. Nonexistent citations are stronger evidence of AI use than the detector score itself.

What to Document

Whatever decision you reach, write it down before it leaves your desk. The documentation makes the decision defensible to an appeals committee, an administrator, or a parent later.

  • The detector score and date, with a screenshot
  • The specific sentences or sections the detector flagged
  • Your reading of the submission as a piece of writing, before consulting the score
  • The student's explanation in their own words, summarized after the conversation
  • Any corroborating evidence (version history, citation issues, comparison to prior work)
  • Your decision and the reasoning

Date and sign it. Keep it. If the case escalates to a hearing or appeal, this documentation is the case.

When to Escalate

Some cases warrant escalation to academic-integrity offices or department heads regardless of how the conversation goes:

  • Fabricated citations the student cannot account for
  • A document version history showing the entire submission pasted in at once with no editing
  • A submission that diverges sharply from the student's prior demonstrated writing ability
  • A pattern across multiple assignments rather than a single instance

Escalation is not the same as adjudication. You are flagging a case that warrants institutional review, not declaring guilt.

When to Step Back

Several signals should make you reconsider the case:

  • The student is a non-native English speaker. Stanford 2023 documented a 61.3% false positive rate on TOEFL essays. Vendor improvements since then are not well-documented for most detectors.
  • The student walks through their writing process credibly and matches what you would expect from someone who did the work.
  • The version history shows extensive iterative editing.
  • The score is the only evidence. No corroborating signal exists.

Treating a high score as the case rather than as a triage signal is where most wrongful findings happen.

Most wrongful AI-detection findings happen when teachers treat a high score as the case rather than as a triage signal, without corroborating evidence and without an honest conversation with the student first.

Editorial summary

The Bottom Line

The hardest version of this situation is when a thoughtful student denies using AI and the detector strongly disagrees. In those cases, the right move is the slower one: gather corroborating evidence, document the process, and only act when the case rests on more than the detector score alone. Students whose first language is not English are at the highest risk of being wrongfully flagged; institutions that have policies in place for that population are positioned better than those that don't.

For the broader appeals workflow, see our appeals guide. For more on what to do when the detector is the only evidence, see Questioning AI Detection.

Have feedback or a topic to suggest? Reach the editorial team at our contact page.

Frequently Asked Questions

What should I do if a student denies using AI but the detector flagged them?

Slow down. Document the score, have a conversation with the student about their process, look for corroborating evidence (version history, citation accuracy, comparison to prior work) before any disciplinary action.

Can a student be falsely accused by an AI detector?

Yes. Stanford's 2023 study (Liang et al.) found a 61.3% average false positive rate on TOEFL essays across seven commercial detectors. False positives are real, especially for non-native English speakers.

What evidence beyond the AI score should I gather?

Document version history, citation accuracy (AI tools frequently fabricate citations), comparison to prior student work, the student's account of their process, and any pattern across submissions.

When should I escalate an AI detection case?

When evidence beyond the score exists: fabricated citations, document pasted in as one block with no editing history, a sharp divergence from prior work, or a pattern across multiple assignments.

How do I document an AI detection case fairly?

Save the score and date with a screenshot, note flagged sections, write the student's explanation in their own words, record corroborating evidence, and document your reasoning. Sign and date the record.