We the editors hold the following position: Turnitin's AI writing indicator is a useful signal that warrants further investigation when it is high. It is not, on its own, sufficient evidence to support a finding of academic dishonesty. Schools and faculty that treat it as the latter create disparate-impact problems, expose themselves to defensible appeals, and undermine the trust that makes integrity policies work in the first place.
Our editorial position
Turnitin's AI score is a useful signal that warrants further investigation when high. It is not sufficient evidence on its own for an academic-integrity finding. Schools that treat it as the latter create disparate-impact problems and undermine the trust that makes integrity policies work.
This piece lays out where the AI score belongs in the decision process and where it does not.

What the Score Is and Isn't
Turnitin's AI indicator is a percentage estimate of how much of a submission was likely generated by a large language model. It is computed by Turnitin's service from statistical signals in the text, primarily perplexity and burstiness. The score is generated without seeing the student's writing history, the assignment context, the rubric, or the conditions under which the work was produced. It looks only at the text in front of it.
That narrow input is what makes the score useful as a triage tool and dangerous as a verdict. The score can identify text that looks statistically similar to AI output. It cannot identify whether the writer intended to deceive, whether the writer had legitimate accommodations that affect their writing style, whether the writer is a non-native English speaker whose prose happens to look AI-like (see the Stanford 2023 study), or whether the writer used permitted AI assistance under the assignment's actual policy.
When the Score Reasonably Drives Action
A high AI score warrants a conversation. That is the lowest-cost, highest-information next step. It is also where most reasonable academic-integrity policies are calibrated to begin. The conversation is not an accusation. It is the moment where the score, the writing, and the student come together to surface whether something further needs to happen.
Beyond a conversation, action is justified when the AI score is corroborated by other evidence:
- A submission that differs substantially in style and vocabulary from the student's prior work, when prior samples are available
- A version history that shows the entire document pasted in as one event, with no iterative editing
- Citations that point to nonexistent sources (a common ChatGPT failure mode)
- A student account of how they produced the work that does not match the work itself
Any one of these in conjunction with a high AI score is a much stronger basis for action than the score alone.
When the Score Should Not Drive Action
A high score on its own, with no corroborating evidence, is not enough. This is true for several practical reasons:
- The base rate of false positives is non-trivial. Independent reviews put the document-level false positive rate in single digits for native English writers and substantially higher for non-native English writers.
- The score is not reproducible across runs. Turnitin updates its model. The same text submitted weeks apart can score differently. An institutional finding cannot rest on a snapshot that may not survive re-checking.
- The methodology is not publicly documented. Turnitin has not published a peer-reviewed paper describing its detection approach. A student appealing a finding cannot effectively contest a methodology they have no access to.
- The disparate-impact risk is real. ESL students, students writing on formulaic topics, and students whose voice happens to be more uniform will be flagged at higher rates regardless of actual AI use.
What Turnitin Itself Recommends
To Turnitin's credit, its own published guidance has shifted toward the same position. Turnitin's support documentation and Educator Network materials are explicit that the AI score should not be used as sole evidence of academic misconduct. Instructors are directed to view the score as a starting point for inquiry. That framing matches our editorial position. Institutions whose policies have not caught up with Turnitin's own guidance are operating beyond what the tool's vendor recommends.
Turnitin's own support documentation and Educator Network materials are explicit that the AI score should not be used as sole evidence of academic misconduct. Instructors are directed to view the score as a starting point for inquiry.
Turnitin published guidance
The Bottom Line
Use Turnitin's AI score for triage. Use a conversation with the student to surface what the score cannot see. Use corroborating evidence to support any finding that goes beyond a conversation. Do not let the score, by itself, become the case. The institutions that adopt this discipline produce fewer wrongful findings, fewer successful appeals, and more durable trust with their students.
For more on how to handle a flagged case when the evidence is contested, see our appeals guide. For the detector that publishes its methodology and reports a 0.5% non-native English false positive rate, see the Proofademic review.
Have feedback or a topic to suggest? Reach the editorial team at our contact page.
Frequently Asked Questions
Should I act on a high Turnitin AI score?
A high score warrants a conversation with the student. It does not, by itself, support a finding of academic dishonesty. Always look for corroborating evidence before any disciplinary action.
What Turnitin AI percentage is too high?
Turnitin doesn't publish an official threshold. Many institutions treat 20% or higher as worth investigating, but the score should be a starting point for inquiry, not a verdict on its own.
Can I fail a student based only on a Turnitin AI score?
Turnitin's own guidance says no, not on the score alone. A finding of academic dishonesty requires corroborating evidence beyond the detector output.
What if my school requires action on AI scores?
Push for a policy revision. Acting on detector scores alone exposes the institution to defensible appeals, disparate-impact problems for ESL students, and undermines trust in the integrity system.
Is Turnitin's AI score reliable?
Reliable as a starting signal, not as conclusive evidence. False positive rates are non-trivial, especially for non-native English speakers (see <a href="https://www.cell.com/patterns/fulltext/S2666-3899(23)00130-7" target="_blank" rel="noopener noreferrer" className="text-primary underline">Stanford 2023</a> findings).


