Teacher reviewing student essays on a laptop in late afternoon light
Back to Blog
AI Detection
June 24, 20267 min read

Does Turnitin Detect ChatGPT? What the Published Tests Show

Turnitin says yes. The published independent testing is more complicated, and the paraphrasing weakness documented in 2023 has not gone away.

Sometimes. Turnitin's AI writing indicator is designed to flag text it identifies as likely generated by large language models including ChatGPT, GPT-4, GPT-5, and similar systems. It produces a percentage score on submitted documents and a sentence-level breakdown for users who view the full report. But how often that score is correct, and how often it misses or false-flags, varies considerably depending on the text it is looking at.

Quick answer

Sometimes. Turnitin catches lazy AI use (ChatGPT output pasted in unedited) most of the time. It misses paraphrased AI text often. And it false-flags some human writing, especially from non-native English speakers.

The honest answer is that Turnitin detects some ChatGPT-generated writing, misses some, and incorrectly flags some human writing. Below is what the publicly available evidence shows.

Teacher reviewing student essays on a laptop in late afternoon light
Turnitin's AI score sits in the assignment view, computed on Turnitin's servers, rendered by your LMS.

Turnitin's Own Claim

Turnitin released its AI writing detection feature in April 2023, integrated into Feedback Studio and Originality. Their published positioning has remained largely consistent: they claim high accuracy on clear-cut AI-generated submissions, acknowledge a false positive rate they describe as "low," and recommend that teachers treat the AI score as one signal among many rather than as definitive evidence of misconduct.

Turnitin has not published a peer-reviewed research paper documenting its detection methodology. The accuracy figures shared in their marketing and support documentation are not paired with disclosed test sets, sample sizes, or methodology details that would let an independent reviewer reproduce the results. This is the industry norm for commercial AI detectors. It is also the central reason teachers should not take vendor accuracy claims at face value.

What Independent Tests Found

The most cited independent evaluation of AI detectors in K-12 and higher education contexts is Liang et al. (Stanford, 2023), which tested seven commercial AI detectors against TOEFL essays written by non-native English speakers. The result that shaped the field: across those seven detectors, the average false positive rate on TOEFL essays was 61.3%. One detector incorrectly flagged 97.8% of those essays as AI-generated. Turnitin's detector shares the underlying perplexity-and-burstiness approach that drove that result.

The Washington Post, EdSurge, and Inside Higher Ed have all run smaller-scale evaluations in the years since, with broadly consistent findings: Turnitin catches clearly AI-generated text reasonably often, but it false-flags writing by non-native English speakers, formulaic academic prose, and student writing that happens to share stylistic features with ChatGPT's output more than the marketing suggests.

Turnitin's own published guidance has shifted to reflect this: the AI score is now framed as an indicator that warrants further investigation, not as proof of misconduct on its own.

The Paraphrase Weakness

Krishna et al. (2023) published a study that has held up well: they showed that running AI-generated text through a paraphrasing tool (their model, DIPPER, but the result generalizes) drops detector accuracy dramatically. Specifically, DIPPER-paraphrased text reduced DetectGPT's accuracy from 70.3% to 4.6% at a 1% false positive rate. Turnitin and most other commercial detectors use signal families similar to DetectGPT's, so the result transfers.

DIPPER-paraphrased text reduced DetectGPT's accuracy from 70.3% to 4.6% at a 1% false positive rate. Turnitin and most commercial detectors use signal families related to DetectGPT, so the weakness generalizes.

Krishna et al., 2023

The practical implication: a student who pastes ChatGPT output into QuillBot, Wordtune, or a similar paraphrasing tool before submitting can defeat Turnitin's AI detection most of the time. This is not a defect in Turnitin specifically; it is an open research problem for AI detection as a category.

Detectors that analyze writing at the sentence level rather than only at the document level offer more resilience here, because per-sentence stylometric signals are harder to uniformly disguise than document-wide statistics. Proofademic documents this approach in its May 2026 research paper, which reports its sentence-level analysis is more robust to paraphrasing than perplexity-only methods. The paper is explicit that humanizer robustness is an open research area, not a solved problem.

What This Means for Teachers

Three practical takeaways:

  • Turnitin will catch lazy AI use. A student who copies ChatGPT's response and pastes it directly into the assignment box is likely to get flagged. Many cases are this simple.
  • Turnitin will miss careful AI use. A student who runs ChatGPT through a paraphrasing tool, or who edits the output substantially before submitting, will often defeat the detector.
  • Turnitin will false-flag some human work. Non-native English writers and students whose writing happens to be structurally similar to ChatGPT's baseline output are at elevated risk. The Stanford 2023 result remains the single most important data point teachers should know.

The right institutional posture is to treat the AI score as one piece of evidence in a broader judgment process, never as the sole basis for an academic-integrity accusation. Our appeals guide walks through how to document the decision when the score is high but the case is uncertain.

The Bottom Line

Does Turnitin detect ChatGPT? Yes, sometimes. Often enough to catch students who use AI without disguising it, often enough to false-flag students who didn't use AI at all, and not often enough on paraphrased AI text to be considered reliable on its own. The 2023 published research on detection limitations has not been substantially overturned in 2026. Treat the AI score as a starting point, not a verdict.

For the full evaluation, see our Turnitin review. For an AI detector that publishes its methodology and addresses the non-native English false-positive problem head-on, see the Proofademic review and its May 2026 research paper.

Have feedback or a topic to suggest? Reach the editorial team at our contact page.

Frequently Asked Questions

Can Turnitin detect ChatGPT?

Sometimes. Turnitin's AI score is designed to flag LLM-generated text. It catches lazy AI use (unedited paste-ins) most of the time, misses paraphrased AI output often, and false-flags some human writing.

How accurate is Turnitin AI detection?

Turnitin claims high accuracy but hasn't published peer-reviewed methodology. Independent reviews put document-level accuracy in the 85-92% range, lower than marketing figures.

Can Turnitin detect ChatGPT if you paraphrase?

Often, no. Krishna et al. (2023) showed paraphrasing tools drop detector accuracy from 70% to under 5% in some tests. Turnitin shares this weakness with most commercial detectors.

Does Turnitin detect GPT-4 and GPT-5?

Turnitin updates its model to cover newer LLMs but doesn't publish per-model accuracy. Detection is generally weaker on the newest frontier models than on earlier GPT-3.5-era outputs.

What percentage on Turnitin AI score is concerning?

Turnitin doesn't publish an official threshold. Many institutions treat 20% or higher as worth investigating, but the score should always be a starting point for conversation, not a verdict on its own.