Can Turnitin Detect ChatGPT?

AI Detection · 13 min read · Updated 2026-08-21

Yes. Turnitin runs an AI writing indicator that is entirely separate from its plagiarism similarity score, and Turnitin's published documentation names the specific ChatGPT model versions it is trained to detect, from GPT-4o through the GPT-5 series. But that indicator is a probability estimate over the prose in your document, not proof: Turnitin states it does not make a determination of misconduct, it hides any score between 1 and 19 percent behind an asterisk, and it needs at least 300 words of prose before it produces a report at all. A high score is not evidence that you used ChatGPT, and a low score is not evidence that you did not.

Yes, and Turnitin publishes the list of ChatGPT versions it targets

Yes. Turnitin has run an AI writing indicator since April 2023. In its one-year anniversary press release the company said over 200 million papers had been reviewed, roughly 11% of them containing at least 20 percent AI writing.

What makes Turnitin unusually checkable is that it publishes the list. Its AI writing detection capabilities FAQ names the models the English-language detector is trained to detect, each with a release date: GPT-4o and GPT-4o-mini, the GPT-5 family including GPT-5-mini and GPT-5-nano, and the GPT-5.1 through GPT-5.4 releases, alongside Gemini, Claude, LLaMA, Mistral, DeepSeek and Grok models. The FAQ adds that it covers "tools based on these LLMs as well," which brings ChatGPT the product, not just the bare API, inside the boundary.

Read that list carefully, because it is a claim of coverage and not a measured hit rate. Turnitin publishes no per-model detection figure. "We can detect GPT-5.4" means that model was part of training and evaluation, not that every GPT-5.4 paragraph gets caught.

One practical fact first: the indicator is instructor-and-admin-only. Turnitin states that the AI writing detection indicator and report are not visible to students. You will not see your own number unless a teacher shows you the report.

The AI indicator and the similarity score answer different questions

Turnitin's FAQ is explicit that the two scores are "completely independent and do not influence each other." The similarity score is the percentage of matching text found when your document is compared against Turnitin's content database. The AI writing percentage is the share of qualifying prose the model predicts was generated or modified by AI. Different instrument, different question.

That matters more for ChatGPT than for almost anything else. When ChatGPT writes an essay it is not copying a source; it is generating new sentences. So raw ChatGPT output routinely produces a clean similarity report. Students see a low similarity number and read it as an all-clear. It is the wrong gauge, and it is exactly why Turnitin bolted a second indicator onto the same report.

The reverse holds too. A high similarity score on an AI-assisted paper usually means something ordinary: quoted material, the assignment prompt, a reference list. It says nothing about AI either way.

What Turnitin's model actually reads, and what it skips

Turnitin describes the pipeline plainly. Sentences are extracted from the submission and grouped into overlapping segments. Each segment is classified and given a value between 0 and 1 denoting the probability the text is human or AI-generated. Every qualifying sentence inherits its segment's score, and because segments overlap, one sentence can carry several scores that are pooled into one and aggregated into the document percentage.

"Qualifying text" is narrower than most students assume. Turnitin analyses only prose sentences, which it defines as blocks of text written in standard grammatical sentences, and explicitly excludes lists, bullet points and other non-sentence structures. Its documentation says the model does not reliably detect AI-generated text in the form of non-prose or code. Bibliographies have been excluded since an August 2023 release, which also began processing long-form prose inside tables.

There are hard requirements before any report exists: at least 300 words of prose in long-form format, no more than 30,000 words, a file under 100 MB, a supported language, and one of .docx, .pdf, .txt or .rtf. Miss any and the indicator shows a gray double dash instead of a number.

One consequence trips people up constantly. The percentage is a share of qualifying text, not of your document. Turnitin says so directly: the number "is not necessarily the percentage of the entire submission." A paper that is half bullet points gets judged on the other half.

Does it work on the newest GPT models?

Partly, and the honest answer is that this is a moving target by design. Turnitin's release notes read like a maintenance log: AI paraphrasing detection in December 2023, interactive report categories and the 30,000-word ceiling in July 2024, Japanese in April 2025, bypasser-tool detection in August 2025, then recall improvements in October 2025 and February 2026. The FAQ notes that in July 2026 Turnitin consolidated a multi-model ensemble into a single model while holding its stated false positive rate under 1%.

The underlying mechanism explains why that list keeps growing. A detector like this learns from samples of specific models' output, so when a genuinely new model ships, its writing sits outside the training distribution until the detector is retrained. OpenAI described the failure mode candidly when it ran its own classifier: "Classifiers based on neural networks are known to be poorly calibrated outside of their training data. For inputs that are very different from text in our training set, the classifier is sometimes extremely confident in a wrong prediction." OpenAI withdrew that classifier on 20 July 2023, citing its low rate of accuracy.

So the gap between a model's release and its appearance on a detector's list is real, but it is unpublished and it closes: Turnitin's list already carries GPT versions released in 2026.

One wrinkle: Turnitin says model updates are not applied retroactively, and a scored document's percentage changes only on resubmission. A paper scored in 2024 keeps its 2024 number until somebody runs it again, which an instructor can do at any time.

Does pasting, retyping or changing the file format change your score?

Almost none of the mechanical manoeuvres touch the AI indicator, because it scores the words, not how they arrived. Here is what each does.

Retyping by hand instead of pasting: no effect. The model reads the extracted text of your submission. Identical words produce an identical prediction.

Pasting into Word or Google Docs first: no effect on the indicator, but not nothing. Turnitin sells a separate product, Turnitin Clarity, unveiled at SXSW EDU in March 2025 and sold as a paid add-on from the third quarter of 2025. It shows educators pasted text, typing patterns, construction time and draft history, with video playback. Where an institution licenses it, how a document was written becomes visible, which the AI indicator alone never showed.

Saving as a PDF instead of .docx: no effect on the model's judgment, though it changes how text is extracted. Turnitin's best-practice note asks a class to submit in one format so processing stays uniform.

Swapping words for synonyms: this changes what a detector sees, because these classifiers score how predictable your wording is rather than who typed it. That cuts both ways. It is also why genuine writing with a narrow vocabulary gets flagged, which is the subject of a later section.

Breaking prose into bullet points: an essay rewritten as bullets is a worse essay, and it reads that way to the person marking it. Mechanically, non-sentence structures are excluded from qualifying text, so what shifts is the size of the pool being scored, not the model's opinion of your sentences.

Submitting text as an image or a scan: no extractable prose, so likely no report, just the gray double dash. That is not a clean result. It is a failed submission, and to an instructor it looks like one.

Coming in under 300 words: no report at all. Turnitin's release notes record a December 2023 bug that briefly let sub-300-word submissions through, and note that scores on those were less reliable. Short is not safe. It is unmeasured.

Does translating the text change anything?

There are two answers, and the second should worry honest students more than the first.

On coverage: Turnitin's AI writing detection runs on long-form English, Spanish and Japanese, and its paraphraser and bypasser detection is English-only. If a paper arrives in an unsupported language, Turnitin says the detector will not process the submission and no report is generated. That is a hole in coverage, not a clean bill of health, and which languages an institution accepts is its decision, not yours.

On round-tripping text through a translator, the evidence points firmly the wrong way. In a 2023 study in the International Journal for Educational Integrity, Debora Weber-Wulff and colleagues tested fourteen detection systems, twelve publicly available tools plus the commercial systems Turnitin and PlagiarismCheck, across 54 test documents. On text genuinely written by humans in English, the tools were correct 96% of the time. On text genuinely written by humans in Bosnian, Czech, German, Latvian, Slovak, Spanish or Swedish and then machine-translated into English with DeepL or Google Translate, the paper reports that accuracy dropped by 20%. Their explanation is blunt: machine translation "leaves some traces of AI in the output, even if the original was purely human-written."

The paper warns that using Google Translate or DeepL on your own work can lead to more false positives, leaving second-language students and researchers at risk of being falsely accused. Draft in another language and translate, and you raise your risk rather than lowering it, whether or not AI was involved.

What happens with a document that is part human, part ChatGPT?

This is the most common real situation and the one detectors handle worst.

Turnitin admits it. In a longer document mixing authentic and AI-generated writing, its documentation says, "it can be difficult to exactly determine where the AI writing begins and original writing ends." It positions the highlights as a place to start a conversation, not as a boundary line.

On short papers the behaviour is worse, and Turnitin publishes that too. In documents of only a few hundred words, it says, the prediction will be "mostly all or nothing" because the model works on a single segment with no overlaps to pool. Genuinely mixed text can therefore come back flagged as entirely AI-generated. A 400-word discussion post where you used AI for one paragraph is exactly that shape of document.

At the other end sits a floor: since July 2024, anything scored between 1% and 19% appears as an asterisk with no percentage and no highlights, because Turnitin considers low-range predictions too prone to false positives.

The figure that does appear is deliberately conservative. Turnitin scopes its false positive rate to documents with over 20% AI writing, and says that to hold that rate under 1% it accepts missing some AI writing, so the percentage an instructor sees is a floor rather than a measurement.

Since December 2023 the same indicator has also covered AI text subsequently run through a paraphraser, and since August 2025 through what Turnitin calls bypasser tools. Both are folded into the single AI percentage. That layer has its own answer in can Turnitin detect QuillBot.

Why neither a high nor a low score proves anything

Turnitin's own position is unambiguous, and students rarely hear it: the company "does not make a determination of misconduct; rather, it provides data for the educators to make an informed decision." Its FAQ says the percentage should not be the sole basis for action or a definitive grading measure.

A high score is not proof. Turnitin lists the genuine writing that trips it: content without much structural variation, text that literally repeats itself, and text paraphrased without developing new ideas. That describes a great deal of legitimate coursework. A 2023 paper in Patterns by Weixin Liang and colleagues at Stanford tested seven widely used publicly available GPT detectors, not the institutional tool Turnitin sells, on human-written TOEFL essays and US eighth-grade essays. The detectors handled the US essays well while misclassifying a majority of the non-native-authored ones, and the effect tracked vocabulary range rather than authorship: narrower vocabulary and steadier grammar are more predictable, and predictability is what these classifiers score. How the numbers fall, and how Turnitin's own compare, is the subject of how accurate Turnitin's AI detection is.

A low score is not proof either. Weber-Wulff's team concluded that the tools they tested "are neither accurate nor reliable and have a main bias towards classifying the output as human-written." Their discussion estimates that approximately half of AI-generated text that undergoes some obfuscation would likely be misattributed to humans; the two cases they ran, manual editing and machine paraphrasing, both landed well below the tools' performance on untouched text.

Some institutions decided the trade was not worth making. Vanderbilt University disabled Turnitin's AI detector in August 2023, stating it did "not believe that AI detection software is an effective tool that should be used," and noting that at a 1% false positive rate the 75,000 papers it submitted in 2022 would have meant roughly 750 papers incorrectly labeled.

How to check your own writing before you hand it in

Because the AI report is invisible to students, the only way to know how your prose reads is to run it yourself. No third-party detector predicts Turnitin's number: different model, different training data, different thresholds. What they all measure is the same underlying property, how predictable your sentences are, and that is the one you can act on.

Our free AI detector runs without an account and checks up to 300 words at a time. The limit is a useful coincidence: 300 words is also Turnitin's own minimum for producing a report, so checking in roughly that size mirrors the way the model scores overlapping segments rather than whole documents.

Check your introduction and conclusion first. Turnitin's May 2023 release notes record a higher incidence of false positives in the first few and last few sentences of a document, because openings and closings are so often written generically. Turnitin changed its logic to reduce that, but the cause is still yours to fix: generic scaffolding reads as machine-written because machines write generic scaffolding.

Then look at vocabulary. A short list of connectives and abstractions does a disproportionate share of the flagging, and an AI word checker surfaces them faster than reading for them will. Cutting them is a writing improvement rather than a detection tactic; the gain is a sentence that says the specific thing you meant.

What to do with what you find

Three situations, and they call for different things.

You wrote it yourself and it still reads as machine-written. Do not rewrite it into something artificially strange to satisfy a score. Fix the habits: vary sentence length, cut stock transitions, put in the specifics only you would know, such as the page you actually used or what went wrong in your experiment. If a flag has already landed, that is a separate problem with its own steps, covered in why your essay got flagged as AI.

It is AI-assisted and your course permits revising AI output. Then the work is to make the argument genuinely yours: restructure it, cut what you do not believe, add what only you can say. Our ChatGPT humanizer is built for that pass. Rewriting changes how prose reads; it does not change whether you followed your course's rules.

It is AI-generated and your course forbids it. No tool solves that, and pretending otherwise would be doing you harm. Turnitin runs its paraphrasing and bypassing detection automatically on every submission at institutions with the feature enabled. The options are to write it, or to talk to the person who set the assignment.

Whatever you find, read your actual policy. It varies by institution, by department and often by assignment; plenty of courses permit AI for brainstorming and forbid it for drafting. The score is a signal on somebody else's dashboard. The rule is what you are accountable to.

FAQ

Can Turnitin detect ChatGPT?

Turnitin runs an AI writing indicator trained to flag text from large language models, and its published documentation names specific ChatGPT model versions from GPT-4o through the GPT-5 series. The result is a probability estimate over the prose in your document, not a record of how it was written. Turnitin states that it does not make a determination of misconduct and that the percentage should not be used as the sole basis for action.

Does Turnitin detect GPT-5?

Turnitin's AI writing detection FAQ lists GPT-5, GPT-5-mini, GPT-5-nano and the GPT-5.1 through GPT-5.4 releases among the models its English-language detector is trained to detect. That is a statement of coverage rather than a published hit rate, because Turnitin does not release per-model detection figures. Newly released models generally sit outside a detector's training data until it is retrained, and Turnitin updates its model several times a year.

Does Turnitin's plagiarism score catch ChatGPT?

No, and this is the most common misunderstanding. The similarity score measures matching text against Turnitin's content database, while the AI writing indicator is a separate percentage; Turnitin states the two are completely independent and do not influence each other. Because ChatGPT generates new sentences rather than copying sources, AI-written text often produces a low similarity score, which is precisely why Turnitin added a second indicator.

Does Turnitin re-score an old paper when it updates its AI model?

Not on its own. Turnitin says model updates are not applied retroactively and that a submission's AI writing percentage changes only if the document is resubmitted. An instructor or administrator can resubmit a paper at any time, in which case it is scored by whichever model version is current, so an old score is a snapshot rather than a permanent verdict.

Can students see their own Turnitin AI writing score?

No. Turnitin states that the AI writing detection indicator and report are not visible to students, and that only instructors and administrators can see them. An instructor can download the AI report as a PDF and share it with you, but you cannot check your own Turnitin score in advance.

Does translating text change whether Turnitin flags it?

Turnitin only processes long-form English, Spanish and Japanese, and its paraphraser and bypasser detection is English-only, so an unsupported language produces no report rather than a clean one. Translation also appears to increase false-positive risk: a 2023 study in the International Journal for Educational Integrity reported that detection accuracy dropped by 20% on genuinely human writing that had been machine-translated into English, concluding that machine translation leaves traces detectors read as AI.

What does a 20% AI score on Turnitin mean?

It means the model estimated that roughly a fifth of the qualifying prose in your submission was likely generated or modified by AI. Twenty percent is also the display floor, because Turnitin suppresses scores between 1% and 19% and shows an asterisk instead, considering low-range predictions too prone to false positives. The percentage covers only prose sentences, not bullet points, code or bibliographies, so it is not a share of the whole document.

Can Turnitin tell exactly which sentences were written by AI?

Not precisely. Turnitin highlights the segments its model predicts were AI-written, but states that in a document mixing authentic and AI writing it can be difficult to determine exactly where the AI writing begins and original writing ends. In documents of only a few hundred words the prediction is mostly all-or-nothing, because the model scores a single segment with no overlapping segments to pool, so genuinely mixed text can be flagged as entirely AI-generated.

Is a Turnitin AI flag proof that a student cheated?

No. Turnitin's own documentation states that it does not make a determination of misconduct and that the AI writing percentage should not be used as the sole basis for action. Genuine student writing can be flagged, particularly formulaic prose, repetitive text and writing by non-native English speakers, and some institutions including Vanderbilt University disabled the feature over exactly that concern.

Try Humanit free

Rewrite AI text to read human, then verify with the built-in detector.

Open the humanizer