Do AI Detectors Unfairly Flag Non-Native English Writers?

AI Detection · 13 min read · Updated 2026-08-21

Yes, for the widely used detectors tested in the strongest published research on the question. A 2023 peer-reviewed study in Patterns by Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou at Stanford ran 91 human-written TOEFL essays and 88 US eighth-grade essays through seven commercial AI detectors, and found they flagged an average of 61.22% of the non-native essays as AI-generated against an average of 5.19% of the native-written ones. The mechanism is perplexity: detectors treat predictable word choice and uniform sentence structure as machine-like, and writing carefully in a second language produces exactly those statistics. Vendors including Turnitin, GPTZero and Originality.AI dispute that the bias applies to their own current models, and a 2024 ETS study using purpose-built detectors on GRE essays found no such gap, but the risk is real enough that universities including Vanderbilt and Waterloo have switched their AI detectors off.

Yes, and the strongest study on it put a number on the gap

In 2023, Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou at Stanford published "GPT detectors are biased against non-native English writers" in Patterns, a Cell Press journal. They ran 91 human-written TOEFL essays and 88 US eighth-grade essays through seven commercial AI detectors. The detectors handled the American schoolwork well, misclassifying an average of 5.19% of it. On the TOEFL essays, written by people sitting an English proficiency test, the average false positive rate was 61.22%.

The per-student picture is worse than the average. Eighteen of the 91 TOEFL essays were flagged by all seven detectors at once. Eighty-nine of the 91 were flagged by at least one. If your university licenses the detector that dislikes your particular essay, the average is not what happens to you. The paper is from 2023, its data and code are public, and nothing published since has removed the mechanism underneath it.

What the Stanford study actually tested

The seven detectors were Originality.AI, Quil.org, Sapling, OpenAI's own classifier, Crossplag, GPTZero and ZeroGPT, all accessed on 15 March 2023. The non-native samples were 91 TOEFL essays collected from a Chinese educational forum. The comparison group was 88 US eighth-grade essays from the Hewlett Foundation's Automated Student Assessment Prize dataset, published on Kaggle in 2012.

The authors released every essay they used, so the word counts can be checked directly. The length difference is the most important fact about this dataset and it rarely gets reported. The 91 TOEFL essays run from 62 to 148 words, with a median of 104, and every one of them is under 150 words. The 88 US eighth-grade essays run from 129 to 677 words, with a median of 378.5.

The American essays are three and a half times longer at the median, and the study's critics are right to raise it. Length is one of the strongest inputs to a detector's confidence, and 104 words gives a statistical model very little to work with. The comparison group is a confound too: TOEFL candidates are mostly adults writing to a clock in a second language, while the ASAP essays were written by American thirteen- and fourteen-year-olds in their first.

So read the finding precisely. It is strong evidence that off-the-shelf detectors in early 2023 failed badly on short second-language writing, not a clean measurement of language background alone. The authors said as much; their limitations section notes the small samples and that most detectors tested ran on GPT-2.

Why careful second-language writing looks statistically like a machine

Detectors of that generation scored two things. Perplexity is how surprising each word is given the words before it. Burstiness is how much sentence length and structure vary across a passage. Language models are trained to produce the likely next word, so their output sits at low perplexity and low burstiness: smooth, even, safe.

Now think about writing an academic paragraph in your second language. You reach for constructions you are confident in. You choose the word you know is correct over the word you suspect might be better. You keep sentences to a shape you can control. Every one of those is a competent decision, and every one lowers perplexity.

This is not a new observation, and it did not come out of AI research. The Stanford paper grounds it in applied linguistics going back decades: Laufer and Nation on lexical richness in 1995, Jarvis on lexical diversity in 2002, Lu and Ortega on syntactic complexity, Biber, Gray and Poonpon on grammatical complexity. The detectors did not invent the pattern. They found a real one and mislabelled it.

The label is the whole problem. "Low perplexity" describes a text. "AI-generated" is a claim about a person. A detector converts the first into the second, and that conversion is where second-language writers get hurt.

The experiment that showed it was the language, not the author

The paper's most convincing part is a pair of experiments that moved the bias back and forth on command. The team asked ChatGPT to rewrite the TOEFL essays with the instruction "Enhance the word choices to sound more like that of a native speaker." Same authors, same arguments, same essays, richer vocabulary. The average false positive rate fell from 61.22% to 11.77%, measured perplexity rose significantly, and the number of essays flagged by all seven detectors dropped from 18 to one.

Then they ran it backwards, asking ChatGPT to "Simplify word choices as if written by a non-native speaker" on the American eighth-graders' essays. Misclassification of that group climbed from 5.19% to 56.65%. American schoolwork, made lexically plainer, started tripping the same alarms.

One further result takes AI out of the picture entirely. The team analysed 1,574 papers accepted at ICLR 2023, chosen because both its submission deadline and author response period fell before ChatGPT was released on 30 November 2022. Authors based in non-native-English-speaking countries wrote abstracts with significantly lower perplexity than authors in native-English-speaking countries, and the difference held after controlling for review ratings. The signature detectors read as "AI" was already sitting in human academic writing before the chatbot existed.

What the detector companies say back

Turnitin was not one of the seven, and it says so. Its position is that the study used TOEFL practice essays all under 150 words, and that its own detector returns no prediction below a 300-word minimum. That objection is fair; the counts above confirm it.

Turnitin has published its own research: across close to 2,000 writing samples from English language learners, a false positive rate of 0.014 (1.4%) for ELL writers against 0.013 (1.3%) for native English writers on documents meeting the 300-word minimum, a gap it describes as not statistically significant. Separately it publishes a document-level false positive rate below 1% for documents scored at 20% AI writing or higher, plus a sentence-level rate of around 4%. Those two document-level figures cover different sets of documents, every ELL submission over the 300-word minimum in one case and only documents already scored at 20% AI or above in the other, so 1.4% and the sub-1% number are not comparable and do not contradict each other. All are the vendor's own internal figures, not an independent audit, and they read best alongside what Turnitin's published accuracy numbers do and do not cover.

GPTZero responded directly, publishing "ESL Bias in AI Detection is an Outdated Narrative" in October 2023 and re-running the Stanford study's own code against its updated model. It reports flagging one of the 91 TOEFL essays as AI, a further 6.6% marked uncertain or possible AI, and under 2% false positives on a larger ESL set it assembled itself.

Originality.AI's rebuttal is methodological: the study tested version 1.1 of its model rather than the current one, 91 forum-sourced essays are too small a sample, and comparing TOEFL candidates against American eighth-graders introduces a confound. It reports a 5.04% false positive rate across more than 1,500 human-written essay samples.

Take those responses seriously, then notice what they share. Each is a vendor grading its own product on data it chose, with no external replication. That does not make them wrong. It does mean a student cannot verify it, and neither can a professor.

The study that found no bias, and why it does not settle the question

The literature genuinely disagrees, and that deserves representing. In August 2024, Yang Jiang, Jiangang Hao, Michael Fauss and Chen Li, researchers at ETS, which administers the GRE, published "Detecting ChatGPT-generated essays in a large-scale writing assessment: Is there a bias against non-native English speakers?" in Computers & Education.

They built their own detectors on GRE writing-assessment data, combining linguistic features from ETS's e-rater engine with perplexity features. On human-written essays, 0.12% of native English speakers' essays were misclassified as AI-generated, against 0.006% of non-native speakers'. Native speakers were flagged slightly more often, and the authors found no evidence of bias disadvantaging non-native writers.

That result points the other way, and it is worth looking closely at what it is a result about. These were purpose-built detectors, trained on a corpus from the exact assessment they were deployed on, with fairness engineered in deliberately: the same team's companion paper at the AIED 2024 conference compares three bias-mitigation strategies, balancing the training data, removing sensitive features and adjusting thresholds. That is not what a general-purpose detector does to your coursework. Fair detection is achievable under controlled conditions; that is not the same as your institution's licensed detector having achieved it.

The bias finding also keeps reappearing. In June 2025, Ahmad Pratama published a study in PeerJ Computer Science testing GPTZero, ZeroGPT and DetectGPT on academic abstracts, finding that non-native English speakers faced higher false positive rates, with the most accurate tool showing the strongest bias against particular author groups and disciplines. Accuracy and fairness turned out not to be the same axis.

Who this actually lands on

The affected population is larger than "international students," though they are the most exposed. The Institute of International Education's Open Doors report, released in November 2025, counted 1,177,766 international students in US higher education in 2024/2025, around 6% of total enrolment. Add domestic students who speak another language at home, graduate researchers drafting in English for the first time, applicants writing personal statements in a language they learned at sixteen. The Pratama study was about journal abstracts, not coursework, so the same penalty follows researchers into publication.

The consequences are not evenly distributed. A domestic student flagged for AI use faces an academic-integrity process. An international student on a study visa faces the same process with immigration status attached to the outcome, in a second language, often without family nearby. Same false positive. Very different stakes.

Which universities have switched their AI detectors off

Vanderbilt University disabled Turnitin's AI detector on 16 August 2023 and published its reasoning. Vanderbilt submitted 75,000 papers to Turnitin in 2022, and at the 1% false positive rate Turnitin claimed at launch, roughly 750 would have been wrongly flagged. It also cited the absence of any explanation of how the detector reaches a verdict, and noted that AI detectors more often label text by non-native English speakers as AI-written.

The University of Waterloo discontinued Turnitin's AI detection functionality in September 2025. Its published rationale states that the tools are unreliable and "have also been found to be biased toward students whose first language is not English," citing peer-reviewed work. Its own testing was inconclusive and included cases where, in the university's words, "the product flagged human written text as 100% generated by AI."

Both replaced detection with process: clear AI policies, citation when AI is used, assignments redesigned around in-class writing and course-specific material, and comparing a submission against a student's earlier work. Waterloo put it bluntly. "Time and effort are best spent on education rather than policing misuse of GenAI."

What to do before you submit

If you write in English as a second language, the most useful habit is to stop discovering what a detector thinks of your work at the same moment your professor does. Check it yourself first. Humanit's free AI detector needs no account and takes up to 300 words per check, which is also the minimum Turnitin requires before it will score a document at all. Run your introduction and conclusion through as separate passages: opening and closing paragraphs are where generic framing language lives, and generic framing is statistically smooth in any perplexity-based detector. Turnitin has acknowledged that it saw a higher incidence of false positives in the first and last few sentences of documents, and says it has since adjusted how its model handles them.

Then read the number as a diagnostic, not a verdict. A high score tells you your prose is statistically smooth. It cannot tell you that you did anything wrong. The fixes are ordinary good-writing fixes: vary sentence lengths on purpose; swap generic transitions for ones that carry an argument; name specific things, because names, dates and figures from your sources are unpredictable in a way that abstract summary is not. A stock AI vocabulary checker shows which words in your draft are the ones models lean on.

And keep your process. Draft in something that records revisions, so the essay exists as a dated timeline rather than one finished file.

If you have already been flagged and English is your second language

A flag is a probability estimate about a text, not a finding about a person, and every major vendor says so in its own documentation. The general playbook for responding when your genuine essay is flagged applies to you the same as anyone. But you have one thing to add that a native-speaker classmate does not, and it is worth having ready in plain words rather than improvising under pressure. Something close to this:

"I wrote this essay myself, and I can show you how. Here is the version history from the document, from the first outline through to the final draft. Here are my notes and the sources I used.

I also want to raise something specific to my situation. There is peer-reviewed research on this: a 2023 study in Patterns by Liang and colleagues at Stanford found that AI detectors flag writing by non-native English speakers far more often than writing by native speakers, because these tools score how predictable the vocabulary and sentence structure are, and writing carefully in a second language produces that pattern. English is not my first language. I am not asking you to ignore the score. I am asking that it not be the only evidence, which is what the detector vendors themselves recommend.

If it would help, I am happy to talk through the argument of this essay, or to write on a related question under supervision."

That last offer is the strongest move you have: a person who wrote an essay can explain why they structured it that way. Ask as well for your institution's written policy on how detection scores may be used. Many now say explicitly that a score alone is not sufficient grounds.

The uncomfortable paradox the researchers left us with

Liang and his co-authors ended their paper with a problem nobody has since resolved. If detectors penalise limited linguistic variety, then the way for a second-language writer to avoid being falsely accused of using AI is to make their vocabulary richer, and the fastest way to do that is to use an AI tool. In their words, to evade false detection these writers "may need to rely on AI tools to refine their vocabulary and linguistic diversity."

That is a bad position to put students in, and it is worth being straight about where it leaves a tool like ours. If your institution permits AI help with language and you disclose it under their rules, running your own writing through a rewriter to vary sentence rhythm and word choice is legitimate, no different in kind from a writing-centre appointment. If it does not permit that, using one to disguise authorship is misconduct, and no score from any tool changes that. Where the line sits is worth working out before you are near it, because detection is probabilistic in both directions and a low score is not permission.

The research is public, peer-reviewed and specific enough to cite in a policy meeting, which is the one lever that reaches further than your own draft. Institutions that read it responded by treating a score as the start of a conversation rather than as evidence.

FAQ

Do AI detectors really flag non-native English speakers more often?

The most cited peer-reviewed study on the question, published in Patterns in 2023 by Liang, Yuksekgonul, Mao, Wu and Zou at Stanford, found that seven commercial detectors flagged an average of 61.22% of human-written TOEFL essays as AI-generated, compared with an average of 5.19% of US eighth-grade essays. Detector vendors dispute that this still applies to their current models, and a 2024 ETS study using purpose-built detectors on GRE essays found no such gap. The bias is well documented for the off-the-shelf detectors tested in 2023 and contested for specific products today.

Why does writing in a second language look like AI to a detector?

Most detectors of that generation score perplexity, meaning how predictable each word is given the words before it, and burstiness, meaning how much sentence length and structure vary. Writing carefully in a second language pushes you toward vocabulary and sentence patterns you are confident in, which lowers both measures. Language models also produce low-perplexity text, so the two end up with a similar statistical signature despite having nothing else in common.

Does Turnitin discriminate against English language learners?

Turnitin says no. It has published research testing close to 2,000 writing samples from English language learners and reported a false positive rate of 0.014 (1.4%) for those writers against 0.013 (1.3%) for native English writers on documents meeting its 300-word minimum, which it describes as no statistically significant difference. Turnitin separately publishes a document-level false positive rate below 1%, but that figure covers only documents already scored at 20% AI writing or higher, so the two rates are measured over different sets of documents and should not be compared with each other. All of these numbers come from Turnitin's own internal testing rather than an independent audit.

Was Turnitin included in the Stanford bias study?

No. The seven detectors tested were Originality.AI, Quil.org, Sapling, OpenAI's classifier, Crossplag, GPTZero and ZeroGPT, all accessed in March 2023. Turnitin has pointed out that the 91 TOEFL essays used were all under 150 words and that its own detector requires at least 300 words before it returns a score, so it would not have produced predictions on that dataset.

Have any universities stopped using AI detectors because of this?

Yes. Vanderbilt University disabled Turnitin's AI detector in August 2023, citing the volume of papers it submits each year, the lack of transparency about how the detector works, and research showing that non-native English writing is more likely to be labelled AI-written. The University of Waterloo discontinued the same functionality in September 2025, stating explicitly that the tools are unreliable and biased toward students whose first language is not English. Both replaced detection with assignment redesign and clearer AI policies.

Did any research find the opposite result?

Yes, and it deserves to be taken seriously. In August 2024, Yang Jiang, Jiangang Hao, Michael Fauss and Chen Li of ETS published a study in Computers and Education using GRE writing-assessment data, in which 0.12% of native English speakers' essays and 0.006% of non-native speakers' essays were misclassified as AI-generated. Those were purpose-built detectors trained on the exact assessment they were deployed on, and the same team's companion paper at the AIED 2024 conference compares three strategies for mitigating bias in them. That is not how a general-purpose commercial detector is applied to coursework.

My score is high but I wrote it myself. What should I change?

Treat the score as a diagnostic rather than a verdict, because a high number only tells you your prose is statistically smooth. The changes that help are ordinary writing improvements: vary sentence lengths deliberately instead of keeping them uniform, replace generic transition words with ones that carry an argument, and add concrete specifics such as names, dates and figures from your sources, since specific detail is less predictable than abstract summary. None of that alters your argument, and none of it is a trick.

What should I say if I am accused of using AI and English is my second language?

Bring the evidence of process first: version history, outlines, notes and sources, which together show the essay being built over time. Then raise the research directly, naming the 2023 Patterns study by Liang and colleagues, and explain that detectors score the predictability of vocabulary and sentence structure rather than authorship, which is a pattern careful second-language writing produces. Ask that the score not be the sole evidence, which is what detector vendors themselves recommend, and offer to discuss your argument in person or write on a related topic under supervision.

Is a low detector score proof that my writing is fine?

No. Detection is probabilistic in both directions, so a low score is not evidence of anything and it is certainly not permission. Different detectors disagree with each other on the same text, and their verdicts change as models are updated. Treat any score as one weak signal about style, and treat your institution's written policy as the thing that actually governs what you are allowed to submit.

Try Humanit free

Rewrite AI text to read human, then verify with the built-in detector.

Open the humanizer