
An AI detector score is not evidence. Turnitin says so itself: the tool does not decide misconduct, the instructor does, and its error rate is not zero. If you wrote the work yourself, your defence does not rest on arguing about the technology. It rests on a documented writing process: version history, notes, research extracts and the ability to explain the text out loud.
Why detectors flag text written by humans
Detectors do not measure where text came from. They measure how predictable it is. If a sentence picks expected words in an expected order, it scores high no matter who wrote it.
Academic English is exactly that kind of text. Fixed phrases, hedged formulations, long sentences, no slang and no personal digressions. The better you have learned to write academically, the closer your style sits to what the model considers average.
Turnitin has measured this on its own data. In an update from its chief product officer it reports a sentence-level error rate of roughly 4 %, meaning four in every hundred highlighted sentences may be human. It also admits the error rate is higher in the first and last sentences of a document, which are typically the introduction and conclusion, where writing is at its most general.
How detection works technically is covered in our article on how AI-generated text gets detected. This one is about something else: what to do when the tool got it wrong about you.
Who gets flagged most often
The strongest data on false positives comes from the study GPT detectors are biased against non-native English writers (Liang et al., the journal Patterns, 2023). The authors ran essays through seven common detectors:
- Average error rate on TOEFL essays by non-native speakers: 61.22 %.
- Eighteen of 91 essays, that is 19.78 %, were flagged as machine-written by all seven detectors at once.
- On essays by native speakers, specifically US eighth-graders, the error rate was only 5.19 %.
The gap is not an accident. Writing in a second language means a smaller vocabulary and safer sentence constructions. To a detector that looks the same as machine text.
The Office of the Independent Adjudicator, which handles student complaints against providers in England and Wales, has seen the consequences. In its casework note on AI and academic misconduct, published on 15 July 2025, it tells providers to consider "whether assumptions about AI use could be biased against a student's writing style" if English is not the student's first language.
Two parts of a thesis carry the most risk, because they are formulaic by nature: the abstract, definitions of key terms, the methodology description and the conclusion. That is where the density of stock phrases per page is highest.
What the score actually means, and what it does not
The score says one thing: this text has statistical properties similar to machine-written text. It shows the marker where to look. That is all. It does not say you used AI, it does not say you committed misconduct, and it does not say anything another tool would confirm.
For work with more than 20 % flagged text, Turnitin reports a document-level error rate below 1 %, validated on 800,000 academic papers written before ChatGPT was released. In practice that is one wrongly flagged paper in a hundred. It sounds small until you are the one.
Below the 20 % threshold it is worse, and Turnitin admits it. Since July 2024 it does not display scores in the 1 to 19 % band at all, replacing them with an asterisk so nobody draws conclusions from the number. A 300-word minimum also applies, below which a document is not assessed.
Independent testing came out worse still. The study Testing of detection tools for AI-generated text (Weber-Wulff et al., International Journal for Educational Integrity, 2023) examined 14 tools including Turnitin and concludes that the available detectors are neither accurate nor reliable.
Turnitin does not make a determination of misconduct. We provide data so that educators can make an informed decision. Since our false positive rate is not zero, you will want to apply your professional judgment, knowledge of your students, and the specific context surrounding the assignment.
Turnitin, Understanding false positives within our AI writing detection capabilities
This is the most useful sentence in the whole article and you can quote it to your university. The vendor itself says its output is not a verdict.
Evidence you can prepare before you submit
A false accusation is not defeated by argument, it is defeated by records. Anyone who can show how the text came about barely needs a defence. Prepare it before you need it.
- Write in a document with version history. Google Docs saves version history automatically, Word does the same through OneDrive or SharePoint and you can open previous versions retrospectively. A hundred saved versions across four months is stronger evidence than any statement.
- Keep your notes and extracts. Photos of book pages, database exports, highlighted PDFs, a Zotero library. Sources you actually read can be verified. Invented sources do not exist.
- Supervise by email, not just in person. Send interim drafts to your supervisor by email even when you meet face to face. That gives you a dated trail showing the text grew gradually.
- Keep the raw data. Completed questionnaires, interview transcripts, the response table before processing. Your own data collection is the one thing a model cannot manufacture.
- Do not write the whole thing in one weekend. Even if you can, the evidence trail will look exactly like the thing you will be accused of.
What to do once you have been accused
Ask for the actual report
Request the name of the tool, the date of the check, the full score and the highlighted passages. Without it you do not know whether this is 22 % in one chapter or a flagged introduction. The difference is decisive, and the OIA is explicit that students "should be provided with all relevant evidence, including detection software reports, to allow them to respond effectively to allegations".
Do not confess pre-emptively
The common mistake is saying "maybe I had that one sentence translated" in an attempt to look cooperative. You have just supplied the intent that nobody had until then. Describe how you wrote, nothing more.
Submit the history of how the text was made
Version history, interim emails with your supervisor, notes and raw data. Present it all at once and ordered by date, not piece by piece on request.
Offer an oral check
Propose that you will explain any passage live. Someone who did not write the text cannot say why this particular definition is in it or where it came from. The OIA notes that providers "may decide to carry out a viva or use another mechanism to test the student's understanding of the work that was submitted". The same skill applies at your thesis defence, so you are practising either way.
Know where the burden of proof sits
This is the single most important point and it is not a matter of opinion. The OIA casework note puts it plainly:
The responsibility is on the provider to prove that the student has done what they are accused of doing, not on the student to disprove it.
Office of the Independent Adjudicator, Casework note: Complaints relating to AI and academic misconduct
The OIA also expects decision-makers to "understand the strengths and limitations of detection software, and weigh this evidence carefully against other available information". If a panel cannot point to what evidence beyond a score led to its conclusion, that is a procedural problem you can raise. Follow your provider's internal appeal route first, in the time limit stated in the outcome letter, because the OIA only reviews complaints once the internal process is complete.
What not to do
- Do not run the text through "humanizer" tools. You turn your own writing into writing that carries traces of machine editing. Since 27 August 2025 Turnitin also detects the use of detection-evasion tools and folds it into the score.
- Do not delete or rewrite your version history. Copying into a new document destroys the only evidence you have. Do not edit the original file either until the matter is resolved.
- Do not run your work through public detectors to test it. You hand your unpublished text to a third party and prove nothing. Scores differ between tools and a low number from another detector is not counter-evidence.
- Do not argue about the technology. A lecture on how detectors work will not persuade a panel. Your line is the writing process, not a critique of the tool.
How to lower the risk in future
University rules keep changing. The University of Edinburgh publishes guidance on generative AI for students, UCL sets out three categories of GenAI use in assessment and Cambridge maintains its own position on AI and academic misconduct. What applied two years ago may not apply now.
A practical minimum that costs you a few minutes:
- Ask your supervisor which tool the department uses and at what score work gets referred.
- Read your own institution's policy, not an article about another institution's policy.
- Declare and cite what you actually used. We covered how in our piece on whether students can use ChatGPT.
- Add your own data, specific examples and names. Formulaic writing is what a detector punishes.
A separate question is the watermark that AI tools now embed in their own output. It has nothing to do with detectors, and we explained the difference in our article on AI text watermarks and invisible characters.
Frequently asked questions
Is a high AI detector score proof that I cheated?
No. Turnitin itself says the tool does not determine misconduct, the educator does, and that its false positive rate is not zero. A score is a prompt for a conversation, not a verdict. The OIA is equally clear that the burden is on the provider to prove the allegation, not on you to disprove it.
What should I do first when my supervisor says my work was flagged as AI?
Ask for the actual report: the name of the tool, the date of the check, the overall score and the highlighted passages. Only then reply. Do not confess pre-emptively to anything and do not send any revised version of the text while all you know is that something was flagged.
Why was my text flagged when I did not use AI at all?
Detectors do not measure origin, they measure predictability. Academic language, fixed phrases and hedged formulations look statistically identical to machine text. The most commonly flagged parts are the abstract, definitions of terms, the methodology description and the conclusion, the most formulaic sections of a thesis.
Will version history from Google Docs or Word help me?
Yes, it is the strongest evidence you can produce. It shows the text grew over weeks, with rewriting and revisiting. The condition is that you wrote directly in that document and did not copy everything into a new file at the end.
Can my university expel me based on a detector score alone?
The score alone should not carry a case. Providers must be able to show what evidence supports the conclusion, and the OIA has partly upheld complaints where a panel could not explain what led it to decide AI had been used. You are entitled to see the evidence and respond before any decision is taken.
Are non-native English speakers really flagged more often?
Yes, and the effect is large. In the Liang et al. study the average false positive rate on TOEFL essays by non-native speakers was 61.22 % compared with 5.19 % on essays by US eighth-graders. The OIA expects providers to take this into account when English is not the student's first language.
Is it worth rewriting the text so it passes the detector?
No. Detection-evasion tools leave traces of their own and Turnitin has folded them directly into the score since August 2025. You would also destroy your version history, which is your best evidence. Adding your own data and specific examples is the better move.
