Homework12 min readLast updated

Are AI detectors accurate? What a Turnitin AI score really means for your kid

TL;DR

AI detectors estimate how statistically predictable a piece of writing is and infer AI authorship from that. They produce both false positives (honest students writing formally, and non-native English speakers especially) and false negatives (AI text lightly edited). Treat a detector score as a reason to ask questions, never as evidence. The reliable check is process: drafts, version history, and whether your child can explain their own argument.

What an AI detector actually measures

No tool can look at a sentence and see where it came from. There is no watermark in ordinary AI text and no signature left behind. What detectors like Turnitin's do instead is score how predictable the writing is — how often each word is the statistically likeliest next word given everything before it. Machine-generated text tends to sit in that predictable band. The detector then converts that score into a percentage that looks a great deal more like evidence than it is.

That distinction matters enormously if you are the parent on the receiving end of an email. "87% AI" does not mean 87% of the essay was written by AI. It means the model's confidence that the text falls in the predictable band. Those are different claims.

So how accurate are they?

Accurate enough to be useful as a flag, nowhere near accurate enough to be used as proof. Vendors publish low false-positive rates measured on clean test sets — long, unedited, single-author documents. Real student work is none of those things. It is short, edited, sometimes co-written with a parent or a tutor, often produced in an unfamiliar formal register because a teacher asked for one.

Two failure directions matter to families:

  • **False positives** — honest work that reads as predictable. Formal five-paragraph structures, template-driven lab reports, and the writing of students learning English get flagged disproportionately. Research has repeatedly found non-native English speakers penalized at higher rates.
  • **False negatives** — AI-generated text that a student has lightly rewritten, reordered, or run through a paraphraser passes routinely. The kids most deliberately cutting corners are the least likely to be caught.

The combination is the uncomfortable part: the tool is most likely to catch the honest kid with a stiff writing voice and least likely to catch the one gaming it.

Why your kid can get flagged for work they wrote

The most common causes we hear from parents are mundane. Your child wrote to a rubric, so every paragraph has the same shape. They used a graphic organizer the teacher provided. They were told their last essay was too casual, so this one is stripped of personality. They used Grammarly's rewrite suggestions. They researched with AI legitimately, absorbed its phrasing, and reproduced the cadence from memory.

None of those are cheating. All of them raise a predictability score.

What to do if your child is accused

Second opinion

Run the flagged work through here instead.

Detectors give you a percentage. This gives you what actually looks like your kid — and what doesn't.

Sign in to run your free checkFree account, no card needed. One Homework Check on us.

Stay off the detector's turf. Arguing about the percentage puts you in a debate you cannot win and implicitly accepts the number as the thing being judged. Argue about process instead.

  • **Pull the version history.** Google Docs and Word both keep it. A document that grew over three sittings with revisions, deletions and typos looks nothing like one pasted in whole.
  • **Collect the process artifacts.** Outlines, notes, browser history, the photo of the sticky-note plan on the kitchen table.
  • **Ask for the school's policy in writing.** Specifically: is a detector score sufficient on its own to substantiate an academic integrity finding? Most policies say it is not.
  • **Have your child explain the argument out loud.** Not defensively — for you first. A child who can explain why paragraph three follows paragraph two wrote paragraph three.
  • **Request a conversation, not an appeal, first.** Most teachers are looking for reassurance, not a conviction.

The habit that prevents all of this

Keep the process visible before anyone asks. Drafting in a document with version history on, keeping the outline, and writing a one-line note of what AI was used for — "asked it to explain the causes of the war, wrote the essay myself" — costs nothing and settles almost every dispute before it starts. That disclosure habit is the core of our AI homework rules for families.

For the wider question of where the line between help and cheating actually sits by assignment type, read is using AI for homework cheating? and our parent's guide to AI homework help.

How the main detectors differ

  • **Turnitin AI writing indicator** — used by most schools, bundled with plagiarism checking. Reports a percentage of sentences it reads as AI. Turnitin's own guidance says the score is an indicator for review, not evidence.
  • **GPTZero** — free and widely used by individual teachers. Reports perplexity and burstiness. More sensitive, and correspondingly more false positives on formal student writing.
  • **Copyleaks** — sentence-level highlighting, often used in higher education. Highlighting looks precise, but it is the same predictability scoring underneath.
  • **Originality.ai** — aimed at publishers rather than schools; tuned to catch AI aggressively, which is the wrong trade-off for a 13-year-old's essay.
  • **"Humanizer" tools** — the paraphrasers students use to defeat all of the above. They work, which is the clearest evidence that detector scores measure style rather than authorship.

The email to send the teacher

Short, unemotional, and asking a procedural question rather than defending your child:

Hi [Teacher] — thanks for flagging this. [Child] and I have gone through the assignment together and they've walked me through how they built the argument. We have the Google Docs version history showing the draft as it developed, plus their outline. Before we go further, could you tell me what the school's academic integrity policy requires as evidence in a case like this — specifically whether a detector score on its own is sufficient? Happy to come in and have [Child] talk through the essay with you.

It works because it concedes nothing, offers evidence, and puts the policy question in writing. The full sequence, including what to do if it escalates, is in flagged for AI: what to do.

What the research actually says

Stanford researchers found that detectors classified more than half of TOEFL essays written by non-native English speakers as AI-generated, while near-perfectly clearing essays by US-born eighth graders. Vanderbilt University disabled Turnitin's AI detector for its instructors, citing false positives and the absence of any way to audit a score. OpenAI withdrew its own AI text classifier in 2023 for low accuracy. Those three facts, together, are the honest state of the field: the vendors closest to the models are the least confident in detection.

A faster read than a detector

If you just want to know whether your own child understood the work in front of you, a predictability score is the wrong instrument. Paste the piece into the Homework Check on our homepage and it will tell you, in plain language, what stood out and — more usefully — the two or three questions to ask your kid tonight. It never assigns a percentage, because a percentage was never the point.

ShareXLinkedInFacebook

Frequently asked questions

Are AI detectors accurate?

Not accurate enough to be used as proof. AI detectors score how statistically predictable a text is and infer AI authorship from that. They generate false positives on formal, template-driven, or non-native-English student writing, and false negatives on AI text that has been lightly edited or paraphrased.

Does Turnitin detect AI?

Turnitin's AI writing indicator estimates the proportion of a document that reads as AI-generated based on predictability, not on any detectable signature. Turnitin itself advises that the score is an indicator for further review rather than evidence of misconduct.

What does an 87% AI score mean?

It does not mean 87% of the work was written by AI. It reflects the model's confidence that the text sits in the statistically predictable band associated with machine-generated writing. It is a probability estimate about style, not a measurement of authorship.

Can a student be falsely accused by an AI detector?

Yes, and it happens regularly. Students writing to a rubric, using a provided graphic organizer, writing in an unfamiliar formal register, or learning English are all flagged at elevated rates despite doing the work themselves.

How do I prove my child wrote their own essay?

Use process evidence rather than arguing about the score: document version history from Google Docs or Word, saved outlines and notes, and your child explaining the argument in their own words. Most school policies state that a detector score alone cannot substantiate an academic integrity finding.

Which AI detector is most accurate?

None is accurate enough to be treated as proof. Turnitin is the most widely used in schools and the most conservative, GPTZero is more sensitive and produces more false positives, and Copyleaks and Originality.ai are tuned for publishing rather than student writing. All four score predictability rather than authorship.

Why do AI detectors flag non-native English speakers?

Writers using a second language tend to reach for common, safe phrasing, which is exactly the statistically predictable pattern detectors treat as machine-like. Stanford researchers found detectors misclassified more than half of TOEFL essays by non-native speakers as AI-generated while clearing comparable native-speaker essays.

What should I write to the teacher if my child is flagged?

Send a short, calm email that offers process evidence — version history, outlines, an offer for your child to talk through the essay — and asks in writing what the school's policy requires as evidence, specifically whether a detector score alone is sufficient. Avoid arguing about the percentage itself.

Have a specific piece of homework in mind?

Paste it — or a screenshot of it — into the Homework Check on our homepage and ask your question. You'll get a plain-language read on whether the thinking looks like your kid's, what stood out, and the questions worth asking them. First check is free.

Keep reading