AI Detection

Can Turnitin Detect Claude or Gemini? [Tested]

By Alex Chen

Short answer: yes. Both get caught. We ran Claude and Gemini output through Turnitin in March 2026 and saw ~96% detection on Claude, ~97% on Gemini. Switching from ChatGPT to a different model won't save you. That said, the models aren't identical under the hood, and the small statistical differences between them are worth understanding if you want to know why some text is marginally harder for detectors to pin down.

Last updated: March 26, 2026

How We Tested

30 essays per model. Six models total: Claude 3.5 Sonnet, Claude 4 Opus, Gemini 1.5 Pro, Gemini 2.0, GPT-4o, and GPT-4 Turbo. Each essay tackled a different academic subject (English, history, psychology, biology, business, political science, sociology, philosophy, nursing, computer science) at roughly 1,500 words. Plain prompts, no tricks, no "write like a human" instructions.

Every essay hit Turnitin through an institutional account. We logged the AI detection percentage and classified anything north of 20% (Turnitin's own investigation threshold) as "detected."

Detection Results by Model

AI Model Detection Rate Average AI Score Lowest Score Highest Score
GPT-4o 98% 94% 18% 100%
GPT-4 Turbo 98% 93% 15% 100%
Gemini 2.0 97% 91% 17% 99%
Gemini 1.5 Pro 96% 89% 14% 98%
Claude 4 Opus 96% 88% 12% 97%
Claude 3.5 Sonnet 95% 86% 11% 96%

A few things pop out of this data. GPT-4 models get flagged the most reliably, averaging above 93% AI scores. Gemini lands in the middle. Claude sits at the bottom of the pack with the lowest averages and the widest spread. One Claude 3.5 Sonnet essay scored just 11%, which would clear Turnitin's threshold entirely.

Why Claude Is Slightly Harder to Detect

That 2-3 point gap isn't random. It traces back to measurable statistical quirks in how Anthropic's models generate text.

Higher Perplexity

Claude text averaged a perplexity of 28-32 in our measurements. GPT-4 sat lower, around 20-25. Both are well under the human range of 60-120, but Claude's slightly elevated scores push more of its sentences into the detector's gray zone where confidence drops.

Why the difference? Anthropic's training pipeline seems to produce marginally more varied word selections. Claude might write "examine" where GPT-4 defaults to "look at," or pick "roughly" instead of "about." Individually these swaps are tiny. Across 1,500 words they compound into a detectable (or rather, slightly-less-detectable) pattern.

More Natural Burstiness

Claude's sentence-length standard deviation averaged 5.2 words. GPT-4's: 3.8. Human writers typically hit 8-15. So Claude's burstiness is still recognizably AI-like, but it inches closer to the human range than GPT-4 does.

Claude also drops in more parenthetical asides, fragments, and structural detours than GPT-4 tends to. None of this makes the text sound genuinely human. It does, however, weaken the statistical signal just enough to occasionally push a sentence below the flagging threshold.

Less Rigid Structure

GPT-4 essays are almost aggressively formulaic. Thesis in the intro. Topic sentence opening every body paragraph. Conclusion that parrots the thesis. Claude is more likely to break that template. It might open a section with a question, lead with an example rather than a claim, or vary paragraph lengths more noticeably. Small structural rebellions, but they matter to a classifier trained on predictable patterns.

Why Gemini Looks Like GPT-4 to Detectors

Different company. Different architecture. Different training team. And yet Gemini's output is statistically near-identical to GPT-4 when you put a detector on it. How?

Nearly the same perplexity: Gemini text averages 22-27, overlapping heavily with GPT-4's 20-25. Both models gravitate toward high-probability word sequences at every token position.

Nearly the same burstiness: Gemini's sentence-length standard deviation: 4.1. GPT-4's: 3.8. Both churn out sentences in the 15-25 word range with minimal variation.

The same vocabulary defaults: "A significant factor." "This demonstrates that." Both models reach for the same stock academic phrases at remarkably similar frequencies.

The convergence makes sense once you think about it. Both models trained on massive overlapping slices of the public internet and optimized for the same goal: helpful, fluent answers. That optimization process funnels the output toward similar probability distributions regardless of what the underlying architecture looks like.

Detection Rates Across Subjects

Here's something we didn't expect: subject matter mattered more than model choice. The topic of the essay shifted detection rates more than which AI wrote it.

  • Highest detection rates: History (98%), English literature (97%), sociology (97%). Humanities writing by humans is naturally varied, so AI uniformity sticks out sharply
  • Moderate detection rates: Psychology (95%), business (95%), political science (94%)
  • Lowest detection rates: Computer science (91%), nursing (92%), biology (93%). Technical fields use constrained vocabulary and rigid sentence forms even in human writing, which narrows the gap between AI and human statistical profiles

This held across every model we tested. The least detectable combination in our entire dataset was a Claude-written computer science essay. Claude's inherently higher variance plus technical writing's inherently lower baseline compounded to push detection accuracy down to about 89-91%.

Can You Evade Detection by Switching Models?

No. Three percentage points separate the most detectable model (GPT-4 at 98%) from the least detectable major model (Claude 3.5 at 95%). That's not a strategy. That's a rounding error in your odds of getting caught.

We've seen students try mixing models. Generate with Claude, then "polish" with GPT-4 (or flip it). In practice this backfired: detection rates actually climbed to 97-98%, because the blended output lost whatever marginal variance the single model had. You end up with the statistical fingerprints of two AI systems layered on top of each other.

Open-source models (LLaMA 3, Mistral) with cranked-up temperature settings can pull detection as low as 90-93% on some essays. Still a 90%+ chance of getting flagged. Not viable.

What Actually Works Against Turnitin (Regardless of Model)

Model-hopping is a dead end. The methods that actually reduce detection are the same whether your source was GPT-4, Claude, or Gemini.

Serious manual rewriting: Rewrite 40-50% of the sentences in your own voice and detection drops to 70-85%. This works because your writing genuinely has different statistical properties than any AI. The tradeoff is obvious: it eats almost as much time as writing from scratch.

Purpose-built humanization: Tools like Anti-Turnitin reshape perplexity, burstiness, and entropy distributions to land in the human range. In our tests, Anti-Turnitin hit a 97% pass rate on Turnitin regardless of source model. GPT-4, Claude, Gemini, didn't matter. After humanization the pass rates converged.

Full tool-by-tool breakdown in our AI humanizer comparison.

Will Future Models Naturally Evade Detection?

Probably not. Counterintuitive, but hear this out. Each new model generation (GPT-4 to GPT-5, Claude 3.5 to Claude 4) has actually been more detectable, not less. Why? Better models produce more fluent, more polished text. Fluent and polished means lower perplexity and lower burstiness. Exactly the signals detectors hunt for.

Could a lab train a model specifically to mimic human statistical patterns? Theoretically. Nobody has. And the incentive structure works against it: "less predictable" output is also "less helpful" for most tasks. AI companies want their models to sound good, not to sound sloppy in the specific ways humans do.

Meanwhile, Turnitin keeps updating. Their January 2026 classifier refresh added training data from Claude 4 and Gemini 2.0, closing the minor detection gaps that existed at those models' launch. Expect that cycle to repeat with every major release.

Where This Leaves You

Turnitin catches Claude at ~96% and Gemini at ~97%. A hair below GPT-4's ~98%, but all three sit comfortably above the threshold that triggers an investigation. No publicly available AI model produces text that reliably slips past Turnitin unedited. The 2-3 point gaps between models are interesting trivia and nothing more.

If you're shopping for an AI model specifically to dodge detection, you're asking the wrong question. Which model you used matters far less than what happens to the output afterward. Raw AI text gets flagged regardless of its origin. Only genuine rewriting or purpose-built humanization alters the statistical patterns that detectors actually measure.

Need to make Claude or Gemini text undetectable? Try Anti-Turnitin free. It works on output from any model.

Frequently Asked Questions

Can Turnitin detect Claude AI?
Yes. In our March 2026 testing, Turnitin detected unedited Claude 3.5 and Claude 4 output with approximately 96% accuracy. That is slightly lower than GPT-4 detection (98%) because Claude produces text with marginally higher perplexity and more sentence-length variation. But 96% is still high enough that submitting raw Claude output will almost certainly get flagged.
Can Turnitin detect Google Gemini?
Yes. Turnitin detects unedited Gemini output at approximately 97% accuracy as of March 2026. Gemini's text has statistical patterns very similar to GPT-4, with low perplexity and uniform sentence lengths. The detection rate is only 1 percentage point lower than GPT-4, making Gemini essentially as detectable as ChatGPT for practical purposes.
Is Claude harder to detect than ChatGPT?
Slightly. Claude's detection rate is about 2 percentage points lower than GPT-4 (96% vs 98%). The difference comes from Claude's slightly more varied sentence structure and marginally higher perplexity scores. However, a 96% detection rate is still extremely high — the difference is not large enough to make Claude a meaningfully safer choice for avoiding detection.
Which AI model is hardest for Turnitin to detect?
Among major models, Claude has the lowest detection rate at 96%, followed by open-source models like LLaMA and Mistral at 93-95%. Gemini is at 97% and GPT-4 at 98%. No currently available AI model produces text that reliably bypasses Turnitin on its own. The differences between models (2-5 percentage points) are too small to matter in practice.

Related Posts

Need to humanize AI text?

Paste your text and get it back clean in under 3 seconds. Free to try.

Try Anti-Turnitin Free
AC

Alex Chen

AI detection researcher and founder of Anti-Turnitin. Spent 3 years reverse-engineering how AI detectors classify text.