What Is Token-Type Ratio (TTR)?

Definition

The ratio of unique words (types) to total words (tokens) in a text. A lower TTR indicates more word repetition. AI-generated text often has a distinctive TTR pattern — moderately diverse vocabulary used with unnaturally consistent distribution across a document.

Token-Type Ratio (TTR) Explained

Token-Type Ratio (TTR) is a lexical diversity metric calculated by dividing the number of unique words in a text by the total number of words. A TTR of 1.0 would mean every word is unique, while a lower ratio indicates more repetition. In AI detection, TTR serves as a supplementary signal. AI models tend to produce text with a moderate, consistent TTR across a document — neither too repetitive nor too varied. Human writing shows more variation in TTR across different sections: an introduction might have high lexical diversity while a technical section reuses key terms heavily. This section-to-section TTR variation is a detectable signal. Standard TTR has a known limitation: it decreases naturally as text length increases because longer texts inevitably repeat common words. To address this, researchers use variants like Moving Average TTR (MATTR) or the Measure of Textual Lexical Diversity (MTLD), which control for text length. AI detection systems that incorporate lexical diversity typically use these length-corrected variants. TTR alone is not a strong enough signal for detection, but it contributes to ensemble classifiers as one of many features. Its value lies in being relatively easy to compute and difficult to consciously manipulate while writing.

How Token-Type Ratio (TTR) Relates to AI Detection

TTR is a secondary feature in most ensemble classifiers. It's not strong enough to drive detection on its own, but contributes to the overall confidence score. Originality.ai and Turnitin both analyze lexical diversity as part of their multi-feature models. The uniformity of TTR across document sections is often more telling than the absolute value.

How Anti-Turnitin Handles Token-Type Ratio (TTR)

Anti-Turnitin's post-processing adjusts lexical diversity patterns to match human norms. The algorithm introduces natural variation in vocabulary density across paragraphs — denser terminology in some sections, more varied expression in others — matching the uneven TTR distribution characteristic of human writing.

Related Terms

Frequently Asked Questions

What is a normal TTR for human writing?

TTR varies by text length and type. For a 1,000-word academic essay, human TTR typically ranges from 0.4 to 0.6, with significant section-to-section variation. AI text tends to fall in a similar range but with notably less variation across sections.

Can I improve my TTR to avoid detection?

Manually adjusting TTR is difficult because it requires changing vocabulary patterns throughout an entire document while maintaining readability. More importantly, TTR is just one of many signals in ensemble classifiers. Improving TTR alone will not bypass detection if other signals like perplexity and burstiness remain AI-like.

See how Anti-Turnitin handles these signals

Try Anti-Turnitin Free

No credit card required. Paste your text and get results in under 3 seconds.