Stylometry studies measurable properties of writing. It predates generative AI by decades and has been used for authorship research, literary analysis and forensic linguistics. In AI detection, it is useful as context—not as a magic fingerprint.
Human prose often alternates between short and long constructions, but the amount of variation depends heavily on genre. AI Doc Scan reports sentence-length variation rather than assuming that “uniform = AI”. Technical standards, legal prose and tightly edited documentation can also be highly regular.
Paragraph lengths can reveal templated structure. Repeated blocks of similar size may be useful evidence when combined with other signals, but are common in manuals, reports and web writing.
The scanner reports several complementary views because simple type-token ratio shrinks as documents become longer. MTLD estimates how long a text can continue before vocabulary diversity falls below a threshold. Yule’s K describes vocabulary concentration. Hapax share measures the proportion of vocabulary appearing only once, while normalized entropy describes how evenly words are distributed.
An n-gram is a sequence of n tokens. Repeated 2-grams are common in normal language; repeated 4- and 5-grams can be more informative about formulaic local structure. The scanner reports a repetition profile across multiple phrase lengths.
Repeated two-word openings can reveal structural habits. Adjacent-sentence lexical overlap estimates how much content vocabulary is reused from one sentence to the next, which can expose restatement and repetitive explanation.
Small function words—articles, prepositions, conjunctions and similar high-frequency items—are often useful in traditional stylometry because writers use them unconsciously. Punctuation entropy describes how varied the punctuation profile is. Both are descriptive and highly genre-dependent.
Phrases such as “it is important to note” or repeated transition structures are sometimes associated with generic LLM prose, but humans use them too. AI Doc Scan therefore treats formulaic-language density as a weak supporting signal rather than a standalone detector.
Stylometry can describe a text and sometimes support classification. It does not tell you who legally authored, owned, reviewed or approved the final work.