Stylometry approach to the statistical study of writing style and LLM

September 24, 2026 - Reading time: 8 minutes

Large language models can produce polished, convincing prose in seconds. But figuring out where that prose came from is considerably harder.

One possible approach is stylometry — the statistical study of writing style. Rather than focusing on what a text says, stylometry looks at how it is written. Researchers can measure recurring patterns in vocabulary, grammar, punctuation, sentence structure, and other features, then compare those patterns across authors, documents, or AI systems.

Stylometry is not new. Researchers have used it to investigate disputed authorship for decades, long before today's language models existed. Now, similar techniques are being applied to a newer question: can statistical patterns reveal whether a piece of writing came from a human or a machine?

What Does Stylometry Actually Measure?

Every writer develops habits, whether consciously or not. Some people favor short, direct sentences. Others rely on longer constructions, particular punctuation marks, or recurring grammatical patterns. Stylometry attempts to turn those habits into measurable data.

Researchers can examine features such as:

  • Sentence length across different sections of a document.

  • Grammatical structures and how frequently particular constructions appear.

  • Punctuation habits, including the use of commas, semicolons, dashes, and other marks.

  • Character sequences and recurring combinations of letters or symbols.

  • Vocabulary distribution and the way words are spread throughout a sample.

Modern computational systems can combine hundreds of these measurements. A machine-learning classifier can then look for patterns associated with known examples of human or machine-generated writing.

The important point is that stylometry does not need to understand a text in the same way a human reader does. It can instead search for statistical fingerprints that might otherwise be difficult to notice.

Can Stylometry Detect LLM-Generated Writing?

Recent research suggests that it can — at least under some controlled conditions.

A study published in Expert Systems with Applications in January 2026 examined short samples produced by humans and several language models. The researchers combined lexical, grammatical, syntactic, and punctuation-based features to classify the texts.

The experiments used ten-sentence samples derived from Wikipedia material. Across different testing conditions, the researchers reported binary classification accuracy ranging from 79 percent to 100 percent.

In one experiment comparing Wikipedia text with GPT-4 output, accuracy reached as high as 98 percent on a balanced dataset.

Those numbers are noteworthy, but they should not be interpreted as proof that AI-generated text can always be identified. Laboratory conditions are considerably cleaner than real-world writing. A document might be edited by a person, translated between languages, paraphrased, or assembled from several sources before anyone attempts to analyze it.

Where Does a ChatGPT Detector Fit In?

A ChatGPT detector generally attempts to estimate whether a piece of writing resembles text produced by a language model.

Stylometry can be one component of that process. A detection system may examine measurable writing characteristics and combine them with statistical models, probability estimates, or machine-learning classifiers.

But there is an important distinction between detecting a pattern and proving authorship.

No individual writing characteristic demonstrates that a machine produced a text. Structured sentences, predictable grammar, or repeated punctuation can occur naturally in human writing as well.

There is another complication: language models themselves change. A detection method trained around one generation of models may behave differently when confronted with newer systems. Effective detection therefore requires continual testing against new models and new types of generated text.

Stylometry Is Also Changing How Authorship Is Studied

The human-versus-machine question is only one part of the story. Stylometry has traditionally been concerned with a broader problem: which author is most consistent with a particular text?

A 2025 study published in PLOS ONE explored authorship attribution using language models trained on known samples from individual writers. The researchers then examined unknown documents and measured which author-specific model found each text most predictable.

The researchers reported results that matched or exceeded leading approaches across several standard datasets. Their analysis also found that content-related terms carried substantial information about authorship.

That finding is significant because it challenges a long-standing assumption in some stylometric approaches: that grammatical function words will always provide the strongest authorship signals. Depending on the dataset and method, the words that convey meaning may contain valuable information too.

Why AI Detection Still Has Limits

Even when a detector performs well in a benchmark, its result is best treated as evidence to examine, rather than a final verdict.

Controlled experiments can produce impressive accuracy because researchers know the origin of the samples. Real documents are messier. Human writers may revise AI-generated passages, combine material from different sources, or substantially alter the original text.

Length matters too. A short paragraph contains fewer stylistic signals than a longer sample, making statistical conclusions more difficult. Heavy editing can further alter the punctuation, vocabulary, and sentence structures that a detector relies on.

Language creates another challenge. Stylometric patterns do not necessarily transfer cleanly from one language to another. Research published in PLOS ONE during 2025, for example, identified measurable stylistic differences between human writing and outputs from seven language models in Japanese. This suggests that language-specific analysis can provide useful signals rather than assuming that methods developed for English will work equally well everywhere.

Where Is Stylometry Most Useful?

Stylometry becomes more informative when researchers have enough comparable material to work with. A single short paragraph provides relatively few features; several longer samples can reveal much richer patterns.

Potential applications include:

  • Academic research into patterns found in machine-generated writing.

  • Authorship attribution when comparing texts against known candidate writers.

  • Forensic analysis of disputed or anonymous documents.

  • Model comparison to investigate stylistic differences between language-model families.

  • Paraphrasing research examining how rewriting changes detectable stylistic signals.

Ultimately, stylometry offers something more useful than a simple human-or-AI switch. It provides a statistical framework for examining recurring patterns across texts.

The strongest approach is therefore unlikely to depend on a single telltale feature. By combining many measurements and considering the quality and quantity of the available evidence, researchers can treat authorship detection as a probabilistic analysis rather than a simplistic yes-or-no decision.

That distinction matters. A statistical signal can be informative without being conclusive — and understanding the difference is essential when stylometry is used to investigate real-world writing.

by T.Smalzar (the jWork for curated articles)



Category: