Skip to content
ParaTrace

Unlimited free credit top-ups, for a limited time. Running low? Ask for more credits and we top you up, as often as you need. This offer ends soon, so catch it while it lasts.

Benchmark report · Detector v7 · October 2026

How accurate is ParaTrace?

Every figure on this page was measured on text the detector never saw in training, scored by the production service at one fixed setting, the Strict operating point.

  • 99.6 %

    of AI-written answers detected

    26,082 answers · 90 models · 26 developers

  • 0.8 %

    of human texts wrongly flagged

    22,691 held-out texts, same setting

  • 96.0 %

    still detected under evasion attacks

    11 attack types · 16,500 texts

  • 20 ms

    to score a 300-word text

    on our servers

I · Detection

It catches 99.6 % of AI-written answers

26,082 answers to real user prompts, written by 90 models from 26 developers, at a setting that wrongly flags fewer than 1 in 100 human texts.

Detection rate by developer
DeveloperDetectedAnswers
OpenAI99.4 %4,595
Google99.4 %3,782
Anthropic99.6 %3,485
Meta99.8 %2,138
Mistral99.7 %1,958
Alibaba99.8 %1,834
DeepSeek99.7 %1,393
Reka99.9 %1,200
Microsoft99.8 %1,072
01.AI99.8 %900
Cohere99.4 %822
xAI99.8 %579
Zhipu100 %373
Amazon100 %312
NVIDIA99.7 %308
Nexusflow100 %300
MiniMax (small sample)100 %183
Databricks (small sample)98.8 %168
Moonshot (small sample)100 %132
Tencent (small sample)100 %54
Other developers99.2 %494

† Fewer than 200 answers, so indicative only. “Other developers” groups six smaller developers, a research fine-tune and one anonymous test model. Snapshots and reasoning modes of one model count once.

Late-2025 flagship models

  • Anthropic98.7 % of 318
  • Google99.3 % of 667
  • xAI99.0 % of 102
  • OpenAI99.2 % of 357

Each developer’s leading model as of late 2025. Answers were collected up to October 2025. Newer models are not yet benchmarked.

New models it never saw

98.6 %

of 1,172 answers from late-2025 model versions that were not in its training data.

Developers it never saw

99.8 %

of 3,948 answers from developers with no model at all in its training data.

Clear-cut verdicts

98 %

of AI-written answers score 99 % or higher, far from the cut-off, so most verdicts don’t hinge on a close call.

II · False alarms

Fewer than 1 in 100 human texts wrongly flagged

0.84 % of 22,691 held-out human texts were flagged as AI, and the longer the text, the rarer a false alarm.

False alarms by length
  • Under 50 words in the text, 2.8 % (24 of 846 texts)
  • 50–99 words in the text, 1.6 % (30 of 1,911 texts)
  • 100–199 words in the text, 1.1 % (75 of 6,825 texts)
  • 200–399 words in the text, 0.6 % (54 of 9,021 texts)
  • 400+ words in the text, 0.2 % (7 of 4,088 texts)

Under 50 words, 2.8 % are wrongly flagged, which is why the service warns on short texts. From 400 words it is 0.17 %.

False alarms by kind of writing
WritingFlaggedTexts
Scientific papers0.0 %432
Legal texts (small sample)0.0 %148
Patents (small sample)0.0 %115
Government reports (small sample)0.0 %61
Medical notes (small sample)0.0 %17
Résumés (small sample)0.0 %15
Essays0.2 %5,562
Questions & answers0.4 %3,369
Reviews0.8 %1,602
Forum posts0.9 %1,869
Book excerpts1.1 %443
Speeches and debates (small sample)1.2 %81
Encyclopedia articles1.3 %1,718
Scientific abstracts1.4 %1,816
News articles1.4 %2,733
Other web writing1.5 %2,037
Blogs (small sample)1.6 %189
Recipes2.0 %293
Educational texts (small sample)2.6 %116
Emails (small sample)2.7 %75

† Small sample, so indicative only.

These are held-out documents, never used in training, from the same kinds of sources the detector learned from. Writing unlike those sources can be flagged more often. See Where to be careful below.

III · Evasion

Common evasion tricks cost it little

Across 11 common attacks, 96.0 % of AI texts are still detected, against 97.3 % before any attack. For text from chat assistants it is 98.0 %.

AI texts still detected, by attack
  • AI paraphrasing, 91.5 % of 1,500 texts, 95 % confidence interval 89.9–92.8 %
  • Deleted articles (a, an, the), 93.6 % of 1,500 texts, 95 % confidence interval 92.2–94.7 %
  • Random upper/lower case, 95.2 % of 1,500 texts, 95 % confidence interval 94.0–96.2 %
  • Synonym swaps, 95.9 % of 1,500 texts, 95 % confidence interval 94.8–96.8 %
  • Deliberate misspellings, 96.1 % of 1,500 texts, 95 % confidence interval 95.0–96.9 %
  • British/American spellings, 97.1 % of 1,500 texts, 95 % confidence interval 96.1–97.8 %
  • Extra paragraph breaks, 97.2 % of 1,500 texts, 95 % confidence interval 96.2–97.9 %
  • Look-alike letters, 97.3 % of 1,500 texts, 95 % confidence interval 96.4–98.0 %
  • Invisible characters, 97.3 % of 1,500 texts, 95 % confidence interval 96.4–98.0 %
  • Extra spaces, 97.3 % of 1,500 texts, 95 % confidence interval 96.4–98.0 %
  • Changed numbers, 97.5 % of 1,500 texts, 95 % confidence interval 96.6–98.2 %

1,500 texts per attack. Lines show 95 % confidence intervals, and the axis starts at 88 %.

The hardest attack

91.5 %

of AI texts rewritten by an AI paraphrasing model are still caught.

Character tricks

Look-alike letters, invisible characters and extra spaces are undone before scoring, so they have no effect on the verdict.

Independent documents

The attacked texts come from a public robustness benchmark. None of these documents were used in training.

IV · Beyond training

It holds up beyond its training data

Detection that only works on familiar material is brittle. These tests use genres, models and settings the detector was never shown.

A genre it never studied

90.6 %

of 5,000 AI-written poems detected, although poetry was absent from training. For poems from chat assistants it is 97.3 %.

Models held back on purpose

99.3 %

of 8,118 texts from two widely used model versions that were deliberately kept out of training.

Any sampling setting

94.1 %

for text generated with random sampling, the hardest setting. With a repetition penalty it is 99.4 % to 99.9 %.

V · Formats

Short or long, prose or code

From 50 words up, detection stays high whatever the length or format of the answer.

AI answers detected, by length
LengthDetectedAnswers
50–99 words98.9 %3,076
100–199 words99.6 %5,193
200–399 words99.8 %8,628
400–799 words99.8 %6,766
800 words or more99.4 %2,419
AI answers detected, by content
ContentDetectedAnswers
Prose99.7 %20,998
With code99.6 %2,210
With math notation98.8 %1,564
With tables99.9 %807
With no formatting at all99.2 %6,252

It isn’t just spotting formatting. Answers with no headings, lists or bold text at all are detected 99.2 % of the time.

VI · Speed

Fast enough to check everything

Scoring takes milliseconds, even for long documents.

  • 20 ms

    to score a 300-word text

  • 70 ms

    to score a typical 1,500-word text

Time on our servers from text in to verdict out, text preparation included. Network and queue time are not included.

VII · Limits

Where to be careful

What we measured that falls short, and what we haven’t measured yet.

  • Persuasive essays by adults, a kind of writing kept out of training entirely, were wrongly flagged 18.4 % of the time (32 of 174). Human writing unlike the training sources can be flagged far more often than the averages above.
  • Human poetry was wrongly flagged 2.2 % of the time, across 1,733 poems.
  • Short texts under 50 words were wrongly flagged 2.8 % of the time. The service warns on anything shorter.
  • Synonym-swapping tools can mislead it. Human text run through one was flagged 16.4 % of the time.
  • Humanizer tools, sold to make AI text pass as human, are its weakest spot. See how it compares with Pangram on them.
  • AI-paraphrased human writing counts as AI. By design, 42.8 % of it was flagged, because the final wording came from a model.
  • English only. We have not yet evaluated documents that mix human and AI writing, human text lightly edited with AI, or paraphrasing tools beyond those tested here.
  • Newer models are not yet benchmarked. Answers were collected up to October 2025.

This is a statistical estimate, not proof. Don't use it as the sole basis for decisions about a person.

VIII · Method

How we measured

One setting throughout.
Every figure uses the Strict operating point, calibrated so that about 1 in 100 human validation texts is flagged. Longer texts are judged part by part, as when you submit them.
AI-written answers.
26,082 responses to real user prompts from public chat logs, collected between 2024 and October 2025, plus a small set generated in-house. Prompts used for testing were excluded from training.
Human writing.
22,691 held-out documents, never used in training, from 43 sources that include essays, questions and answers, news, forum posts, reviews, encyclopedia articles, scientific papers and abstracts, legal texts, patents and government reports.
Evasion and new genres.
Documents from a public robustness benchmark, none used in training. Each attack was applied to the same 1,500 AI texts.
Scoring.
Every text was scored by the production service running detector v7, text preparation included, exactly as when you submit it.
Against other detectors.
We set ParaTrace beside Pangram on the public tests Pangram reports on, in a separate comparison, and repeated the check NeurIPS ran on scientific papers.
Rates.
The detection rate is the share of AI-written texts flagged as AI. The false-alarm rate is the share of human-written texts flagged as AI. Ranges are 95 % confidence intervals (Wilson).