Benchmark report · Detector v7 · October 2026
How accurate is ParaTrace?
Every figure on this page was measured on text the detector never saw in training, scored by the production service at one fixed setting, the Strict operating point.
99.6 %
of AI-written answers detected
26,082 answers · 90 models · 26 developers
0.8 %
of human texts wrongly flagged
22,691 held-out texts, same setting
96.0 %
still detected under evasion attacks
11 attack types · 16,500 texts
20 ms
to score a 300-word text
on our servers
I · Detection
It catches 99.6 % of AI-written answers
26,082 answers to real user prompts, written by 90 models from 26 developers, at a setting that wrongly flags fewer than 1 in 100 human texts.
| Developer | Detected | Answers |
|---|---|---|
| OpenAI | 99.4 % | 4,595 |
| 99.4 % | 3,782 | |
| Anthropic | 99.6 % | 3,485 |
| Meta | 99.8 % | 2,138 |
| Mistral | 99.7 % | 1,958 |
| Alibaba | 99.8 % | 1,834 |
| DeepSeek | 99.7 % | 1,393 |
| Reka | 99.9 % | 1,200 |
| Microsoft | 99.8 % | 1,072 |
| 01.AI | 99.8 % | 900 |
| Cohere | 99.4 % | 822 |
| xAI | 99.8 % | 579 |
| Zhipu | 100 % | 373 |
| Amazon | 100 % | 312 |
| NVIDIA | 99.7 % | 308 |
| Nexusflow | 100 % | 300 |
| MiniMax (small sample) | 100 % | 183 |
| Databricks (small sample) | 98.8 % | 168 |
| Moonshot (small sample) | 100 % | 132 |
| Tencent (small sample) | 100 % | 54 |
| Other developers | 99.2 % | 494 |
† Fewer than 200 answers, so indicative only. “Other developers” groups six smaller developers, a research fine-tune and one anonymous test model. Snapshots and reasoning modes of one model count once.
Late-2025 flagship models
- Anthropic98.7 % of 318
- Google99.3 % of 667
- xAI99.0 % of 102
- OpenAI99.2 % of 357
Each developer’s leading model as of late 2025. Answers were collected up to October 2025. Newer models are not yet benchmarked.
New models it never saw
98.6 %
of 1,172 answers from late-2025 model versions that were not in its training data.
Developers it never saw
99.8 %
of 3,948 answers from developers with no model at all in its training data.
Clear-cut verdicts
98 %
of AI-written answers score 99 % or higher, far from the cut-off, so most verdicts don’t hinge on a close call.
II · False alarms
Fewer than 1 in 100 human texts wrongly flagged
0.84 % of 22,691 held-out human texts were flagged as AI, and the longer the text, the rarer a false alarm.
- Under 50 words in the text, 2.8 % (24 of 846 texts)
- 50–99 words in the text, 1.6 % (30 of 1,911 texts)
- 100–199 words in the text, 1.1 % (75 of 6,825 texts)
- 200–399 words in the text, 0.6 % (54 of 9,021 texts)
- 400+ words in the text, 0.2 % (7 of 4,088 texts)
Under 50 words, 2.8 % are wrongly flagged, which is why the service warns on short texts. From 400 words it is 0.17 %.
| Writing | Flagged | Texts |
|---|---|---|
| Scientific papers | 0.0 % | 432 |
| Legal texts (small sample) | 0.0 % | 148 |
| Patents (small sample) | 0.0 % | 115 |
| Government reports (small sample) | 0.0 % | 61 |
| Medical notes (small sample) | 0.0 % | 17 |
| Résumés (small sample) | 0.0 % | 15 |
| Essays | 0.2 % | 5,562 |
| Questions & answers | 0.4 % | 3,369 |
| Reviews | 0.8 % | 1,602 |
| Forum posts | 0.9 % | 1,869 |
| Book excerpts | 1.1 % | 443 |
| Speeches and debates (small sample) | 1.2 % | 81 |
| Encyclopedia articles | 1.3 % | 1,718 |
| Scientific abstracts | 1.4 % | 1,816 |
| News articles | 1.4 % | 2,733 |
| Other web writing | 1.5 % | 2,037 |
| Blogs (small sample) | 1.6 % | 189 |
| Recipes | 2.0 % | 293 |
| Educational texts (small sample) | 2.6 % | 116 |
| Emails (small sample) | 2.7 % | 75 |
† Small sample, so indicative only.
These are held-out documents, never used in training, from the same kinds of sources the detector learned from. Writing unlike those sources can be flagged more often. See Where to be careful below.
III · Evasion
Common evasion tricks cost it little
Across 11 common attacks, 96.0 % of AI texts are still detected, against 97.3 % before any attack. For text from chat assistants it is 98.0 %.
- AI paraphrasing, 91.5 % of 1,500 texts, 95 % confidence interval 89.9–92.8 %
- Deleted articles (a, an, the), 93.6 % of 1,500 texts, 95 % confidence interval 92.2–94.7 %
- Random upper/lower case, 95.2 % of 1,500 texts, 95 % confidence interval 94.0–96.2 %
- Synonym swaps, 95.9 % of 1,500 texts, 95 % confidence interval 94.8–96.8 %
- Deliberate misspellings, 96.1 % of 1,500 texts, 95 % confidence interval 95.0–96.9 %
- British/American spellings, 97.1 % of 1,500 texts, 95 % confidence interval 96.1–97.8 %
- Extra paragraph breaks, 97.2 % of 1,500 texts, 95 % confidence interval 96.2–97.9 %
- Look-alike letters, 97.3 % of 1,500 texts, 95 % confidence interval 96.4–98.0 %
- Invisible characters, 97.3 % of 1,500 texts, 95 % confidence interval 96.4–98.0 %
- Extra spaces, 97.3 % of 1,500 texts, 95 % confidence interval 96.4–98.0 %
- Changed numbers, 97.5 % of 1,500 texts, 95 % confidence interval 96.6–98.2 %
1,500 texts per attack. Lines show 95 % confidence intervals, and the axis starts at 88 %.
The hardest attack
91.5 %
of AI texts rewritten by an AI paraphrasing model are still caught.
Character tricks
Look-alike letters, invisible characters and extra spaces are undone before scoring, so they have no effect on the verdict.
Independent documents
The attacked texts come from a public robustness benchmark. None of these documents were used in training.
IV · Beyond training
It holds up beyond its training data
Detection that only works on familiar material is brittle. These tests use genres, models and settings the detector was never shown.
A genre it never studied
90.6 %
of 5,000 AI-written poems detected, although poetry was absent from training. For poems from chat assistants it is 97.3 %.
Models held back on purpose
99.3 %
of 8,118 texts from two widely used model versions that were deliberately kept out of training.
Any sampling setting
94.1 %
for text generated with random sampling, the hardest setting. With a repetition penalty it is 99.4 % to 99.9 %.
V · Formats
Short or long, prose or code
From 50 words up, detection stays high whatever the length or format of the answer.
| Length | Detected | Answers |
|---|---|---|
| 50–99 words | 98.9 % | 3,076 |
| 100–199 words | 99.6 % | 5,193 |
| 200–399 words | 99.8 % | 8,628 |
| 400–799 words | 99.8 % | 6,766 |
| 800 words or more | 99.4 % | 2,419 |
| Content | Detected | Answers |
|---|---|---|
| Prose | 99.7 % | 20,998 |
| With code | 99.6 % | 2,210 |
| With math notation | 98.8 % | 1,564 |
| With tables | 99.9 % | 807 |
| With no formatting at all | 99.2 % | 6,252 |
It isn’t just spotting formatting. Answers with no headings, lists or bold text at all are detected 99.2 % of the time.
VI · Speed
Fast enough to check everything
Scoring takes milliseconds, even for long documents.
20 ms
to score a 300-word text
70 ms
to score a typical 1,500-word text
Time on our servers from text in to verdict out, text preparation included. Network and queue time are not included.
VII · Limits
Where to be careful
What we measured that falls short, and what we haven’t measured yet.
- Persuasive essays by adults, a kind of writing kept out of training entirely, were wrongly flagged 18.4 % of the time (32 of 174). Human writing unlike the training sources can be flagged far more often than the averages above.
- Human poetry was wrongly flagged 2.2 % of the time, across 1,733 poems.
- Short texts under 50 words were wrongly flagged 2.8 % of the time. The service warns on anything shorter.
- Synonym-swapping tools can mislead it. Human text run through one was flagged 16.4 % of the time.
- Humanizer tools, sold to make AI text pass as human, are its weakest spot. See how it compares with Pangram on them.
- AI-paraphrased human writing counts as AI. By design, 42.8 % of it was flagged, because the final wording came from a model.
- English only. We have not yet evaluated documents that mix human and AI writing, human text lightly edited with AI, or paraphrasing tools beyond those tested here.
- Newer models are not yet benchmarked. Answers were collected up to October 2025.
This is a statistical estimate, not proof. Don't use it as the sole basis for decisions about a person.
VIII · Method
How we measured
- One setting throughout.
- Every figure uses the Strict operating point, calibrated so that about 1 in 100 human validation texts is flagged. Longer texts are judged part by part, as when you submit them.
- AI-written answers.
- 26,082 responses to real user prompts from public chat logs, collected between 2024 and October 2025, plus a small set generated in-house. Prompts used for testing were excluded from training.
- Human writing.
- 22,691 held-out documents, never used in training, from 43 sources that include essays, questions and answers, news, forum posts, reviews, encyclopedia articles, scientific papers and abstracts, legal texts, patents and government reports.
- Evasion and new genres.
- Documents from a public robustness benchmark, none used in training. Each attack was applied to the same 1,500 AI texts.
- Scoring.
- Every text was scored by the production service running detector v7, text preparation included, exactly as when you submit it.
- Against other detectors.
- We set ParaTrace beside Pangram on the public tests Pangram reports on, in a separate comparison, and repeated the check NeurIPS ran on scientific papers.
- Rates.
- The detection rate is the share of AI-written texts flagged as AI. The false-alarm rate is the share of human-written texts flagged as AI. Ranges are 95 % confidence intervals (Wilson).