← All articles

Is ZeroGPT Accurate? We Tested It on 30 Short Emails (2026)

Is ZeroGPT accurate on short email? On this corpus it errs cautious rather than trigger-happy: ZeroGPT scored every one of the 15 emails in our non-AI set at 0% AI, and it missed 2 of the 15 AI-written emails at the default ≥50% line. It was also the only detector in the test whose two sets did not overlap at all — 0% on all 15 non-AI emails, 26.2% or higher on all 15 AI ones — and it did that on 64–86-word emails, every one of them under the 150-word minimum its own FAQ asks for. Limato ran all 30 emails through five detectors — ZeroGPT, GPTZero, QuillBot, Grammarly and Limato's own AI Detector — on 2026-08-15, at two thresholds, with confidence intervals. The result that should worry you is not any single tool's error rate: it is that the tools disagree with each other so hard that 4 of the 30 emails were scored 0% by one detector and 100% by another.

This page also carries a correction. An earlier version of this study, run the day before, found that detectors flagged 63–100% of non-native business emails as AI. That finding was wrong, it has been withdrawn, and the section explaining why is below rather than buried. Every number on this page comes from the replacement run.

We keep measuring this because Limato ships a free AI Detector of its own and it faces the identical risk: scoring a non-native writer's own words as machine-generated. On the new data it fails that test worse than any tool we compared it against at the ≥50% line this post applies to everyone — and better than that count suggests at the ≥70 boundary where the product actually prints the word AI. Both numbers, and the one that is worse for us than either, are in that section.

Method

30 short business emails, five detectors, one run per detector per item, run date 2026-08-15. 15 items are AI-written; 15 are not.

Lengths, in words:

Set n Min Median Max
non-AI set, all 1515647486
— of those, hand5647278
— of those, model-draft-hand-edited106474.586
AI-written15737679

Why short matters, and where that puts us relative to each tool's own floor. A real business email is short. Several detectors state a minimum input length that this entire corpus is under — which is the point of the study rather than a defect in it, but it has to be on the record before any result is read.

Detector Stated minimum input length Items below it (of 15 non-AI) Items below it (of all 30)
ZeroGPT150 words (“At least 150–200 words”, own FAQ)15 of 1530 of 30
GPTZero250 characters (~50 words), own benchmark pages0 of 150 of 30
QuillBot80 words (“A minimum of 80 words is required to scan”, own product page)12 of 1527 of 30
Grammarlynone stated for the free web tool
Limato AI Detectornone documented

A stated minimum is not always an enforced one: in a length probe outside this sample, both GPTZero and QuillBot returned a real score for a 5-word, 21-character input.

Two thresholds. Every score is evaluated at the common default of ≥50% AI and at a stricter ≥20% AI cutoff. The strict cut is not a straw man — it is how a suspicious reader actually behaves, treating any non-zero flag as a mark against you. Reporting both shows how much of a detector's reputation rests on threshold choice alone.

One run per text per detector. No retries, no best-of-N. Detector scores are known to be non-deterministic; a single run understates that variance, and that is listed as a limitation below.

One wrinkle in that run date. The 15 non-AI items were scored on 2026-08-15. The 15 AI items were scored on 2026-08-14 and carried over unchanged when the human half of the corpus was replaced — they were never re-run, apart from one Grammarly cell re-collected on 2026-08-15. results.csv stamps the rebuild date, 2026-08-15, on all 120 of its rows; the raw *.jsonl collector logs carry the true per-call dates, and anyone checking the numbers should use those.

The result we had to throw away

The first version of this study used a different human set: 30 non-native business email drafts from esl-email-2026-08/corpus.json, the same corpus behind our earlier rewrite measurement. Run on 2026-08-14, it produced a striking result.

Detector Flagged (≥50%) Rate 95% CI Flagged (≥20%)
ZeroGPT19/3063.3%[45.5%, 78.1%]23/30
GPTZero30/30100.0%[88.6%, 100.0%]30/30
QuillBot23/3076.7%[59.1%, 88.2%]26/30
Grammarly24/3080.0%[62.7%, 90.5%]24/30
Limato0/300.0%[0.0%, 11.4%]19/30

Those numbers mean nothing, and we are publishing them anyway so the record is complete.

The 30 drafts were model-generated. They were built to imitate non-native business email, but a model wrote every word of them and no person was involved in producing them at any point — not writing, not editing. So when ZeroGPT flagged 19 of them, ZeroGPT was correct. When GPTZero flagged all 30, GPTZero was correct. A false-positive rate computed over machine-written text is not a false-positive rate; it is a detection rate with the label flipped. The number that looked like an indictment of the detectors was, read properly, the detectors doing their job.

The consequence is worth stating plainly, because it cuts against a claim this site has repeated and that circulates widely: "AI detectors punish non-native writers" did not reproduce once the non-AI half of the corpus was written by a person rather than by a model. On the replacement set the same detectors flag 0 to 2 of the 15 non-AI items at ≥50%. If there is a penalty on non-native English at this length, this study did not find it — on 15 emails from one author with one first language, which is thin enough that "did not find it" is the strongest form the claim can take.

Two caveats on the withdrawn table itself. Those 30 items are no longer in results.csv — the numbers above were recovered from the raw per-call collector logs, which still hold one line per item per detector, all dated 2026-08-14. For 9 of the 150 item-by-detector cells (8 ZeroGPT, 1 GPTZero) only a repeat-run reading survives; the repeat reading was used rather than dropping the item. The recovered ZeroGPT, QuillBot and Grammarly counts — 19/30, 23/30 and 24/30 — reproduce the figures the first version of the study reported exactly, which is the main reason to trust the reconstruction.

And the comparison is not clean. The old set and the new one differ in length (63–126 words, median 69.5, versus 64–86, median 74), in how they were constructed, and in whether a person was involved in producing them at all. This is a documented before-and-after of our own methodology. It is not a controlled A/B on identical text, and nobody should read the swing as an experiment about corpus construction.

False positives on the replacement set

A false positive here means a detector scored an item above the threshold when it was not AI-written. All 15 non-AI items count, at both thresholds.

Detector Flagged (≥50%) FP rate 95% CI Flagged (≥20%) FP rate 95% CI
ZeroGPT0/150.0%[0.0%, 20.4%]0/150.0%[0.0%, 20.4%]
GPTZero2/1513.3%[3.7%, 37.9%]5/1533.3%[15.2%, 58.3%]
QuillBot0/150.0%[0.0%, 20.4%]0/150.0%[0.0%, 20.4%]
Grammarly2/1513.3%[3.7%, 37.9%]2/1513.3%[3.7%, 37.9%]
Limato4/1526.7%[10.9%, 52.0%]10/1566.7%[41.7%, 84.8%]

Read the confidence intervals before the point estimates. At n=15 a 0/15 result still admits a true rate as high as 20.4%, and 2/15 spans 3.7% to 37.9%. Nothing in this table separates the third-party tools from each other at any useful resolution. What it does support is a negative: none of them produced the mass flagging the withdrawn corpus appeared to show.

Does it matter how the email was produced?

Five of the 15 non-AI emails were typed in English from scratch; the other ten started from a model draft the author then edited. Splitting the same measurement along that line is the obvious robustness check, so here it is.

Threshold ≥50%

Detector hand (n=5) 95% CI model-draft-hand-edited (n=10) 95% CI
ZeroGPT0/5 — 0.0%[0.0%, 43.4%]0/10 — 0.0%[0.0%, 27.8%]
GPTZero0/5 — 0.0%[0.0%, 43.4%]2/10 — 20.0%[5.7%, 51.0%]
QuillBot0/5 — 0.0%[0.0%, 43.4%]0/10 — 0.0%[0.0%, 27.8%]
Grammarly1/5 — 20.0%[3.6%, 62.4%]1/10 — 10.0%[1.8%, 40.4%]
Limato1/5 — 20.0%[3.6%, 62.4%]3/10 — 30.0%[10.8%, 60.3%]

Threshold ≥20%

Detector hand (n=5) 95% CI model-draft-hand-edited (n=10) 95% CI
ZeroGPT0/5 — 0.0%[0.0%, 43.4%]0/10 — 0.0%[0.0%, 27.8%]
GPTZero1/5 — 20.0%[3.6%, 62.4%]4/10 — 40.0%[16.8%, 68.7%]
QuillBot0/5 — 0.0%[0.0%, 43.4%]0/10 — 0.0%[0.0%, 27.8%]
Grammarly1/5 — 20.0%[3.6%, 62.4%]1/10 — 10.0%[1.8%, 40.4%]
Limato2/5 — 40.0%[11.8%, 76.9%]8/10 — 80.0%[49.0%, 94.3%]

On ZeroGPT and QuillBot the two groups are identical: 0 flags in either group, at either threshold. Grammarly flags exactly one item in each group at both thresholds — hw-01, typed from scratch, and hw-12, from a model draft. The two groups separate only on GPTZero and Limato, and they separate the same way: both flag the model-drafted-then-edited items more often (GPTZero 2/10 against 0/5 at ≥50%, Limato 8/10 against 2/5 at ≥20%). With 5 and 10 items the intervals overlap almost completely, so that is a direction, not a difference this data can size.

False negatives on AI-written email

The 15 AI items were model-written and not edited. A detector scoring one below the threshold is a false negative — the opposite failure, and the one that matters if your goal is catching machine text rather than not accusing a person.

Detector Missed (<50%) FN rate 95% CI Missed (<20%) FN rate 95% CI
ZeroGPT2/1513.3%[3.7%, 37.9%]0/150.0%[0.0%, 20.4%]
GPTZero0/150.0%[0.0%, 20.4%]0/150.0%[0.0%, 20.4%]
QuillBot2/1513.3%[3.7%, 37.9%]1/156.7%[1.2%, 29.8%]
Grammarly1/156.7%[1.2%, 29.8%]1/156.7%[1.2%, 29.8%]
Limato0/150.0%[0.0%, 20.4%]0/150.0%[0.0%, 20.4%]

Unedited model output at email length is easy to catch: worst case here is 2 of 15 missed. The interesting failure at this length is not missing AI text. It is what happens on everything that is not obviously AI text.

ZeroGPT specifically

ZeroGPT is the tool most people mean when they search "is ZeroGPT accurate", so it gets its own read beyond the tables.

It flagged nothing on the non-AI set. 0 of 15 at ≥50%, 0 of 15 at ≥20%. And it did not merely stay under the line: it returned exactly 0% on every single one of the 15 items. Not 3%, not 11% — zero, fifteen times.

It missed 2 of 15 AI-written emails at ≥50% (13.3%, 95% CI 3.7–37.9%), the joint-worst false-negative rate in the set alongside QuillBot. At the strict ≥20% cut it missed none.

And all 15 non-AI items are below its own stated floor. ZeroGPT's own FAQ answers “What text length works best?” with “At least 150–200 words”; the longest item in the non-AI set is 86. This is the part worth thinking about. A tool operating below its documented minimum could behave in either direction — it could guess wildly, or it could refuse to commit. ZeroGPT does neither, and the result is the cleanest number in the study: exactly 0% on all 15 non-AI items, and 26.2% or higher on all 15 AI ones. The two sets do not overlap at all. It is the only detector here that separates them — GPTZero, QuillBot, Grammarly and our own each score at least one non-AI item above their own lowest AI item.

Read carefully what is being separated, though. The line it draws is between unedited generic business prose and short email in a non-native register — and it draws it the same way whether the email was typed from scratch or edited out of a model draft, returning a flat 0% on all 15 either way. So a 0% from ZeroGPT is not a statement about who wrote the words. It is ZeroGPT saying the text does not read like unedited model output at this length.

The detectors do not agree with each other

Every one of the 30 items has a score from all five detectors, which makes disagreement directly measurable. It is large.

At the ≥50% line, where the score collapses to a yes/no, pairwise agreement across all 30 items runs from 80.0% to 93.3%:

PairAgreeRate
ZeroGPT vs QuillBot28/3093.3%
ZeroGPT vs Grammarly27/3090.0%
QuillBot vs Grammarly27/3090.0%
ZeroGPT vs GPTZero26/3086.7%
GPTZero vs QuillBot26/3086.7%
GPTZero vs Limato26/3086.7%
GPTZero vs Grammarly25/3083.3%
Grammarly vs Limato25/3083.3%
ZeroGPT vs Limato24/3080.0%
QuillBot vs Limato24/3080.0%

80% agreement between two tools sounds high until you notice that a corpus which is half AI-written by construction makes agreement cheap: both tools flagging the same obvious machine text costs them nothing. The disagreement concentrates exactly where a person's reputation would be at stake.

Three worked examples

Unedited detector output, one run each, each item labelled with how it was produced. Scores are each tool's own 0–100% AI likelihood.

Example 1 — hw-01, typed in English from scratch, 78 words

Text (hw-01 — typed from scratch)

"Good evening Maria. Text you about next delivery, which should be arrive on the warehouse in Valencia at 12 of this month. Logistic company informs that in the documentation absent actual invoice, without it the cargo does not release it from the terminal. Could you please provide the fixed document on this address during current day? If you need any additional information from our side, I'm ready to provide it immediately. Thank you for cooperation. I appreciate it."

DetectorScoreVerdict at ≥50%
ZeroGPT0%Not flagged (correct)
GPTZero8.8%Not flagged (correct)
QuillBot19.2%Not flagged (correct)
Grammarly100%AI (false positive)
Limato52%AI (false positive)

This is the clearest false positive in the study. The email reads as a non-native speaker's own English — "should be arrive", "in the documentation absent actual invoice", "during current day" are the author's first-language syntax wearing English words — and it was typed out from scratch by the person whose first language is doing the interference. Three detectors read it correctly. Grammarly returns 100%, its maximum. Our own detector returns 52% — over this post's ≥50% line, though our own interface labels that score mixed rather than AI.

Example 2 — hw-12, edited from a model draft, 72 words

Text (hw-12 — model draft, author-edited)

"Good afternoom Klara. During the check of documents we found out that in the invoice for july the incorrect sum is indicated: instead of agreed two thousand one hundred euro in the document stands two thousand six hundred. The mistake happened by our fault during the transfer of data from the system. The corrected invoice I attach to this email. I'm sorry for caused inconveniences and for the necessity of repeated approval."

DetectorScoreVerdict at ≥50%
ZeroGPT0%Not flagged
GPTZero42.2%Not flagged
QuillBot0%Not flagged
Grammarly100%AI (false positive)
Limato28%Not flagged

Look at the row as a whole: 0, 42.2, 0, 100, 28 for one email. This is one of the 4 items scored 0 by one detector and 100 by another. Two tools see nothing, one sees a coin flip, one is certain. Whatever these numbers are, they are not five estimates of the same quantity. The typo in the first word ("afternoom") survived the hand edit; whether any of the five keyed on it is not something one run per text can answer.

Example 3 — hw-07, edited from a model draft, 83 words

Text (hw-07 — model draft, author-edited)

"Dear Duran, write you about the invoice number 2214 from 18 june for the sum three thousand four hundred euro. According to conditions of our contract the payment should be come during 30 days, but for today the money are not arrived on our account yet. Maybe the payment sent but not reflected in the system. Could you please check the status and inform us about result. If any difficulties with the payment appear, we are ready to discuss the schedule of repayment."

DetectorScoreVerdict at ≥50%
ZeroGPT0%Not flagged
GPTZero100%AI (false positive)
QuillBot0%Not flagged
Grammarly0%Not flagged
Limato78%AI (false positive)

Another 0-to-100 split, running the other way from Example 2: GPTZero is maximally certain this is machine-written while ZeroGPT, QuillBot and Grammarly all return a flat 0%. GPTZero's 100% and Limato's 78% are two of the false positives counted in the tables above. Five tools, one email, and no majority in either direction: at least three of the five are wrong here, and one run per text cannot tell you which three.

Our own detector, reported against itself

Limato's AI Detector was the worst false-positive offender in this study. On the 15-item non-AI set it flagged 4 at ≥50% (26.7%, 95% CI 10.9–52.0%) — twice the rate of GPTZero and Grammarly, against 0 of 15 for ZeroGPT and QuillBot. At ≥20% it flagged 10 of 15 (66.7%, 95% CI 41.7–84.8%), while ZeroGPT and QuillBot still flagged none. One of the four it flagged at ≥50% is hw-01 in Example 1 above.

The mirror image is also true, and we are not publishing one without the other. On the withdrawn 30-item model-generated corpus, Limato's detector was the only one of the five that flagged nothing at all: 0 of 30 at ≥50%, where the four third-party tools flagged between 19 and 30. A day earlier that was written up as evidence our detector avoided a failure mode the others had. It was nothing of the kind. On text a model actually wrote, our detector missed everything the others caught. At ≥20% on that same corpus it flagged 19 of 30 (63.3%) — so its apparent perfect record lived entirely inside one threshold convention.

Put the two together and the shape is consistent: our detector is the most willing to flag on the current corpus and was the least willing on the old one, which mostly says its scores sit in a narrow middle band that different thresholds cut in different places. It did get the AI side right on the current data — 0 of 15 missed at both thresholds — but so did GPTZero, and every tool here catches unedited model output easily. That is the cheap half of the problem.

One number this page owes you, because it is the one we ship. The ≥50% and ≥20% cuts used throughout this post are analyst conventions, applied identically to all five tools so the columns compare to each other. They are not what our product prints. Limato's detector returns a raw 0–100 score and a verdict, and the verdict boundary is fixed in our API: at 70 and above it says AI, from 30 to 69 mixed, below 30 human. Two of the four non-AI emails counted against us at ≥50% scored 52 and 65 — our own interface calls those mixed. Counted by the label a user actually sees, rather than by a threshold we do not use:

Verdict Limato printed non-AI set (n=15) AI set (n=15)
AI (score ≥70)115
mixed (30–69)60
human (<30)80

That is a better result and we are not going to pretend it is not. It is also not the exoneration it looks like, for two reasons. First, we cannot extend the same courtesy to the other four: we do not have their shipped verdict boundaries in a form we could apply consistently, so every cross-tool number on this page stays at ≥50% and ours is quoted there too. Comparing our shipped boundary against their re-thresholded scores would be the exact trick this post exists to document. Second, and worse for us: no threshold separates our two sets at all. Our highest-scoring non-AI email scored 78; our lowest-scoring AI email scored 72. The classes overlap, so 1-of-15 is a cut sitting inside that overlap rather than a clean boundary — move it up to 79 and we would clear the false positive at the cost of missing 11 of the 15 AI emails. ZeroGPT, on this corpus, has no overlap to cut into. That is the comparison that matters, and we lose it.

If you use Limato's AI Detector, carry both numbers: on this data it flagged roughly a quarter of non-AI short emails at the ≥50% threshold this post applies to everyone, and 1 of 15 at the ≥70 boundary where it actually prints the word AI. Every caveat we would apply to a competitor's score applies to ours.

What this data does not show

Reproduce it

Every file behind this page is public, in limato-app/limato-public under data/detector-study-2026-08/: sample.json (the 30 item ids with word counts, the detector list and both thresholds), corpus-hw.json (the 15 non-AI items verbatim, each carrying a provenance label recording how it was produced), results.csv (one row per third-party detector per item — 120 rows), limato.jsonl (our own detector's raw score and the verdict it printed for each item — the verdict-boundary table above is a straight count of that verdict field, not a re-thresholding of the scores), and score.js, which recomputes the false-positive, false-negative and coverage tables from those files and carries its own Wilson confidence-interval implementation plus a --selftest flag. Download the four data files into one directory and node score.js reprints every table on this page offline, no dependencies and no network. Only the 15 non-AI items are stored there verbatim; the 15 AI items live in esl-email-2026-08/corpus.json in the same repository, published with the earlier study, and are referenced by id rather than copied.

Two things to know before you run it. score.js does not split on provenance: it prints the n=15 rows only, and those match this page's headline table exactly — the hand and model-draft-hand-edited columns in the robustness check were computed over those two subsets with the same Wilson formula. And limato.jsonl still carries the 30 superseded esl-* rows, deliberately, because this page cites them; they are not in sample.json, so the script prints a run of WARNING: row for unknown item_id lines before the first table rather than scoring rows it was not asked about. The rest of the withdrawn corpus's scores are recoverable the same way, from the esl-* lines in the per-detector zerogpt.jsonl, gptzero.jsonl, quillbot.jsonl and grammarly.jsonl logs in that directory. The collectors that produced those logs are not published; the logs are their verbatim output, and every number on this page is derived from them by the script above.

Reuse. The data files in data/detector-study-2026-08/ are released under CC BY 4.0. Republish the numbers, the tables or the corpus anywhere, including commercially, as long as you credit Limato and link back to this page.

Related measurements from the same corpus family: what a rewrite model actually cuts from non-native English email — measured on the same 30-item corpus withdrawn above — and our comparison with QuillBot as a writing tool rather than a detector. Limato also ships a Humanizer for restructuring AI-sounding drafts; it is a separate tool and was not part of this measurement.

Frequently asked questions

Is ZeroGPT accurate?

On short email it errs cautious rather than trigger-happy. Limato ran 30 emails of 64–86 words through five detectors on 2026-08-15. ZeroGPT scored all 15 items in the non-AI set at 0% AI — 0 of 15 flagged at its default ≥50% threshold, 95% CI 0.0–20.4% — and missed 2 of the 15 AI-written emails at that same line (13.3%, 95% CI 3.7–37.9%). Note that all 15 non-AI items fall below ZeroGPT's own stated 150-word minimum input length. It was the only detector of the five whose two sets did not overlap: 0% on every non-AI item, 26.2% or higher on every AI one.

Does ZeroGPT flag non-native English writers?

Not in this sample. ZeroGPT flagged 0 of the 15 items in our non-AI set — all 15 of them written by a non-native English speaker — and returned exactly 0% on every one of them. The sample is one author with one first language, so this cannot be generalised to other language backgrounds.

Do AI detectors punish non-native writers?

That claim did not reproduce here once the non-AI half of the corpus was written by a person rather than by a model. An earlier version of this study measured false-positive rates of 63.3% (ZeroGPT), 100.0% (GPTZero), 76.7% (QuillBot) and 80.0% (Grammarly) at ≥50% — but its 30 "non-native" drafts were model-generated end to end, so flagging them was correct behaviour, not a false positive. On the replacement set ZeroGPT and QuillBot flag 0 of 15 and GPTZero and Grammarly flag 2 of 15 (13.3%) at the same threshold. Fifteen emails from one author with one first language is far too thin to generalise from.

What ZeroGPT score is considered AI?

ZeroGPT's interface treats 50% and above as AI-generated by default. We scored every item at that default and at a stricter ≥20% cutoff. For ZeroGPT the choice changed nothing on this corpus: it returned exactly 0% on all 15 non-AI items and 0 of 15 were flagged at either line.

Is ZeroGPT or GPTZero more accurate on short email?

They fail in opposite directions. On the 15 non-AI items at ≥50%, ZeroGPT flagged 0 and GPTZero flagged 2 (13.3%, 95% CI 3.7–37.9%). On the 15 AI-written items, ZeroGPT missed 2 (13.3%) and GPTZero missed 0 (95% CI 0.0–20.4%). With n=15 per side the confidence intervals overlap heavily, so treat this as directional.

Can free AI detectors be trusted on short text like emails?

They disagree with each other too much to be treated as a verdict. Across the 30 items, 9 (30.0%) had at least two detectors scoring 50 percentage points or more apart, and 4 (13.3%) were scored 0% by one detector and 100% by another. Two tools also state minimum input lengths this corpus is under: all 15 non-AI items are below ZeroGPT's 150 words, and 12 of 15 are below QuillBot's 80.

How does Limato's own AI Detector score on this data?

It was the worst false-positive offender in the set. On the 15 non-AI items Limato's AI Detector flagged 4 at ≥50% (26.7%, 95% CI 10.9–52.0%) and 10 at ≥20% (66.7%, 95% CI 41.7–84.8%), against 0 of 15 for ZeroGPT and QuillBot. It missed 0 of the 15 AI-written items at both thresholds, and it was the only detector that flagged 0 of 30 items in the withdrawn model-generated corpus at ≥50%. Those cuts are analyst conventions applied to all five tools alike; at the ≥70 boundary where Limato's own product prints the verdict "AI", it labelled 1 of the 15 non-AI emails as AI and caught all 15 AI-written ones. That is not a clean result either: its two sets overlap (non-AI max 78, AI min 72), so no threshold separates them the way ZeroGPT's did.

Write the email in your own words, faster

Limato rewrites, translates and fixes tone directly in the browser — for people whose English is their second language.

Add to Chrome →