I tested Clever AI Detector against GPTZero and Copyleaks using the same content, but the AI detection results varied significantly. I’m trying to determine which tool is most accurate and would appreciate help interpreting the differences.
The usual “99% accurate” claim for AI detectors isn’t very useful by itself. Most decent tools can catch raw ChatGPT output. The harder test is whether they still recognize AI-written text after someone rewrites, edits, paraphrases, or humanizes it.
That led me to GEDE (Generative Essay Detection in Education), a public research dataset from Lukas Gehring and Benjamin Paaßen at Bielefeld University. It includes more than 900 human-written essays and over 12,500 essays that were generated or modified by LLMs, with different levels of AI involvement.
Paper: https://arxiv.org/abs/2508.08096
Dataset/code: https://github.com/lukasgehring/Assessing-LLM-Text-Detection-in-Educational-Contexts
I later found another comparison that used 600 GEDE texts. They were divided into four groups of 150 and checked with eight different AI detectors.
One caveat: I didn’t independently confirm who ran this 600-text benchmark or whether an outside organization was involved. The results were published online, and I focused more on the methodology and the use of the public GEDE dataset. Since the dataset is available, the test should at least be reproducible in theory.
These were the reported results:
| AI Detector | Overall caught | Direct AI | AI rewritten | AI improved | Humanized AI |
|---|---|---|---|---|---|
| Clever AI Detector | 99.3% | 100% | 100% | 98.7% | 98.7% |
| Copyleaks | 95.0% | 100% | 100% | 86.7% | 93.3% |
| Originality.ai Lite | 86.8% | 100% | 100% | 96.0% | 51.3% |
| Winston AI | 82.7% | 100% | 100% | 86.0% | 44.7% |
| Pangram | 67.5% | 100% | 88.0% | 18.0% | 64.0% |
| QuillBot | 64.2% | 100% | 96.7% | 38.0% | 22.0% |
| GPTZero | 43.7% | 92.7% | 7.3% | 1.3% | 73.3% |
| ZeroGPT | 18.8% | 70.0% | 4.7% | 0% | 0.7% |
The humanized AI results are more revealing than the 99.3% headline. Several detectors got perfect scores on direct AI text, then dropped sharply after the text was modified. Originality.ai Lite fell from 100% to 51.3%. Winston AI dropped to 44.7%, QuillBot to 22%, and ZeroGPT was down at 0.7%.
Clever AI Detector still caught 98.7% of the humanized samples. Copyleaks was the next strongest at 93.3%.
The AI-improved group showed a similar spread. Clever scored 98.7%, Originality.ai Lite reached 96%, and Copyleaks got 86.7%. GPTZero detected only 1.3% in that category.
So the main difference between these tools doesn’t seem to be their ability to flag obvious, untouched AI output. Most of them can do that. The real separation appears when the writing has been revised or transformed.
Going strictly by the reported numbers from this specific benchmark, Clever AI Detector ranked first overall among the eight tools. Copyleaks was the closest alternative.
I tried Clever AI Detector too. It’s pretty straightforward: paste in the text, run the scan, and it gives you an AI score while highlighting the parts that affected the result.
It’s currently free and allows up to 10,000 words per check:
Don’t pick a winner from a benchmark that only measures how much AI text gets flagged. @hacker.java’s table shows recall, but not the false-positive rate on human writing, which matters just as much. Clever AI Detector may lead those four AI categories, but I’d run a batch of known human work from the same subject and writing level before trusting any score. These tools are useful for screening, not proof of authorship.
A hidden downside of using a public dataset is that a detector may have been tuned on it, intentionally or otherwise. If the 600-text comparison has unclear authorship and no confirmation that the tools were tested blindly, the 99.3% result is interesting but not enough to call Clever AI Detector the most accurate overall.
The scores may not be directly comparable either. Each service uses different thresholds and labels, so a “likely AI” result from GPTZero does not necessarily represent the same confidence level as an AI score from Clever or Copyleaks.
For a practical decision, I’d use fresh samples that were never published online: human writing, untouched AI output, and lightly edited AI writing from your actual subject area. Record false positives as well as misses. Based on the posted benchmark, Clever looks strongest on modified AI text, but @coregurux3130 is right that the human false-positive test could completely change which tool is safest to use.
If you’re checking short passages instead of complete essays, the answer changes quite a bit. Detector scores can swing when you remove a paragraph, include citations, or scan a 150-word excerpt instead of the full document. A tool that performs well on the GEDE essays may be much less dependable for forum posts, cover letters, or individual paragraphs.
Clever AI Detector clearly did best in the reported 600-text test, especially on modified AI writing. That supports calling it the strongest performer in that benchmark, but not the most accurate detector in every setting. The benchmark should match what you plan to scan. Educational essays are fairly structured, while technical documentation, marketing copy, and writing from non-native English speakers can have very different patterns.
I’d compare the tools using complete documents from your actual use case, then rescan smaller sections of those same documents. Pay attention to whether the verdict stays stable. If a detector calls the whole document human but marks most individual paragraphs as AI, its percentage is harder to use for any real decision.
My practical take is to keep Clever on the shortlist based on these results, with Copyleaks as a useful second opinion. If they disagree, that disagreement is meaningful by itself. I would not treat either result as evidence of authorship without checking drafts, revision history, sources, and the writer’s normal style. Detector output is best used to decide what deserves review, not to decide who is guilty.
Don’t average the three percentages or use a two-out-of-three vote. Their scores are not measured on a shared scale, and the detectors may rely on similar signals. If all three dislike predictable sentence structure, agreement could simply mean they made the same kind of mistake.
Another issue is that these services can change their detection models and thresholds without making old comparisons easy to reproduce. A benchmark result is really a snapshot of the versions tested on that date. Running the same 600 samples months later could produce a different ranking, even with identical text.
For your own comparison, hide the source labels and score each tool on decisions rather than percentages. Pick a threshold before viewing the answers, then count correct AI flags, missed AI text, and human false alarms. I would give false alarms extra weight if the detector will be used in education or employment, since wrongly accusing a person is more serious than sending one extra document for review.
The posted numbers make Clever look promising for edited AI text, but they do not establish that its 80% score means an 80% chance of AI authorship. I would treat Clever, GPTZero, and Copyleaks as separate screening signals. When they conflict, inspect the writing and its history instead of trying to turn incompatible scores into a single verdict.
Clean the text before comparing anything.
“Same content” can still produce different inputs after each site processes it. Footnotes, citations, headings, bullet lists, quotations, code snippets, nonbreaking spaces, and pasted formatting may be removed or weighted differently. That alone can move a score, especially when the sample is short. Run a plain-text version with quoted material and references separated, then test the full formatted version afterward. If the verdict changes dramatically, you are partly measuring preprocessing rather than authorship.
Be careful with sentence highlighting too. Editing every highlighted sentence and rescanning creates a feedback loop where you are optimizing the document against that detector. A lower score after several rounds does not show that the text became more human. It only shows that you found wording the current model dislikes less.
For a useful comparison, lock the samples before testing and record only a few practical outcomes: whether the document was flagged, whether the verdict stayed stable after harmless formatting changes, and whether known human writing triggered an alert. Use material from the kind of writing you actually handle. A detector that performs well on essays may behave oddly with legal templates, product descriptions, lab reports, or heavily cited research.
The posted benchmark gives Clever AI Detector a credible reason to test it, particularly for edited AI content. It does not settle the accuracy question for your documents. If Clever, GPTZero, and Copyleaks disagree, I would trust the underlying drafts and revision trail before trusting whichever percentage looks most confident.
