I need a reliable AI content detector that can identify AI-generated writing after it has been heavily edited by a person. The tools I’ve tested give conflicting results, so I’m looking for accurate recommendations based on real experience.
The usual “99% accurate” claim for AI detectors isn’t very useful when the test only involves untouched ChatGPT output. Raw AI text is the easy case. I was more interested in what happens after someone rewrites, edits, paraphrases, or humanizes it.
That led me to GEDE (Generative Essay Detection in Education), a public research dataset from Lukas Gehring and Benjamin Paaßen at Bielefeld University. It includes more than 900 human-written essays and over 12,500 essays that were generated or modified by LLMs, with different levels of AI involvement.
Paper: https://arxiv.org/abs/2508.08096
Dataset/code: https://github.com/lukasgehring/Assessing-LLM-Text-Detection-in-Educational-Contexts
I later found a comparison that tested eight AI detectors against 600 texts from GEDE. The sample was divided into four groups of 150: direct AI, AI-rewritten, AI-improved, and humanized AI.
One caveat here: I wasn’t able to independently verify who ran this specific 600-text benchmark or whether an outside organization was involved. The results were published online, and the underlying GEDE dataset is public, so the test should at least be reproducible. Still, I’d treat this as one benchmark rather than a final verdict on every detector.
These were the reported detection rates:
| AI detector | Overall caught | Direct AI | AI rewritten | AI improved | Humanized AI |
|---|---|---|---|---|---|
| Clever AI Detector | 99.3% | 100% | 100% | 98.7% | 98.7% |
| Copyleaks | 95.0% | 100% | 100% | 86.7% | 93.3% |
| Originality.ai Lite | 86.8% | 100% | 100% | 96.0% | 51.3% |
| Winston AI | 82.7% | 100% | 100% | 86.0% | 44.7% |
| Pangram | 67.5% | 100% | 88.0% | 18.0% | 64.0% |
| QuillBot | 64.2% | 100% | 96.7% | 38.0% | 22.0% |
| GPTZero | 43.7% | 92.7% | 7.3% | 1.3% | 73.3% |
| ZeroGPT | 18.8% | 70.0% | 4.7% | 0% | 0.7% |
The overall score isn’t the part that stood out to me. The humanized AI results are much more revealing.
Several detectors scored 100% on direct AI and then dropped sharply after the writing was humanized. Originality.ai Lite fell to 51.3%, Winston AI to 44.7%, QuillBot to 22%, and ZeroGPT to 0.7%. Clever AI Detector stayed at 98.7%, while Copyleaks reached 93.3%.
The AI-improved group showed a similar gap. Clever scored 98.7%, Originality.ai Lite got 96%, and Copyleaks reached 86.7%. GPTZero detected only 1.3% of that category.
So the practical takeaway is that most decent detectors can flag obvious, untouched AI writing. The real differences start showing up after the text has been edited or transformed.
Going strictly by the numbers from this specific benchmark, Clever AI Detector ranked first overall among the eight tools. Copyleaks was the closest alternative.
I tried Clever AI Detector too. The interface is pretty straightforward: paste in your text, run the check, and it gives you an AI score while highlighting the sections that affected the result.
It’s currently free and allows up to 10,000 words per check:
Don’t trust a detector that only reports how much AI text it catches without showing its false-positive rate on edited human writing. That benchmark makes Clever AI Detector look strong, but no score should be treated as proof of authorship, especially after heavy editing.
Heavy editing can remove most of the patterns a detector is looking for, so the more “human” the final text becomes, the less meaningful the label gets. At some point the detector is no longer identifying who wrote it. It is estimating whether the prose resembles material in its training data.
That makes “best detector” very dependent on what you check. Essays, marketing copy, technical documentation, and writing by non-native English speakers can behave quite differently. Text length matters too. A result from 1,500 words deserves more attention than a result from two polished paragraphs, although neither proves authorship.
Based on the benchmark @kevin86 posted, Clever AI Detector is reasonable to put on a shortlist, especially because it apparently held up better on transformed text. Copyleaks would be the obvious comparison. I still wouldn’t choose either from that table alone. The missing test is whether they falsely flag the kind of genuine human writing you actually deal with.
A practical evaluation would be:
- Collect 20 to 30 known human samples from the same setting and similar writers.
- Add known AI samples that have been edited to the degree you care about.
- Test complete documents rather than isolated sentences.
- Record false positives as carefully as successful detections.
- Repeat the checks later, since detector models and thresholds can change.
If a tool catches 95% of edited AI but flags 15% of your human samples, that is probably unacceptable for disciplinary, hiring, or publishing decisions. If you are only screening a large pile of low-risk submissions for closer review, the same tool may still be useful.
For anything consequential, I’d use the detector as a triage signal and then look for stronger evidence: document revision history, sudden changes in style, missing drafts, unverifiable citations, or whether the writer can explain and revise the work. Heavily edited AI writing simply does not leave a dependable forensic fingerprint. Clever may be the strongest candidate mentioned here, but no percentage output should be treated as a verdict.
No detector can reliably identify authorship once a person has substantially rewritten the text. Clever AI Detector may be worth testing, but the better check is whether it correctly handles your known human samples, especially polished or non-native writing.
A hidden problem with that table is that “caught” reduces very different scoring systems to a yes-or-no result. One detector may call anything above 50% AI, while another may require much stronger confidence. That can make two similar models look far apart simply because their default thresholds differ.
For a fair comparison, I’d want to see each detector measured at the same acceptable false-positive rate. For example, set every tool so it falsely flags no more than 2% of known human documents, then compare how much edited AI it catches. Without that calibration, the detector with the highest “caught” number may just be the most aggressive one.
Mixed authorship creates another issue. If a person keeps the AI outline, rewrites most sentences, adds original examples, and changes the argument, what is the correct label? Calling the entire document AI-generated is questionable even if a detector finds traces of the original draft. A percentage that looks precise can hide the fact that the category itself is fuzzy.
Clever AI Detector and Copyleaks seem like the two worth comparing from the posted results, but I would ignore their final labels at first. Save the raw scores for the same set of documents and check whether the scores separate your known-human samples from your edited-AI samples. Pay attention to overlap. If genuine human work regularly scores 60% and edited AI regularly scores 65%, the detector is not useful for your case, regardless of its benchmark rank.
I’d favor the tool that gives stable scores, exposes which passages affected the result, and has a sensible uncertain range instead of forcing every document into “human” or “AI.” A detector willing to say “inconclusive” is more useful than one that confidently labels everything.
Before comparing detectors, make a clean copy of the document and remove quotations, references, assignment instructions, standard disclaimers, and other pasted boilerplate. I initially assumed you should scan the file exactly as submitted, but that can muddy the result. A bibliography or repeated template language may influence the score even though it says nothing about who wrote the main text.
For the actual prose, Clever AI Detector and Copyleaks seem like sensible first checks based on the benchmark posted. I would run the full document, then run a few substantial sections separately. If the overall score is high only because one formulaic section gets flagged, that is very different from consistent flags throughout the argument. Passage highlighting seems more useful here than a single percentage.
What confused me most about edited AI writing is that there may no longer be a clean “AI or human” answer. Someone could replace half the sentences, keep the structure, add personal examples, and correct the facts. A detector might recognize patterns from the remaining text, but it cannot tell you how much genuine work the person contributed.
So I would pick whichever of those two tools gives the clearest passage-level feedback and treats uncertain text as uncertain. If they disagree, I would record that as inconclusive rather than choosing the result I prefer. For heavily edited material, conflicting scores are probably information in themselves: the text is near the point where detector labels stop being dependable.
