Just a few years after generative AI entered the mainstream, the struggle to identify machine-produced text remains one of the most vexing problems in publishing and academia. In a new round of tests, 11 standalone AI content detectors were given five writing samples: two created by a human author and three produced by ChatGPT. The task was straightforward: declare whether each sample was written by a person or an AI. The results were anything but consistent.
Only three dedicated detectors identified every sample correctly. Another popular service, which had earned a perfect score in an earlier round, collapsed to a 20 percent accuracy rate. Meanwhile, several general-purpose chatbots outperformed most of the tools built specifically for AI detection. The findings suggest that content detectors should be treated with caution, and that a chatbot might be as useful as almost any specialized product.
Why AI-generated plagiarism matters
The definition of plagiarism is central to this debate. Plagiarism means taking the words or ideas of another person and passing them off as one’s own without crediting the source. Using an AI writing tool such as ChatGPT does not involve theft in the traditional sense, but when the output is presented without attribution, it still fits that definition. Students, journalists, and marketers increasingly rely on AI-generated text, and that has created demand for tools that can flag it.
Yet detection tools are not neutral referees. They are statistical models that try to estimate the likelihood a text came from a large language model. Because they rely on probability, they make mistakes. The latest tests show that those mistakes are not marginal. Human-written columns were flagged as AI, some AI-generated pieces were judged human with high confidence, and one detector even refused to make a call.
How the test was constructed
The evaluation used five blocks of text. Two were drafted by the author; three were generated by ChatGPT. Each test block was submitted separately to every detection service, and the service response was compared with the known origin of the text. A result above 70 percent confidence was treated as the tool’s final answer, whether that answer indicated human or AI authorship. That made scoring simple: each correct identification earned one point, with a maximum of five points per detector.
The full group of standalone services included BrandWell, Copyleaks, GPT-2 Output Detector, GPTZero, Grammarly, Originality.ai, Pangram, QuillBot, Undetectable.ai, Writer.com, and ZeroGPT. One previously tested service, Monica, was removed because its free tier limited tests to 250 words and then required a costly upgrade. Pangram joined the test for the first time and immediately delivered a perfect result.
What the latest results show
Across 55 individual tests, accuracy ranged from 100 percent to only 20 percent. The three perfect performers were Pangram, QuillBot, and ZeroGPT. Pangram, built by former Google and Tesla engineers, also earned a perfect score in its very first appearance in the series. QuillBot, which had been wildly inconsistent in earlier tests, delivered a perfect score for the second consecutive round. ZeroGPT improved from an earlier score of 80 percent to a perfect result and has maintained that level in this latest run.
At the other end of the scale, Undetectable.ai scored 20 percent. It identified human-written text as 60 percent likely to be AI, then judged three ChatGPT-created samples to be 75 percent, 76 percent, and 77 percent likely human. Writer.com also struggled, assigning human origin to every sample despite three of the five having been written by ChatGPT. BrandWell and Grammarly each managed only 40 percent accuracy.
Several established names landed in the middle. Copyleaks scored 80 percent, but it misclassified the first human-written sample as 100 percent AI authored. GPTZero also scored 80 percent, yet it was unable to render a judgment on one human sample and identified another human sample as AI. Originality.ai earned 80 percent but made a striking error: it declared a human-written essay to be 100 percent AI generated.
Do chatbots work better
Because the test results were uneven, the evaluation also considered whether everyday AI chatbots could serve as AI detectors. The same five text blocks were submitted to ChatGPT free tier, ChatGPT Plus, Microsoft Copilot, Google Gemini, and Grok, with a simple prompt asking whether each block was written by a human or an AI.
ChatGPT Plus, Copilot, and Gemini all returned perfect scores. The free version of ChatGPT missed one human-written sample but correctly identified the other four. Grok, despite its reputation for strong reasoning, missed three of five samples and concluded that every block was human-written.
There was one especially surprising result. When the first human-written passage was submitted to the free version of ChatGPT in a private and unsigned session, the chatbot not only classified it as human text but also identified the author by name. The model drew on patterns in online articles that the author had written, demonstrating that context can sometimes help a model reach a conclusion that a specialized detector cannot.
Why accuracy is so volatile
AI detection scores change for many reasons. Models update, services add restrictions, and language models improve. In this round, some companies that had performed well in the summer declined after introducing limits on free usage. Others changed their scoring thresholds. The result is a moving target for schools, media organizations, and anyone who manages a website.
False positives are especially dangerous. A writer who chooses words carefully, or a non-native speaker who follows formal grammar rules, can be wrongly accused of using artificial intelligence. The tests suggest this risk is very real. One detector claimed human writing was fully AI-generated; another did the opposite and cleared all three AI samples as human.
What this means for editors and teachers
No single tool is safe enough to be the only line of defense. The original human samples sometimes looked too polished to a detector, while the ChatGPT samples were not polished enough to trip other detectors. Editors should treat a detection report as a starting point for discussion rather than as proof. Teachers should ask students to explain their work and preserve early drafts. Publishers should combine software checks with human review.
It is also worth remembering that the most common AI detectors are not academically certified instruments. Many are commercial marketing products. Their accuracy claims often overstate real-world performance, and independent checks continue to show that their scores can be influenced by content length, formatting, and the specific language model used to create the text.
Winners in a crowded market
Pangram stands out as a useful new option. Its five free scans per day fit the needs of occasional users, and it delivered correct verdicts for all samples. QuillBot, after a rocky start in earlier evaluations, now looks far more reliable. ZeroGPT has matured from a bare-bones website into a proper service while also improving its scoring. For people who need a second opinion on a piece of writing, these three are the strongest standalone candidates in the current field.
Among chatbots, Copilot, Gemini, and ChatGPT Plus all earned perfect results in this test. Since many people already pay for or have access to these assistants, they offer a no-extra-cost method of checking text for AI origins. The free version of ChatGPT is a reasonable fallback, but it is not perfect.
Because AI writing tools continue to evolve, these rankings are not permanent. Services that score perfectly today may degrade tomorrow after an update, just as several did during this testing cycle. New entrants, such as Pangram, can quickly rise to the top. The most reliable approach is to compare several tools, understand their limitations, and never rely on a probability number alone. As the tests continue, the most important lesson remains: both dedicated detectors and general chatbots can help, but they should be used as assistants to human judgment, not replacements for it.
Source: ZDNET News