TL;DR
- Pangram ranked first overall with 93.1% accuracy, correctly classifying 149 out of 160 samples.
- GPTZero came second at 90.6%, only four correct classifications behind Pangram.
- Pangram correctly identified 87.8% of AI-written samples and 98.7% of human-written samples.
- Only 1 out of 78 human samples was wrongly flagged as AI by Pangram.
- QuillBot produced zero false accusations in this dataset, but it missed 32 of the 82 AI texts.
- Based on this test, Pangram is the best all-round detector of the six. However, the gap over GPTZero is small enough that this test alone cannot prove Pangram is universally better in every situation.
What Did We Test?
The comparison included Pangram, GPTZero, Originality.ai, QuillBot, Winston AI, and ZeroGPT. Each CSV contained 160 labeled samples. The real source of each sample was already known, so the detector's answer could be checked against the correct answer.
The dataset was nearly balanced: 82 samples were AI-written and 78 were human-written. That matters because a detector should not be able to look accurate simply by favoring one label most of the time.
For five of the six detector files, the same 160 texts matched after normalizing spacing and line breaks. The Originality.ai file matched 138 of Pangram's 160 texts after the same normalization, so its overall result is still useful, but its head-to-head comparison is slightly less controlled.
Also Read: [STUDY] How accurate is GPTZero? - An Independent Analysis!
Pangram Had the Highest Overall Accuracy
Pangram finished at the top with 93.1% accuracy. It got 149 of 160 classifications right. GPTZero followed at 90.6%, with 145 correct answers. The remaining detectors were noticeably lower: Originality.ai reached 83.8%, QuillBot 80.0%, Winston AI 79.4%, and ZeroGPT 73.8%.
In simple terms, Pangram made 11 mistakes in the test. GPTZero made 15. ZeroGPT made 42.
That makes Pangram the clear numerical winner in this dataset. But accuracy by itself can hide an important problem. A detector could achieve a decent score while being excellent at recognizing human writing but poor at catching AI, or the other way around.
Also Read: [HOT] Is ZeroGPT A Good AI Detector?
Can Pangram Catch AI Without Flagging Human Writing?
Pangram correctly detected 72 of 82 AI texts, giving it an AI detection rate of 87.8%. It also correctly recognized 77 of 78 human texts, or 98.7%.
This balance is one of the strongest arguments in Pangram's favor. QuillBot actually did slightly better on human writing because it did not falsely flag any of the 78 human samples. However, QuillBot detected only 50 of the 82 AI samples. In other words, it was very cautious about accusing human writers, but that caution came with a major cost: it missed almost four out of every ten AI texts.
GPTZero was much closer to Pangram. It also wrongly flagged only one human sample, while catching 68 of 82 AI samples. Pangram caught four more AI samples while keeping the same number of false accusations.
Also Read: How accurate is Winston AI? The Surprising Results
The Error Types Matter More Than a Single Score
For students, a false positive is especially important. This happens when a detector labels genuinely human writing as AI-generated. Pangram and GPTZero each produced only one false positive in the 78 human samples. Originality.ai produced two, Winston AI produced five, and ZeroGPT produced 16.
ZeroGPT's result is a useful warning against trusting a detector only because it gives a confident-looking percentage. In this dataset, more than one in five human texts were incorrectly labeled as AI by ZeroGPT.
A false negative is the opposite mistake: AI-generated text is labeled human. Pangram had 10 false negatives, the fewest of all six tools. GPTZero had 14, Originality.ai 24, ZeroGPT 26, Winston AI 28, and QuillBot 32.
Is Pangram Definitely Better Than GPTZero?
This is where the answer becomes more interesting. Pangram beat GPTZero by 2.5 percentage points, but 160 samples are not enough to treat that gap as permanent proof. On the matching texts, there were 14 cases where Pangram and GPTZero differed in whether they were correct. Pangram won nine of those disagreements, while GPTZero won five.
That favors Pangram, but not by a huge margin. The confidence ranges for the two detectors also overlap. A confidence range is simply an estimate of where the detector's true long-run accuracy might fall if we repeated similar tests many times.
So the data supports saying Pangram performed best in this experiment. It does not support saying Pangram will always beat GPTZero on every essay, every writing style, every language, or every future model.
What the Real Detector Screens Look Like
These interfaces can make the result look extremely certain. That is useful for readability, but students should remember that a detector's percentage is not the same thing as scientific proof that a particular person used AI.
So, Is Pangram the Best AI Detector?
Among the six detectors in this dataset, yes. Pangram had the highest overall accuracy, the highest AI detection rate, and one of the lowest false-positive rates. It offered the strongest balance between catching AI writing and protecting human writing from incorrect flags.
GPTZero was a very close second and deserves that distinction. QuillBot was the safest if the only goal was avoiding false accusations, but it missed too much AI-generated text to rank first overall. Originality.ai, Winston AI, and ZeroGPT all made substantially more mistakes in this test.
The most important lesson, though, is bigger than the ranking. An AI detector should be treated as evidence, not a verdict. Even the best performer here was wrong 11 times out of 160. If a teacher, university, or student treats a detector result as unquestionable proof, the detector does not need to fail often for the consequences to become serious.
Pangram was the best detector we tested. It was not perfect, and that difference matters.