AI detectors are everywhere now, especially around schools and universities. If you are a student, you may have already pasted an essay into a detector just to see whether it thinks your writing looks “AI-generated.” The problem is that two detectors can look at the same text and give very different answers.
To see how large that difference can be, we tested ZeroGPT and Pangram on the same set of 160 samples: 82 were AI-written and 78 were human-written. Both datasets recorded an AI score from 0 to 100, where a higher score means the detector believes the text is more likely to be AI-written. In the recorded results, a score of 50 or above always matched an “AI” verdict, so we used the tools' recorded verdicts when calculating accuracy.
The short answer
In this test, Pangram was substantially more accurate. It classified 149 of 160 samples correctly (93.1%), while ZeroGPT classified 118 correctly (73.8%).
The difference becomes more important when you look at human-written work. Pangram correctly recognized 77 of 78 human samples (98.7%). ZeroGPT correctly recognized 62 of 78 (79.5%). In other words, ZeroGPT incorrectly flagged 16 genuinely human samples as AI, compared with only one for Pangram.
Also Read: Stealthwriter vs ZeroGPT: I Tested 100 Rewrites, and the Results Were Complicated!
Why false positives matter so much for students
A false positive is a human-written piece that a detector wrongly labels as AI. For students, this is probably the most worrying type of error. You can write something yourself and still receive a high AI score.
In our test, Pangram produced 1 false positive. ZeroGPT produced 16. ZeroGPT also missed more AI-written text: it labeled 26 AI samples as human, compared with 10 for Pangram. That second mistake is called a false negative.
Looking at the paired results gives another useful picture. Both tools were correct on 113 samples. Pangram was correct while ZeroGPT was wrong on 36 samples. ZeroGPT was correct while Pangram was wrong on only 5. Both tools were wrong on 6 samples.
Also Read: [STUDY] ZeroGPT vs Winston AI: Which AI Detector Performs Better? 160-Sample Test
The scores themselves behave very differently
Another striking difference was how the tools used the 0–100 scale. Pangram was extremely decisive: 158 of its 160 scores were either exactly 0 or exactly 100. Only two results landed somewhere in between.
ZeroGPT was much more gradual. It returned 85 intermediate scores between 0 and 100. That can make the score look more nuanced, but a more detailed-looking number does not automatically mean a more accurate result. In this dataset, ZeroGPT's extra score variation came with more classification mistakes.
Also Read: [STUDY] Turnitin vs. ZeroGPT: How do they compare?
The average AI score also showed stronger separation for Pangram. Real AI samples averaged 87.8 with Pangram, while human samples averaged just 1.4. With ZeroGPT, AI samples averaged 72.3, but human samples still averaged 30.3. A detector is easier to interpret when its scores for human and AI writing stay farther apart.
What the interfaces are like
Pangram's interface is more report-like. The screenshots show an overview panel, a clear Human Written or AI Generated verdict, a percentage gauge, AI highlighting, and separate Details and Evidence areas. That makes it easier to inspect why a piece received a particular result instead of seeing only a percentage.
ZeroGPT is simpler. You paste text into a large box, run the detector, and receive a percentage with a short interpretation such as “likely human written” or “mixed signals.” For quick checks this is straightforward, although the test results show that the percentage should not be treated as a precise measurement of how much of a document was written by AI.
Also Read: Which is Better GPTZero or ZeroGPT? An In-Depth Comparison!
So, which one performed better?
On these 160 samples, the numbers clearly favored Pangram. Its overall accuracy was about 19 percentage points higher, and the biggest practical difference for students was its much lower false-positive rate. ZeroGPT gave more finely graded percentages, but it also had much more overlap between scores for real human writing and AI writing.
That does not mean any AI detector can prove who wrote a piece of text. This was one dataset, detector models can be updated, and performance can change with writing style, text length, editing, subject area, or newer AI models. A detector result is best treated as one signal, not a final verdict.