Originality.ai and Pangram are both popular AI writing detectors, so it is reasonable to wonder whether they give roughly the same answer. In this dataset, the answer is: they are similar in some important ways, but they are not interchangeable.
Both detectors were very good at leaving human-written text alone. The bigger difference was on AI-written text. Pangram caught substantially more of it, while Originality.ai missed a much larger share.
TLDR
- Originality.ai accuracy: 83.8%
- Pangram accuracy: 93.1%
- Originality.ai false positive rate: 2.6%
- Pangram false positive rate: 1.3%
- Originality.ai AI recall: 70.7%
- Pangram AI recall: 87.8%
- On the 138 exact texts shared by both datasets, the two detectors agreed on 114 texts (82.6%).
- When the two detectors disagreed, Pangram was correct 19 times while Originality.ai was correct 5 times.
For students and writers, the most reassuring result is that false positives were uncommon in this test. For AI detection itself, Pangram was clearly stronger.
How the comparison was done
Each CSV contains 160 texts. Both datasets have the same class balance: 82 AI-written texts and 78 human-written texts. The ground-truth column is labeled Written By, while each detector has its own verdict column.
There is one important detail. The two CSVs are not perfectly identical sets. Only 138 texts appear in both files exactly. That means the overall accuracy figures below are calculated using each detector's full 160-row dataset, while direct detector-to-detector agreement is calculated only on the 138 shared texts.
This distinction matters. Comparing accuracy can be done across the two full files because the class totals are the same, but asking whether the detectors made the same decision on the same piece of writing requires the shared subset.
Also Read: Is Pangram Better Than GPTZero?
Overall accuracy: Pangram was ahead
| Metric | Originality.ai | Pangram |
|---|---|---|
| Correct results | 134 / 160 | 149 / 160 |
| Accuracy | 83.75% | 93.12% |
| AI recall | 70.73% | 87.80% |
| Human recall | 97.44% | 98.72% |
| False positive rate | 2.56% | 1.28% |
| False negative rate | 29.27% | 12.20% |
| Precision when calling text AI | 96.67% | 98.63% |
Pangram correctly classified 149 of 160 texts, compared with 134 of 160 for Originality.ai. That is a gap of about 9.4 percentage points.
The reason for that gap becomes obvious when the AI-written samples are separated from the human-written samples.
Also Read: [STUDY] Is Pangram AI Detector Similar to Turnitin?
The biggest difference was missed AI text
Originality.ai correctly detected 58 of 82 AI-written texts. It missed the other 24, producing a false negative rate of 29.3%.
Pangram detected 72 of 82 AI-written texts and missed 10. Its false negative rate was therefore only 12.2%.
This is the main reason the two tools should not be treated as equivalent. Originality.ai was much more likely to let an AI-written sample pass as human in this dataset.
Pangram was not simply labeling everything as AI either. Its performance on the human-written samples remained very strong, which is especially important for students and professional writers.
What about false positives on human writing?
A false positive happens when a detector labels genuinely human-written work as AI. For a student, this is usually the most worrying type of mistake.
Originality.ai incorrectly flagged 2 of 78 human-written samples, giving it a false positive rate of 2.56%. Pangram falsely flagged 1 of 78, for a false positive rate of 1.28%.
So both were conservative with human writing in this test. Pangram was slightly better, but with only 78 human samples, the difference between one false positive and two should not be exaggerated.
The more useful takeaway is that both detectors produced far fewer false positives than false negatives. Their main weakness in this dataset was failing to catch some AI writing, not accusing large amounts of human writing.
Also Read: Is Originality AI Similar to Turnitin? A Data-Based Comparison
How often did Originality.ai and Pangram actually agree?
There were 138 exact texts that could be matched between the two CSVs. The detectors gave the same verdict on 114 of them, an agreement rate of 82.6%.
That sounds fairly similar, but it also means they disagreed on 24 texts, or roughly one out of every six shared samples.
The disagreement cases are particularly revealing. Pangram was right while Originality.ai was wrong on 19 shared texts. Originality.ai was right while Pangram was wrong on only 5.
Across all 138 shared texts, both were correct on 108 and both were wrong on 6.
Looking only at the shared AI-written texts
The shared subset contains 82 AI-written texts. Both detectors correctly caught 54 of them. Pangram caught another 18 that Originality.ai missed, while Originality.ai caught only 4 that Pangram missed. Both failed on 6.
This is strong evidence that Pangram's advantage in this dataset came from better AI recall, not from becoming recklessly aggressive with human writing.
Also Read: [100 Samples Test] Can BypassGPT Really Bypass Originality.ai?
Examples from the tests
Pangram correctly labeled the Android customization sample as AI-generated:
It also classified the social engineering sample as AI-generated:
But Pangram missed the AI-written classic books sample and labeled it human-written. This is one of the false negatives in the dataset:
Originality.ai also shows both sides of the problem. It confidently identified some AI-written samples:
And it also returned strongly human results on other writing:
Another AI result from Originality.ai shows how strongly its interface can express confidence:
Be careful when comparing their scores
The score columns in these CSVs run in opposite directions. The Originality.ai column is explicitly labeled so that 100% means human-written. Pangram's score behaves in the opposite direction, with high values corresponding to AI verdicts.
So a raw score of 100 does not mean the same thing in the two files. The verdict columns are the safer way to compare the detectors directly.
Pangram's scores in this dataset were also unusually polarized. Almost every result was either 0 or 100, with only two intermediate values. Originality.ai had more intermediate scores, although many of its results were also close to one extreme. That makes the two scoring systems look less similar than their final Human/AI labels might suggest.
So, is Originality.ai similar to Pangram?
Yes in broad behavior, no in actual performance.
Both tools were highly accurate on human-written samples and both had low false positive rates. They also agreed on about 82.6% of the texts that could be matched directly.
But an 82.6% agreement rate is not high enough to assume they are substitutes for each other. On about one in six shared texts, they gave different answers. More importantly, Pangram detected AI writing much more reliably in this test, with 87.8% AI recall versus 70.7% for Originality.ai.
If your main concern is avoiding a false accusation on human writing, both performed well here. If your goal is catching AI-written content, Pangram was clearly the stronger detector in this dataset.
For students and writers, the practical lesson is not to treat any single detector score as proof of authorship. These tools can disagree with each other, and even a confident result can still be wrong. A detector result is better treated as one signal rather than a final judgment.
Final verdict
Originality.ai and Pangram are similar enough that they often reach the same conclusion, especially on clearly human-written text. But their behavior diverges much more on AI-written samples. Based on these 160-sample tests and the 138 texts that can be matched directly, Pangram was more accurate overall, missed far fewer AI texts, and maintained a similarly low false positive rate.









