We believe it’s important that institutions can rely on Pangram’s high accuracy, therefore we encourage third-party verification on our quality metrics (false-positives & false-negatives). Below, we’ll highlight evaluations of Pangram from researchers at University of Chicago (UChicago) and University of Maryland (UMD) and commercial reviewers.
Key Takeaway: Pangram’s internal testing holds up to scrutiny from third-parties.
In a peer reviewed paper published in June 2026, a group of researchers at the Vrije Universiteit Brussel tested four AI detection tools – Pangram, GPTZero, Turnitin, and Copyleaks – on 160 academic papers over 4000 words each. The dataset was written in English, and evenly split between human-written by ESL students, AI-generated, hybrid (mixed human and AI text), and humanised AI text.
The researchers found that Pangram was the only tool to reliably detect AI content: "from the four AI detection tools studied here, at this moment only [Pangram] produced satisfactory results."
Pangram was the only tool to reliably detect both fully AI-written and humanised papers (97.5% and 95%). On the fully AI-generated set the other three caught essentially none, all scoring 0%. None of them handled humanised text reliably either, though Turnitin managed about half and Copyleaks a quarter.
| Tool | Detection rate (fully AI) | Detection rate (hybrid) | Detection rate (humanised) | False positive rate |
|---|---|---|---|---|
| Pangram | 97.5% | 95% | 95% | 0% |
| Turnitin | 0% | 60% | 52.5% | 0% |
| Copyleaks | 0% | 32.5% | 25% | 0% |
| GPTZero | 0% | 0% | 2.5% | small positive bias (only tool not at zero on human text) |
Earlier research found other detectors produced false positives for ESL writers, but this independent study indicates Pangram is not substantially biased against ESL text.
At UChicago’s Becker Friedman Institute for Economics, researchers compared four AI detectors: Pangram, GPTZero, Originality AI, and RoBERTa (an open-source AI detector). The study used each detector to analyze 1,992 human text(s) written pre-2020 and 1,992 AI-generated texts across different genres and word counts. They looked at two types of errors in AI detection: False Positive Rates and False Negative Rates. These rates were compared for multiple thresholds. The detectors also classified AI generated text from popular LLMs like ChatGPT, Claude and Gemini. Researchers created multiple FPR Policy Caps among detectors to note the changes in FNR.
From the study, Artificial Writing and Automated Detection by Brian Jabarian and Alex Imas on August, 2025:
Pangram dominates the other detectors across all thresholds.
Pangram is the only detector that meets a stringent policy cap (FPR ≤ 0.005) without compromising the ability to accurately detect AI text.
Pangram remains the low-cost leader in all genres and on average: $0.0228 per correctly flagged AI passage versus $0.0416 for OriginalityAI and $0.0575 for GPTZero, making Pangram the most cost-efficient detector for both full-length passages and stubs.
The study showed that:
Pangram achieves essentially zero false positive rates and false negative rates on medium-length to long passages.
Pangram’s high accuracy was acclaimed across different genres of text such as: blogs, reviews, resumes, news and novels. In shorter-form text, the false positive and false negative rates slightly increase “but remain well below reasonable policy thresholds".
Average false positive rate (flagging human text as AI) and false negative rate (missing AI text) for each detector, across the six genres and four LLMs. Lower is better on both.
| Tool | False positive rate | False negative rate | Robust to StealthGPT humanizer? |
|---|---|---|---|
| Pangram | 0.001 | 0.01 | Yes (near-100% detection) |
| Originality.ai | 0.002 | 0.035 | Partly (up to 0.21 FNR on short text) |
| GPTZero | 0.007 | 0.06 | No (FNR 0.50+) |
| RoBERTa (open-source) | 0.50 | unreliable (flags most text as AI) | n/a (unsuitable) |
Pangram no longer predicts for text under 50 words, but as noted in the study,
Pangram’s performance largely holds up on very short passages (< 50 words) and is robust to “humanizer” tools (e.g., StealthGPT), the performance of other detectors becomes case-dependent.
In Experiment 1 of this UMD study, annotators with various levels of knowledge on LLMs were used to predict whether or not a text was AI-generated. After observing that one annotator was nearly perfect at identifying AI text, four additional expert annotators with similar backgrounds in LLM usage were used to classify the same sample of 60. The results from expert votes were compared with commercial detectors like Pangram, Pangram Humanizer, and GPTZero, as well as open source tools like Fast-DetectGPT. During this process, Pangram as compared to other detectors.
Share of AI articles caught (TPR) and false positive rate across all conditions, plus the hardest condition — humanized o1-pro text. Only Pangram kept pace with expert humans.
| Detector | Overall detection (TPR) | False positive rate | Detection on humanized text |
|---|---|---|---|
| Expert humans (majority vote) | 99.3% | 0% | 100% |
| Pangram Humanizers | 99.3% | 2.7% | 96.7% |
| Pangram (base) | 98.0% | 2% | 90.0% |
| GPTZero | 85.3% | 0.7% | 46.7% |
| Fast-DetectGPT | 80.0% | 7.2% | 23.3% |
| Binoculars (Accuracy mode) | 66.7% | 1.3% | 6.7% |
| RADAR | 15.3% | 2% | 0% |
Pangram can accurately detect humanized AI-generated text. This is corroborated by computer scientists at UMD who have noted that Pangram scored highest overall for detecting humanizers and paraphrased text, outperforming other AI detection software with 99.3% accuracy.
Learn more about how Pangram holds up against humanizers
Amanda Caswell at Tom’s Guide stated in an article that after trying dozens of AI detection tools, Pangram “outperformed the others I tried”. Pangram was also shown to be diligently working on reducing the already low incidents of false positives.
David Gewirtz at ZDNET describes Pangram as “a newcomer to our tests that immediately soared into the winners' circle.”
Because AI usage in research papers has increased, there is a concern that this is an indicator of misconduct. Adam Day’s Medium article used Pangram’s AI detection for reliable results on the prevalence of AI content while also concluding that there are legitimate use cases for generative AI in research. Day recommends using Pangram to conduct research, saying: “if someone wants to do a survey of genAI usage in the published literature, I think there’s a great opportunity to do that with Pangram’s tools.”
At Pangram, we believe transparency is essential to trust. We’d love to partner with you to bring AI transparency to your organization.

Destiny is Pangram's Research Analyst Intern. She is also a student at the NYC College of Technology, studying Applied Mathematics and Chemistry. Destiny's work at Pangram is has contributed greatly to investigating AI slop on the internet. Outside of work and education, Destiny is passionate about creative writing and fictional horror.






