Pangram 4 produced zero false positives across all 25,996 human-written essays in the PERSUADE 2.0 corpus, including every subgroup by disability status, gender, grade level, ethnicity, ELL status, and economic disadvantage.
We evaluated Pangram 4 on the PERSUADE 2.0 dataset of diverse student essays and found zero false positives across all 25,996 essays. Pangram produced zero false positives in every reported subgroup.
It is important to understand how Pangram performs across diverse kinds of authors. Datasets of human-written essays that include author demographic information are not easy to come by, but evaluating Pangram's performance on text produced by different kinds of writers — including students with disabilities, people of color, non-native English speakers, and economically underprivileged groups — is essential to understanding whether the model treats every writer fairly.
The PERSUADE dataset helps address this. Developed to advance the creation of unbiased algorithms to assess student writing, PERSUADE ("Persuasive Essays for Rating, Selecting, and Understanding Argumentative and Discourse Elements") is a labelled collection of nearly 26 thousand essays from American middle and high school students written between 2010 and 2020. The dataset covers students from diverse backgrounds and ability levels, and each essay is labeled with demographic information about the author, such as grade level, gender, ethnicity, and socioeconomic background.
The PERSUADE dataset offers an excellent opportunity to evaluate Pangram for two reasons. Firstly, Pangram does not train on the PERSUADE dataset, so we can use it to test how Pangram generalizes to previously unseen student writing. Secondly, PERSUADE's demographic data — including writing samples from vulnerable populations — allows us to examine whether Pangram performs fairly across different demographic groups.
Key takeaway: We evaluated Pangram 4 on the PERSUADE 2.0 dataset of diverse student essays and found zero false positives across all 25,996 essays. We also examined the results by disability status, gender, grade level, race and ethnicity, English language learner status, and economic disadvantage. Pangram produced zero false positives in every reported subgroup.
The breakdown below describes how researchers defined each demographic attribute in the PERSUADE dataset, and shows the number of essays represented in each group.
Disability status in the PERSUADE 2.0 dataset refers to students that have an IEP or 504 plan, and could encompass students with a wide range of learning disabilities or health conditions.
| Group | n | FPR |
|---|---|---|
| Identified as having a disability | 2,688 | 0% |
| Not identified as having a disability | 18,135 | 0% |
| Missing/not reported | 5,173 | 0% |
Gender in the PERSUADE dataset is self-identified, roughly a 50/50 split between male and female. Pangram demonstrates zero observed false positives on both gender groups.
| Group | n | FPR |
|---|---|---|
| Female (F) | 13,142 | 0% |
| Male (M) | 12,854 | 0% |
Pangram 4 produced zero false positives across every observed grade level, which demonstrates Pangram's robustness for many different levels of writing proficiency.
| Group | n | FPR |
|---|---|---|
| Grade 6 | 1,372 | 0% |
| Grade 8 | 9,629 | 0% |
| Grade 9 | 2,066 | 0% |
| Grade 10 | 8,274 | 0% |
| Grade 11 | 3,083 | 0% |
| Grade 12 | 404 | 0% |
| Missing/not reported | 1,168 | 0% |
The majority of writers identified in the PERSUADE database identified as White (45%), followed by Hispanic (25%), then Black (19%), which roughly mirrors the demographic spread of the US student population.
| Group | n | FPR |
|---|---|---|
| White | 11,571 | 0% |
| Hispanic/Latino | 6,560 | 0% |
| Black/African American | 4,959 | 0% |
| Asian/Pacific Islander | 1,743 | 0% |
| Two or more races/Other | 1,022 | 0% |
| American Indian/Alaskan Native | 141 | 0% |
While some early research had indicated that old AI detection programs may have shown bias against English as a second language and English language learners, Pangram has demonstrated essentially no bias against these authors, and has a 0% observed FPR on ELL students in the PERSUADE dataset.
| Group | n | FPR |
|---|---|---|
| ELL | 2,244 | 0% |
| Not ELL | 22,451 | 0% |
| Missing/not reported | 1,301 | 0% |
Economic disadvantage in the PERSUADE corpus is defined as student eligibility for certain federal assistance, such as free or reduced-price school lunch and Supplemental Nutrition Assistance Programs.
| Group | n | FPR |
|---|---|---|
| Economically disadvantaged | 9,643 | 0% |
| Likely not economically disadvantaged | 11,116 | 0% |
| Missing/not reported | 5,237 | 0% |
Across all 25,996 human-written essays in the PERSUADE 2.0 corpus, Pangram 4 produced zero false positives across every subgroup, including disability status, gender, grade level, ethnicity, English language-learner status, and economic disadvantage. This result indicates that Pangram generalizes well to diverse author groups without disproportionately misclassifying these writers' work as AI-generated.
We aim to make false positives as rare as possible across all author groups, and continued evaluation is important. However, within the scope of PERSUADE 2.0, Pangram 4 demonstrates zero false positives across every demographic subgroup.

Katherine Thai is the Founding AI Research Scientist at Pangram Labs, an AI detection startup. She completed her PhD in Computer Science under the supervision of Mohit Iyyer at the University of Massachusetts Amherst in December 2025, where her work was focused on evaluating LLMs on tasks related to literary analysis.







