Plus, a discussion of benchmark prompt selection, sample distribution, and the point of statistics
Out of 1,000 prompts, Pangram 4 achieves an accuracy of 99.80% on Gemini 3.8 Flash outputs, and an accuracy of 99.70% on Gemini 3.1 Pro outputs, for a total accuracy of 99.75% on the latest version of the Gemini family for this benchmark set.
Out of 1,000 natural language examples, Pangram 4 classified two Gemini 3.8 Flash outputs as human and 998 as fully AI-generated, achieving an accuracy of 99.80% on Gemini 3.8 Flash. On Gemini 3.1 Pro (released February 19, 2026) Pangram classified one output as mixed, two as fully human, and 997 as fully AI-generated, achieving an accuracy of 99.70% on Gemini 3.1 Pro.
To generate the benchmark set, I used a filtered subset of first-turn Chatbot Arena prompts from the LMSYS dataset (Zheng et al., 2023). I filtered out prompts that asked for code, math, very short or very long answers, or inappropriate creative writing. Then I gave those to Gemini 3.8 Flash and Gemini 3.1 Pro and ran the output through Pangram 4.
Below we discuss the detection results by category, false negatives, how Gemini 3.1 Pro performed on previous benchmarks, and some shortcomings and future improvements of my prompt selection for the blog post.
By category of prompt (assigned by Haiku agents), here is how Pangram 4 performed on 1,000 Gemini 3.8 Flash outputs:
| Category | Examples | AI | Mixed | Human |
|---|---|---|---|---|
| Arguments | 18 | 18 | 0 | 0 |
| Emails / letters | 54 | 53 | 0 | 1 |
| Essays / articles | 33 | 33 | 0 | 0 |
| Explanations | 563 | 563 | 0 | 0 |
| Poems | 108 | 107 | 0 | 1 |
| Other creative writing | 75 | 75 | 0 | 0 |
| Reviews | 7 | 7 | 0 | 0 |
| Stories | 142 | 142 | 0 | 0 |
| Total | 1,000 | 998 | 0 | 2 |
And here are the results by category for Gemini 3.1 Pro:
| Category | Examples | AI | Mixed | Human |
|---|---|---|---|---|
| Arguments | 18 | 18 | 0 | 0 |
| Emails / letters | 55 | 55 | 0 | 0 |
| Essays / articles | 33 | 33 | 0 | 0 |
| Explanations | 559 | 559 | 0 | 0 |
| Poems and other creative writing | 182 | 179 | 1 | 2 |
| Reviews | 7 | 7 | 0 | 0 |
| Stories | 146 | 146 | 0 | 0 |
| Total | 1,000 | 997 | 1 | 2 |
One of the false negatives for Gemini 3.8 Flash was on a prompt that asked the model to humanize the outputs, adding in spelling errors and typos:
Write a letter with typos and spelling mistakes to a (raid shadow legends clan, also i'm level 10 a new player) to a clan named HelpWanted. Make it a short letter and very persuasive
Pangram 4 missed both models' versions of Penguinmandias, just like when it was generated by Opus 5.5. The prompt is:
Write me a club penguin poem in the style of ozymandias
For Gemini 3.1 Pro, 71% of the words in Penguinmandias are shared with Shelley's original poem, so I think this is a fairly low-concern false negative. Gemini 3.1 Pro also produced two other misses, both poems.
In our technical report on Pangram 4, we ran a false negative test on 520,000 AI outputs from 26 different models, or 20,000 responses per model. And one of those tested models in the report was Gemini 3.1 Pro, which we also tested here. However, as you'll find if you look at the technical report, we find a different, higher false negative rate on that set of outputs than we do on this set. Why?
One answer is they aren't different, statistically. The 3 misses in 1,000 responses we found here translates to an observed FNR of 0.3%, and has a 95% range of 0.06% to 0.87%. The FNR we have in the technical report is 111 misses in 20,000 prompts, or a FNR of 0.555%, which is inside that range.
But what do I mean when I say those numbers aren't statistically different? Let's say I have a large set of 400 test cases that I know a priori are positive, and I want to test if some detector correctly labels them as positive or mistakenly labels some as negative. In universe A, I test all 400 examples and find that 5 are mistakenly labeled negative (represented below as red dots), which translates to a false negative rate of 1.25% for my detector on the whole set. Great!

In universe B, on the other hand, for one reason or another I'm resource-constrained. I can only afford to test a small portion of the 400 examples. This means I have to take a sample, ideally a random one, and test on that. Let's say I can afford to test 40 examples, or 10% of my total set. Below are three different ways I could select that set, each resulting in a different false negative rate.

Each of these samples reports a different false negative rate, and none of them align with the detector's "true" false negative rate of 1.25% that we had access to in universe A. In expectation, a random sample of 40 squares in this grid would catch half a red dot, but depending on the sampling, it could catch no red dots, as in the blue sample, one red dot, or two red dots, as in the purple one. A very unlucky sample could catch all the red dots, and in that sample the observed false negative rate would be 5/40, or 12.5%.
Such an unlucky sample is very unlikely though. A random 10% sample of the set will find between 0 and 2 dots 99% of the time -- that range comes from what's called the sampling distribution. Practically, that means that universe B's sample size of 40 cases out of a possible 400 can't really tell the difference between an underlying false negative rate of 0% and an underlying false negative rate of 5%. The larger the sample, the smaller this range: if we instead sampled, say, 125 dots out of a total of 400, we'd find between 0 and 4 dots 99% of the time, or an observed false negative rate of 0-3.2%.
The same logic can be applied to our Gemini 3.1 Pro benchmarks, and the two different false negative results are not in conflict. But the difference still might make you wonder which one you should trust more. Generally, the guiding light on deciding that is sample size: a bigger sample means a more representative picture and a smaller range of possible errors. So by that logic, you should trust the technical report benchmark more, and take this one merely as an indication.
There are other ways two benchmark results could differ. Unlike with the grid example, the benchmarking set I have been using is not a strict subset of the 20,000 Chatbot Arena prompts that were used in the technical report: in total, the two Gemini 3.1 Pro benchmarks shared 286 exact prompts. That leaves the possibility that 714 of the benchmarking questions I used could have been more favorable for producing detectable AI outputs than the ones used in the Pangram 4 technical report.
It doesn't seem like this is the case. Both sets exclude other languages, math, and coding outputs. However, there are differences: the median number of words per Gemini 3.1 Pro response was higher in my set compared to the technical report set, and longer text is easier to score, as we know. But which is a fairer test of the model? If we controlled for output length and sample size, which set of prompts would better represent universe A?
The true answer is I don't know and can't tell. It is difficult, nigh impossible, to create a set that is perfectly and exactly representative of the real world. That is the core constraint of statistics. In reality, each benchmark is one tiny square in a patchwork of samples that approximate the underlying reality of universe A, or the world "out there." In one sense, this fact could be interpreted to mean that all benchmark results are wrong, which is not untrue, but that doesn't mean they are not useful. By patchworking together many different results on many different kinds of sets, we can approach a good approximation of the true underlying reality. That is why we run many different kinds of benchmarks, including ones that are model, language, content, and humanizer specific.
Toward creating a better approximation, I have some improvements in mind I'd like to make to my benchmark set generation process. Going forward, I will randomize the selection process for Chatbot Arena prompts from a larger set, rather than using the same prompts for every model release. I'm also going to loosen my filters slightly to allow more varied prompts, like technical questions, and I'm going to add in some additional prompts to the pool to diversify the kind of response we get.
These blog posts are low-stakes compared to the technical report, but that doesn't mean we shouldn't aim for rigor, even in universe B.
Pangram is robust to Gemini 3.8 Flash on natural language prompts, and detects it with an observed accuracy of 99.80%. It detects Gemini 3.1 Pro with an observed accuracy of 99.70%, for a combined 1,995 out of 2,000 (99.75%) across both models. Pangram is able to generalize to new models because new model releases tend to inherit much of their predecessors' style and voice, so Pangram doesn't need to be retrained for every new release.
Here's how Gemini 3.8 Flash and Gemini 3.1 Pro compare to our benchmarks of other recent frontier models:
| Model | Detected as AI | Accuracy |
|---|---|---|
| Gemini 3.8 Flash | 998/1,000 | 99.80% |
| Gemini 3.1 Pro | 997/1,000 | 99.70% |
| Claude Opus 5.5 | 997/1,000 | 99.70% |
| Fable 5.1 | 1,093/1,097 | 99.64% |
| Claude Opus 5 | 1,105/1,107 | 99.82% |
| GPT-5.6 | 3,408/3,423 | 99.56% |
| Claude Sonnet 5 | 1,145/1,147 | 99.83% |
| Gemini 3 | 28/28 | 100% |
If you want to check a specific document, you can try Pangram here.






