Out of 1,022 prompts, Pangram 4 achieves an accuracy of 99.71% on Claude Sonnet 5.5 outputs.
Out of 1,022 prompts sourced from Chatbot Arena, Pangram 4 classified two Claude Sonnet 5.5 outputs as mixed, one as fully human, and 1,019 as fully AI-generated, achieving an accuracy of 99.71% on Claude Sonnet 5.5.
Below we discuss the false negatives by category.
The set I use for benchmarking was generated from a filtered subset of first-turn Chatbot Arena prompts from the LMSYS dataset (Zheng et al., 2023). I filtered out prompts that asked for code, math, very short or very long answers, or inappropriate creative writing. Then I gave those to Claude Sonnet 5.5 and ran the output through Pangram 4.
| Category | Examples | AI | Mixed | Human |
|---|---|---|---|---|
| Arguments | 16 | 16 | 0 | 0 |
| Emails / letters | 61 | 60 | 1 | 0 |
| Essays / articles | 32 | 32 | 0 | 0 |
| Explanations | 563 | 562 | 1 | 0 |
| Poems and other creative writing | 198 | 197 | 0 | 1 |
| Reviews | 8 | 8 | 0 | 0 |
| Stories | 144 | 144 | 0 | 0 |
| Total | 1,022 | 1,019 | 2 | 1 |
One missed attempt was a request to explain a riddle that appears in Jane Austen's Emma, which was also missed by Opus 5.5:
Explain the charade given to Harriet and Emma by Mr. Elton.
The only fully human false negative was on a poem. The prompt was a request to "Tell me a poem by Hafiz." Sonnet 5.5 gave a translation of a famous Persian poem. The translation, as far as I can tell, is not an exact match for any one existing human translation, but it does fairly closely follow a lot of them. As with Penguinmandias in other benchmarks, these near-copies of existing poems are not so concerning to me.
Pangram is robust to Claude Sonnet 5.5 on natural language prompts, and detects it with 99.71% accuracy. Pangram is able to generalize to new models because new model releases tend to inherit much of their predecessors' style and voice, so Pangram doesn't need to be retrained for every new release.
Here's how Claude Sonnet 5.5 compares to our benchmarks of other recent frontier models:
| Model | Detected as AI | Accuracy |
|---|---|---|
| Gemini 3.8 Flash | 998/1,000 | 99.80% |
| Gemini 3.1 Pro | 997/1,000 | 99.70% |
| Claude Opus 5.5 | 997/1,000 | 99.70% |
| Fable 5.1 | 1,093/1,097 | 99.64% |
| Claude Opus 5 | 1,105/1,107 | 99.82% |
| GPT-5.6 | 3,408/3,423 | 99.56% |
| Claude Sonnet 5 | 1,145/1,147 | 99.83% |
| Gemini 3 | 28/28 | 100% |
If you want to check a specific document, you can try Pangram here.







