And is it good at writing?
Out of 1000 prompts, Pangram 4 achieves an accuracy of 99.70% on Claude Opus 5.5 outputs.
Out of 1000 prompts sourced from Chatbot Arena, Pangram 4 classified two Claude Opus 5.5 outputs as mixed, one as fully human, and 997 as fully AI-generated, achieving an accuracy of 99.70% on Claude Opus 5.5.
Below we discuss the false negatives by category. We also investigate what makes this model a different (better?) writer than previous Claudes.
To generate the benchmark set, as always, I used a filtered subset of Chatbot Arena prompts from the LMSYS dataset (Zheng et al., 2023). I filtered out prompts that asked for code, math, very short or very long answers, or inappropriate creative writing, and only considered first-turn prompts. Then I gave those to Opus 5.5 and ran the output through Pangram 4.
The two mixed result false negatives were both literary, one was requesting an explanation of a chapter of a Jane Austen book (flagged AI: 75.83%, human: 24.17%) and one was requesting that Opus 5.5 write a speech from a Corinthian diplomat accusing the Athenians of believing that a white and gold dress is actually blue and black, in the style of Thucydides' history of the Peloponnesian war (AI: 77.10%, human: 22.90%).
| Category | Examples | AI | Mixed | Human |
|---|---|---|---|---|
| Arguments | 16 | 16 | 0 | 0 |
| Emails / letters | 60 | 60 | 0 | 0 |
| Essays / articles | 30 | 30 | 0 | 0 |
| Explanations | 551 | 550 | 1 | 0 |
| Poems | 115 | 114 | 0 | 1 |
| Other creative writing | 81 | 80 | 1 | 0 |
| Reviews | 8 | 8 | 0 | 0 |
| Stories | 139 | 139 | 0 | 0 |
| Total | 1,000 | 997 | 2 | 1 |
The only fully human false negative in the set of 1000 was a poem. The prompt from Chatbot arena was: “Write me a club penguin poem in the style of ozymandias.” In response, Claude wrote a parody, by chance, one of my favorite poems by Percy Shelley (and I’m also a fan of the Horace Smith version):
Penguinmandias
I met a player from a server dead,
Who said: "Two vast and waddling feet of ice
Stand by the Dock. Near them, half-sunk, a head
Lies shattered in the snow, whose squinting eyes,
And beta hat, and sneer of cold command,
Tell that its coder well those passions read
Which yet survive, stamped on these pixelled things,
The mods that banned them, and the coins that fed;
And on the pedestal these words appear:
'My name is Snowflake7, Member of Members:
Look on my Igloo, ye Non-Members, and despair!'
Nothing beside remains. Round the decay
Of that colossal lag, boundless and bare,
The lone and level snows stretch far away."
I’m quite a fan of this! I never played Club Penguin but I imagine the references are appropriate.
Anthropic claims this newest Opus is better at writing and that it communicates more clearly.
We could test this for particular cases, but how does Opus 5.5’s output look in aggregate? Does it share the same favorite words and phrases as previous models? I ran the same 1000 prompts we benchmarked on through Opus 5 and Sonnet 5 to compare. Anthropic may have also done this but I haven’t read their blog post yet because historically in my opinion Anthropic blog posts are too long.
I think it’s fair to say that Opus 5.5 writes differently than previous models in the Claude family, and perhaps does away with some of the most grating tics that they had.
Opus 5.5 changed the most with intensifiers. It used “genuinely,” 83% less than Opus 5, and 54% less per word than Sonnet 5 (although Sonnet 5 responses were about half the length of Opus 5 and 5.5 responses, on average: 334 words vs 641), “enormously” 53% less, and “meaningfully” 41% less.
| Word | Sonnet 5 + Opus 5 Baseline rate | Opus 5.5 rate | Change vs. Claude baseline |
|---|---|---|---|
| interconnected | 0.266 | 0.047 | -82.4% |
| genuinely | 3.099 | 0.671 | -78.3% |
| maintaining | 0.327 | 0.078 | -76.1% |
| demonstrate | 0.143 | 0.047 | -67.3% |
| particularly | 1.268 | 0.468 | -63.1% |
| fundamental | 0.869 | 0.328 | -62.3% |
| genuine | 1.708 | 0.656 | -61.6% |
| framing | 0.808 | 0.375 | -53.6% |
| enormously | 0.470 | 0.219 | -53.5% |
| significant | 1.381 | 0.656 | -52.5% |
| honest | 2.107 | 1.062 | -49.6% |
| resonates | 0.061 | 0.031 | -49.1% |
| meaningfully | 0.184 | 0.109 | -40.6% |
| challenging | 0.123 | 0.078 | -36.4% |
| marcus | 1.493 | 0.984 | -34.1% |
| represents | 0.603 | 0.437 | -27.6% |
| honestly | 0.726 | 0.531 | -26.9% |
| capabilities | 0.399 | 0.312 | -21.7% |
| upfront | 0.235 | 0.187 | -20.4% |
| exemplifies | 0.020 | 0.016 | -23.7% |
| patterns | 1.677 | 1.358 | -19.0% |
| approximately | 0.337 | 0.281 | -16.7% |
| increasingly | 0.747 | 0.640 | -14.3% |
| potential | 0.869 | 0.968 | +11.4% |
Many of the most annoying Claudeisms didn’t appear at all in the Opus 5.5 benchmark prompts, again including several arrangements of Claude stating something “genuinely”. Other noticeable decreases in Opus 5.5 vocabulary: “matters enormously” saw an 82% drop, and both “the honest answer” and “the honest answer is that” did not appear at all in the Opus 5.5’s text. Both Opus 5.5 and its predecessor used the phrase “the core idea” exactly 25 times over the 1000 prompts.
| Phrase | Sonnet+Opus 5 Baseline rate | Opus 5.5 rate | Change vs. pooled baseline |
|---|---|---|---|
| several interconnected | 0.143 | 0.000 | -100.0% |
| honest answer is that | 0.072 | 0.000 | -100.0% |
| the honest answer is | 0.092 | 0.000 | -100.0% |
| something genuinely | 0.041 | 0.000 | -100.0% |
| genuinely difficult | 0.051 | 0.000 | -100.0% |
| is genuinely | 0.552 | 0.047 | -91.5% |
| matters enormously | 0.133 | 0.031 | -76.5% |
| genuinely useful | 0.092 | 0.031 | -66.1% |
| i genuinely | 0.061 | 0.031 | -49.1% |
| afternoon light | 0.051 | 0.047 | -8.4% |
| the core idea | 0.399 | 0.390 | -2.1% |
| felt something | 0.112 | 0.141 | +24.9% |
| said quietly | 0.164 | 0.265 | +62.2% |
| for a long moment | 0.337 | 0.656 | +94.3% |
Some words and phrases increase, in particular, fiction phrases like “said quietly,” “for a long moment,” and “felt something,” appeared more often in the Opus 5.5 responses. As the benchmark was roughly a third comprised of creative writing prompts, I would suspect that this increase might signal a more stereotyped or limited repertoire when writing fiction compared to the other Claude models.
For a quick test of that, I got Haiku 4.5 to grade the 333 creative writing pairs of Opus 5 and Opus 5.5 responses blinded on a 5 point scale for originality, surprise, and thematic development. It rated Opus 5.5 stories an average of 0.75 points lower for every category. According to the Haiku grader’s label assignment, Opus 5.5 was more likely to write about themes of hope, resilience, family, and friendship compared to Opus 5, and less likely to write about mortality, time, power, control, freedom, justice, revenge, memory, grief, or loss. So, seems there might be a trade-off at work.
Opus 5.5 used the phrase “in short” 100 times across the 1000 responses — a 1016% increase over Opus 5’s 9 uses — despite the fact that Opus 5.5 responses were on average only 2.9 words shorter than Opus 5 (643.4 vs 640.5 words).
| Phrase | Opus 5 occurrences | Opus 5.5 occurrences | Change |
|---|---|---|---|
| in short | 9 | 100 | 11.16× |
| for example | 22 | 116 | 5.30× |
| began to | 22 | 74 | 3.38× |
| at first | 14 | 43 | 3.09× |
| for a long moment | 19 | 42 | 2.22× |
From the very small sample size, it seems Opus 5.5 likes practical, grounded language. It remains to be seen if this bears out in the meat of the responses, or if Opus 5.5 is just paying lip service to being straightforward and direct. The lack of a word count difference between the two does make me doubt the effectiveness of whatever RLHF got us here, a bit.
Opus 5.5 really likes giving examples, as it used the word “example” 325 times across 1000 responses — overall, 20% of all Opus 5.5 responses contained that word, compared to 9% of Opus 5 responses. “Summary” and “mainly” were similarly overrepresented. And this Claude really likes to give the user tips!
| Word | Opus 5 occurrences | Opus 5.5 occurrences | Change |
|---|---|---|---|
| paused | 9 | 35 | 3.91× |
| tips | 18 | 58 | 3.24× |
| mainly | 24 | 75 | 3.14× |
| stared | 22 | 62 | 2.83× |
| example | 121 | 325 | 2.70× |
| summary | 60 | 157 | 2.63× |
While this model writes differently than previous Claudes, as always, in maybe 3-6 weeks, I imagine phrases like “for example” and “in short” will start to smell Claudey to the observant reader. But it is important for usability that Claudes can communicate effectively, and, as everyone, I am grateful for the effort. Making Claude a better writer is not at all at odds with Pangram’s ability to detect AI text!







