firsthand stance

Xbench
Measuring the Mandate of Heaven
Sentiment

Switching
completed moves
Where people actually movedi
Head-to-head
direct comparisons
Comparison matrix
global preference
XbenchPrefi
Harnesses


head-to-head
Codex vs Claude Codei
completed moves · —
all three
Harness matrix and ratingi
Reasons
The reasons for people's sentiments, preferences, and switches.
the shape of each one
Where the shapes overlapi
Models
Harnesses
one at a time
Why each one is liked and dislikedi
Methods, limitations, and open questions
How this was read, where it is weak, and what I would do next.
limitations
Where I think this is wrong
Happiness and bias
To generalise, people on X are upset with Anthropic right now, and that carries into how they talk about its models.
People talking their book
There’s two trillion dollars of SpaceX stock making its weight felt on twitter. While I root for Elon, It’s hard to take Gavin Baker seriously when he’s on a book tour calling Grok Bot another “Claude Code moment”.
I feel that the love for Grok 4.6 and Grok Build is mostly legitimate, but the Grok Bot heat feels possibly a bit manufactured.
Price really matters
The smartest model is not the most preferred model. Stated preference tracks price more than intelligence. Cheap and good beats best and expensive, at least on X, at least this week.
open questions
What I would do differently
More days
The pulls so far cost about $130 in X API reads for a rolling seven-day window. The same spend every day for a month would give a far steadier signal, and I would expect some of the rankings to move. It is also the main reason the XbenchPref error bars are as wide as they are.
Not all signal is equal
I would like to try an SF-only cut, sorry New York! Posts mostly carry no location, so every author would need a profile read at about a cent each, and the sample would need to be roughly fifty times larger to hold up.
More models
DeepSeek would be a fun one to add.
classification
What counted, and what did noti
Counts toward the charts
Stances from described use · cost, quota and rate-limit complaints · completed first-person switches · a reply that adds its own experience.
Shown but not scored
Recommendations and rankings without use · benchmark reposts · one-word poll answers · agreeing with someone else's experience.
Left out entirely
Pre-release speculation · vendor and reseller promotion · news and fact-checks · jokes · reply accounts writing as an assistant · anything whose target cannot be named.

