Xbench
rolling 7d

Xbench

Measuring the Mandate of Heaven

posts readi
people heardi
firsthand opinionsi
reviewer correctionsi
01

Sentiment

firsthand stance

Positive / mixed / negativei

positivemixednegativenet = % positive − % negative
0255075100%
nnet
02

Switching

completed moves

Where people actually movedi

the actual tweetsopens on X ↗
03

Head-to-head

direct comparisons

Comparison matrix

read across each row · select a cell for the postsi
the actual tweetsopens on X ↗

global preference

XbenchPrefi

firsthand votes
Opponent-adjusted preferencebars show uncertainty
Bradley-Terry, centred on 1000 matchups
04

Harnesses

head-to-head

Codex vs Claude Codei

completed moves ·

the actual tweetsselect a ribbon · opens on X ↗

all three

Harness matrix and ratingi

the three with enough posts
the actual tweetsopens on X ↗
05

Reasons

The reasons for people's sentiments, preferences, and switches.

the shape of each one

Where the shapes overlapi

models and harnesses side by side

Models

Harnesses

hollow point: fewer than five people · models under 30 authors left offselect up to four · hover a name to preview it

one at a time

Why each one is liked and dislikedi

dimensions with fewer than five people are dimmed
06

Methods, limitations, and open questions

How this was read, where it is weak, and what I would do next.

limitations

Where I think this is wrong

my read, not the data's

Happiness and bias

To generalise, people on X are upset with Anthropic right now, and that carries into how they talk about its models.

People talking their book

There’s two trillion dollars of SpaceX stock making its weight felt on twitter. While I root for Elon, It’s hard to take Gavin Baker seriously when he’s on a book tour calling Grok Bot another “Claude Code moment”.

I feel that the love for Grok 4.6 and Grok Build is mostly legitimate, but the Grok Bot heat feels possibly a bit manufactured.

Price really matters

The smartest model is not the most preferred model. Stated preference tracks price more than intelligence. Cheap and good beats best and expensive, at least on X, at least this week.

rolling seven days, one labeler pass, about ten percent human review · * GPT-6 Astra reached Pro and Codex users on Sep 4, so its sample is one day old

open questions

What I would do differently

next run

More days

The pulls so far cost about $130 in X API reads for a rolling seven-day window. The same spend every day for a month would give a far steadier signal, and I would expect some of the rankings to move. It is also the main reason the XbenchPref error bars are as wide as they are.

Not all signal is equal

I would like to try an SF-only cut, sorry New York! Posts mostly carry no location, so every author would need a profile read at about a cent each, and the sample would need to be roughly fifty times larger to hold up.

More models

DeepSeek would be a fun one to add.

every counted post links back to X

classification

What counted, and what did noti

labels v2

Counts toward the charts

Stances from described use · cost, quota and rate-limit complaints · completed first-person switches · a reply that adds its own experience.

Shown but not scored

Recommendations and rankings without use · benchmark reposts · one-word poll answers · agreeing with someone else's experience.

Left out entirely

Pre-release speculation · vendor and reseller promotion · news and fact-checks · jokes · reply accounts writing as an assistant · anything whose target cannot be named.