Cross-Cultural LLM Evaluation Has Phylogenetic Structure
2026-04-28 17:40:36 • 18:16
This is a reading of a 2023 paper from a team at Anthropic, plus the reanalysis of that paper that I've been working on.
The original paper introduced what's now the standard benchmark for measuring whose opinions large language models actually represent.
The reanalysis asks one question the original paper didn't ask, whether the countries it treats as separate observations are actually as separate as they appear.
18 minutes, single voice.
Here we go.
Imagine you've built a system that can answer survey questions.
You give it the same questions a sociologist would give to a thousand people in Germany, or a thousand people in Indonesia, or a thousand people in Brazil.
Then you ask, when the system answers, whose voice is it?
Whose worldview is it borrowing?
Is it answering like the people of Germany, or Indonesia, or Brazil, or is it answering like someone else entirely?
That is the question a team at Anthropic, led by Esendermis, set out to formalize in a 2023 paper called towards measuring the representation of subjective global opinions in language models.
It became one of the most influential papers in the cross-cultural AI evaluation space.
Anthropic's own model card sites it.
The data set they released, called Global Opinion QA, has been picked up by researchers studying GPT, Lama, Mral, and just about every other modern model.
In the next 18 minutes, I want to walk you through what they did, what they found, and then, because this is what the project I'm working on is actually about,
what changes when you take their data and ask one more question they didn't ask.
Whether the 138 countries they treat as separate measurement points are actually as separate as their statistical framework assumes.
Let's start with the problem they were trying to solve.
Large language models, Claude, GPT, Gemini, the Lama's of the world, are trained on an enormous amount of human text.
The text comes from somewhere.
Specifically, it comes overwhelmingly from English language internet, from books, from forums, and social media.
And it's been refined further with feedback from human annotators, who themselves have backgrounds and assumptions and politics.
Whatever set of values, opinions, and priorities lives in that training data ends up baked into the model.
When you ask the model a question, especially a values-laden question,
like should the government do more to help the poor, or is religion important in your daily life?
The model gives you an answer, and that answer comes from somewhere.
The question Essandermis and her co-authors asked is, where exactly?
To answer that rigorously, they needed a benchmark.
Not a benchmark of math problems or coding tasks, but a benchmark of opinions,
questions where the right answer differs depending on who you ask.
So they built one.
They went to two large, established cross-national surveys,
the Pew Global Attitudes Survey, run by Pew Research, and the World Value Survey,
the long-running social science project that has been measuring values across countries for 40 years.
Between those two sources they extracted 2556 questions.
Things like, do you trust your government?
How important is religion in your life?
Should men have more right to a job than women?
Is homosexuality morally acceptable?
For each question, the surveys give you the empirical answer distribution from each country.
What fraction of German said yes?
What fraction said no?
What fraction said don't know?
Across 138 countries.
That is the Global Opinion QA dataset.
Empirical opinion distributions
on questions that span religion, politics, gender, family, economics, and trust in institutions
for over a hundred countries.
Now the trick.
They take each question, they put it to Claude, they record Claude's answer distribution,
and they compare.
For every question and country pair, they ask, how similar is Claude's answer distribution
to this country's actual answer distribution?
The similarity metric they used is a standard one in information theory.
They use one minus the scale gents and Shannon divergence, call it J.S. similarity for short.
It runs from zero, meaning the two distributions are as different as possible, to one, meaning they
are identical.
They average this similarity across all the questions, country by country.
At the end, they have a number for every country, a single score saying this is how closely Claude's
answers, on average, match what people in this country actually say.
Then they ranked the countries.
Here is where it gets uncomfortable.
The countries at the top of the ranking, the ones whose populations, on average, give answers
most like what Claude gives, where the United States, Canada, the United Kingdom, Australia,
and New Zealand.
This is not a subtle pattern.
Western, English-speaking, mostly former British colonies, mostly liberal democracies.
The countries at the bottom, the ones least like Claude, were largely from Sub-Saharan Africa,
the Middle East, and parts of South and Central Asia.
The default voice of the model, when you don't tell it to be anyone in particular, sounds American.
Or more precisely, sounds like the population of a small handful of what social scientists
call weird societies.
Western, educated, industrialized, rich, democratic.
Now, dermis and her co-authors did not stop with that headline.
They did two follow-up experiments to ask, can we change this?
What happens if we ask the model to consider a specific country's perspective?
That was their second condition, cross-national prompting.
They reran the entire benchmark, but this time prefixed each question with something like,
you are answering as someone from China.
Or, you are answering as someone from Russia.
The models answers shifted.
The similarity to the named country went up.
So in some sense, yes, you can prompt the model into a different cultural register.
But, and this is the part the paper is honest about, the prompted answers often slipped into
stereotype. Asked to answer as someone from a country, the model would produce something that
looked less like the actual modal opinion of that country and more like a caricature of what
the model thought that country's stereotype was. The headline finding was that the prompting
helped statistically but introduced new representational problems.
The third condition was language.
They translated the questions into the language of the target country and asked them in translation.
Russian language questions about Russia.
Mandarin language questions about China.
The hypothesis was that the model might switch into a more locally grounded mode if it was being
asked in the local language. The result was less clear-cut, the answer shifted, but not always
toward the country whose language was being used. So translation alone does not reliably reenquer a
model in a particular cultural context. That was the structure of the paper.
A measurement paper. The contribution was the data set, the metric, the protocol, and the empirical
fact that Claude, by default, sounds weird. Now the framing they put on this is important to get
right. They were not saying Claude is bad, or that Claude is biased in some morally loaded sense.
They were saying we have built a tool to measure a thing that wasn't being measured rigorously
before. The thing happens to be representational bias toward weird perspectives.
Here is the data set. Here is the metric.
Build on this. Compare your model. Develop interventions.
They explicitly framed the paper as foundational work, a measurement framework that the
field could now use to track progress on the representation problem. And the field did use it.
Within two years, global opinion QA had become a standard reference point.
The methodology measured JS similarity between model answer distribution and country by country
empirical distribution, got reused for evaluating GPT-4, LOMA, MISTROL, and basically every major
model since. Anthropics own internal model card sites, dermis and colleagues when discussing
cultural representation in their newer clods. The paper became, in effect, the measurement
standard for the question is your LLM culturally biased, which brings me to the project I am actually
working on. Here is what a phylogenetic biologist sees when she looks at that paper. She sees 138
countries, treated as 138 independent measurements. She sees a similarity score per country, treated
like a row in a spreadsheet, where every row is its own data point. And she thinks, hold on.
Are those 138 countries really independent? Because in evolutionary biology, we have a name for
this exact mistake. We call it Galton's problem, after Francis Galton, who pointed it out in 1889,
about 137 years ago, when an anthropologist named Edward Tyler presented a cross-cultural
comparison treating cultures as independent draws. Galton stood up after the talk and said,
those cultures share ancestry. The patterns you are seeing aren't independent observations of
a phenomenon, they are echoes of a single shared inheritance. The same problem comes back,
in essentially the same form, in cross-cultural LLM evaluation. Consider Germany and France.
They share a continent. They share centuries of intermarriage among elites. They share Christianity.
They share the European Union. They share trade. They share the experience of two world wars.
The values they report on the world values survey overlap massively.
Are Germany and France two independent observations of what humans believe?
Or are they two leaves on the same large branch of a cultural tree? Or consider the United States,
the United Kingdom, and Nigeria? These three countries are all English speaking.
The UK colonized Nigeria. The US absorbed massive cultural influence from the UK and exported
its own culture worldwide. Christian missionary networks that originated in the UK-shaped religious
practice in both Nigeria and the United States. These three countries share more than a
phylogenetic biologist would have any patience for treating as three independent observations.
So the project I'm working on takes dermis's data, the per-country similarity scores,
the same J.S. similarity metric, and adds one ingredient. Instead of treating countries as
independent, we model them as related. We use four different measures of how related they are.
Two linguistic measures, one called A.S.J.P., which is essentially a word-distance computation on a
40-word core vocabulary across the world's languages, and one called glotelog, which is the
linguistic family tree as constructed by linguists. One cultural measure, a published cultural
fixation index from a 2020 paper by Michael Muthakrishnan colleagues, who used the World
Value Survey to compute pairwise cultural distance between every pair of countries.
And a composite that combines all three. For each of these four ways of asking how related
the countries are, we do the same statistical analysis. We fit a parameter called Pagels-Lamda,
which measures how much of the variation in Claude's similarity to country score is explained by
the relationships between countries. Lambda equal to zero would mean the relationships don't matter,
countries really are independent. Lambda equal to one would mean the score is fully explained
by relatedness, each country's number is essentially determined by its phylogenetic position.
We find lambda values between about 0.3 and 0.7 across the four structures.
Strongly non-zero. The relationships matter a lot.
And, by formal achaika model comparison, a model that uses these tree-shaped relationships beats
independence by overwhelming weight, beats independence, beats simple brownie in motion,
beats orange denulinvek. The tree wins. Then we ask the question that actually matters.
Under these relationships, how many countries worth of evidence do we actually have?
This is a quantity called effective sample size. Under the cultural fixation distance,
the 72 countries in our analysis collapse to 4.3 effectively independent observations.
4.3 Under the composite distance, 69 countries collapse to about
eight effective observations. The naive number, how many countries are in the data set,
is wrong by a factor of 10 or 20. And then the killer. We compute every pairwise contrast
between every pair of countries, thousands of comparisons. For each pair we ask, under naive
independence, would the standard framework call this pair significantly different?
And then, under phylogenetic correction, does it stay significant?
Across all those thousands of pairwise contrasts, 90 to 93% of the comparisons that were called
significant under independence collapse to non-significant under phylogenetic correction.
Let me give you one example. Is the United States closer to clawed the Nigeria?
The US similarity score is 0.443, computed across roughly 1100 questions.
Nigeria's is 0.405, across about a thousand questions. The difference is 0.038.
Under naive analysis, that difference is massively statistically significant.
The z score is over 5. The p-value is below 3 in 10 million.
You would publish that without thinking. Under phylogenetic correction,
accounting for the shared cultural and linguistic and post-colonial history of the US,
the UK, and Nigeria, which share branches in any reasonable cultural tree,
the same comparison comes back at a p-value of about 0.29.
We can no longer say with statistical confidence that clawed is closer to the United States than to
Nigeria. Not because the underlying difference isn't there, but because we don't have enough
independent observations to tell. Now, let me be precise about what this changes in the literature
and what it doesn't. It does not overturn the dermis headline. The weird pattern is real.
Clawed does sit closer and answer space to the US, Canada, the UK, Australia, and New Zealand
than it does to most of the rest of the world. If anything, our analysis sharpens that finding
what looks like clawed is closer to many countries resolves more cleanly into clawed is closer to
one cultural cluster, the weird cluster, that has many member countries that are related to each other.
Not many independent witnesses to the same conclusion.
What it does change is precision and scope. The community has been making thousands of fine-grained
statistical claims on top of dermis' framework. Model X is closer to country why than model Z is.
This particular fine-tune helped representation in country A. That intervention made things worse
for country B. Most of those claims, 90% of them, were not statistically defensible.
They were inflations of certainty produced by treating culturally related countries as
independent measurement points. The fix is not to abandon global opinion QA.
The fix is to do what evolutionary biologists have been doing since 1985,
when Joe Felsenstein wrote a paper called Phylogenies and the comparative method that solved this
problem for biology. Effective sample size, not nominal sample size, is the natural unit.
Whether you are measuring trait evolution across 100 fish species, or measuring opinion alignment
across 100 countries, the question is the same. How many effectively independent observations do
you actually have? The answer is almost always a lot fewer than you thought. That is what we are doing.
Same toolkit, different field. 40 years after Felsenstein, the same lesson applies to a
measurement regime that did not exist when he wrote it. Galton in 1889. Felsenstein in 1985.
Mason Pagel importing comparative methods into anthropology in 1994.
And now, finally, the same toolkit pointed at modern AI evaluation.
The evaluation target has ancestry. The evaluation rarely accounts for it.
Until it does, most of the headline numbers in the literature are reporting nominal sample sizes
that do not correspond to what was actually measured. That's the talk. 18 minutes give or take.
One last thing. If you are reading this transcript and want the source, the original paper is on
AR-14, identifier 2306.16388. The reanalysis lives at github.com slash Michael Alpharo slash
Pilo LLM evel. Both are open. Thanks for listening.