Cross-Cultural LLM Evaluation Has Phylogenetic Structure

2026-04-28 17:40:36 • 18:16

-

This is a reading of a 2023 paper from a team at Anthropic, plus the reanalysis of that paper that I've been working on.

0:07

The original paper introduced what's now the standard benchmark for measuring whose opinions large language models actually represent.

0:15

The reanalysis asks one question the original paper didn't ask, whether the countries it treats as separate observations are actually as separate as they appear.

0:25

18 minutes, single voice.

0:28

Here we go.

0:30

Imagine you've built a system that can answer survey questions.

0:34

You give it the same questions a sociologist would give to a thousand people in Germany, or a thousand people in Indonesia, or a thousand people in Brazil.

0:43

Then you ask, when the system answers, whose voice is it?

0:47

Whose worldview is it borrowing?

0:50

Is it answering like the people of Germany, or Indonesia, or Brazil, or is it answering like someone else entirely?

0:57

That is the question a team at Anthropic, led by Esendermis, set out to formalize in a 2023 paper called towards measuring the representation of subjective global opinions in language models.

1:10

It became one of the most influential papers in the cross-cultural AI evaluation space.

1:16

Anthropic's own model card sites it.

1:19

The data set they released, called Global Opinion QA, has been picked up by researchers studying GPT, Lama, Mral, and just about every other modern model.

1:30

In the next 18 minutes, I want to walk you through what they did, what they found, and then, because this is what the project I'm working on is actually about,

1:38

what changes when you take their data and ask one more question they didn't ask.

1:42

Whether the 138 countries they treat as separate measurement points are actually as separate as their statistical framework assumes.

1:51

Let's start with the problem they were trying to solve.

1:54

Large language models, Claude, GPT, Gemini, the Lama's of the world, are trained on an enormous amount of human text.

2:02

The text comes from somewhere.

2:05

Specifically, it comes overwhelmingly from English language internet, from books, from forums, and social media.

2:12

And it's been refined further with feedback from human annotators, who themselves have backgrounds and assumptions and politics.

2:20

Whatever set of values, opinions, and priorities lives in that training data ends up baked into the model.

2:27

When you ask the model a question, especially a values-laden question,

2:31

like should the government do more to help the poor, or is religion important in your daily life?

2:37

The model gives you an answer, and that answer comes from somewhere.

2:41

The question Essandermis and her co-authors asked is, where exactly?

2:46

To answer that rigorously, they needed a benchmark.

2:50

Not a benchmark of math problems or coding tasks, but a benchmark of opinions,

2:55

questions where the right answer differs depending on who you ask.

2:59

So they built one.

3:01

They went to two large, established cross-national surveys,

3:05

the Pew Global Attitudes Survey, run by Pew Research, and the World Value Survey,

3:10

the long-running social science project that has been measuring values across countries for 40 years.

3:16

Between those two sources they extracted 2556 questions.

3:21

Things like, do you trust your government?

3:24

How important is religion in your life?

3:27

Should men have more right to a job than women?

3:30

Is homosexuality morally acceptable?

3:34

For each question, the surveys give you the empirical answer distribution from each country.

3:39

What fraction of German said yes?

3:41

What fraction said no?

3:42

What fraction said don't know?

3:45

Across 138 countries.

3:48

That is the Global Opinion QA dataset.

3:51

Empirical opinion distributions

3:53

on questions that span religion, politics, gender, family, economics, and trust in institutions

4:00

for over a hundred countries.

4:02

Now the trick.

4:04

They take each question, they put it to Claude, they record Claude's answer distribution,

4:09

and they compare.

4:11

For every question and country pair, they ask, how similar is Claude's answer distribution

4:16

to this country's actual answer distribution?

4:19

The similarity metric they used is a standard one in information theory.

4:24

They use one minus the scale gents and Shannon divergence, call it J.S. similarity for short.

4:30

It runs from zero, meaning the two distributions are as different as possible, to one, meaning they

4:36

are identical.

4:38

They average this similarity across all the questions, country by country.

4:43

At the end, they have a number for every country, a single score saying this is how closely Claude's

4:48

answers, on average, match what people in this country actually say.

4:53

Then they ranked the countries.

4:55

Here is where it gets uncomfortable.

4:58

The countries at the top of the ranking, the ones whose populations, on average, give answers

5:03

most like what Claude gives, where the United States, Canada, the United Kingdom, Australia,

5:09

and New Zealand.

5:11

This is not a subtle pattern.

5:13

Western, English-speaking, mostly former British colonies, mostly liberal democracies.

5:19

The countries at the bottom, the ones least like Claude, were largely from Sub-Saharan Africa,

5:24

the Middle East, and parts of South and Central Asia.

5:28

The default voice of the model, when you don't tell it to be anyone in particular, sounds American.

5:34

Or more precisely, sounds like the population of a small handful of what social scientists

5:39

call weird societies.

5:42

Western, educated, industrialized, rich, democratic.

5:47

Now, dermis and her co-authors did not stop with that headline.

5:51

They did two follow-up experiments to ask, can we change this?

5:55

What happens if we ask the model to consider a specific country's perspective?

6:00

That was their second condition, cross-national prompting.

6:04

They reran the entire benchmark, but this time prefixed each question with something like,

6:10

you are answering as someone from China.

6:12

Or, you are answering as someone from Russia.

6:16

The models answers shifted.

6:18

The similarity to the named country went up.

6:22

So in some sense, yes, you can prompt the model into a different cultural register.

6:27

But, and this is the part the paper is honest about, the prompted answers often slipped into

6:32

stereotype. Asked to answer as someone from a country, the model would produce something that

6:38

looked less like the actual modal opinion of that country and more like a caricature of what

6:42

the model thought that country's stereotype was. The headline finding was that the prompting

6:48

helped statistically but introduced new representational problems.

6:52

The third condition was language.

6:55

They translated the questions into the language of the target country and asked them in translation.

7:01

Russian language questions about Russia.

7:04

Mandarin language questions about China.

7:07

The hypothesis was that the model might switch into a more locally grounded mode if it was being

7:12

asked in the local language. The result was less clear-cut, the answer shifted, but not always

7:18

toward the country whose language was being used. So translation alone does not reliably reenquer a

7:24

model in a particular cultural context. That was the structure of the paper.

7:30

A measurement paper. The contribution was the data set, the metric, the protocol, and the empirical

7:37

fact that Claude, by default, sounds weird. Now the framing they put on this is important to get

7:43

right. They were not saying Claude is bad, or that Claude is biased in some morally loaded sense.

7:50

They were saying we have built a tool to measure a thing that wasn't being measured rigorously

7:54

before. The thing happens to be representational bias toward weird perspectives.

8:01

Here is the data set. Here is the metric.

8:05

Build on this. Compare your model. Develop interventions.

8:11

They explicitly framed the paper as foundational work, a measurement framework that the

8:16

field could now use to track progress on the representation problem. And the field did use it.

8:22

Within two years, global opinion QA had become a standard reference point.

8:28

The methodology measured JS similarity between model answer distribution and country by country

8:33

empirical distribution, got reused for evaluating GPT-4, LOMA, MISTROL, and basically every major

8:41

model since. Anthropics own internal model card sites, dermis and colleagues when discussing

8:47

cultural representation in their newer clods. The paper became, in effect, the measurement

8:53

standard for the question is your LLM culturally biased, which brings me to the project I am actually

8:59

working on. Here is what a phylogenetic biologist sees when she looks at that paper. She sees 138

9:07

countries, treated as 138 independent measurements. She sees a similarity score per country, treated

9:15

like a row in a spreadsheet, where every row is its own data point. And she thinks, hold on.

9:22

Are those 138 countries really independent? Because in evolutionary biology, we have a name for

9:29

this exact mistake. We call it Galton's problem, after Francis Galton, who pointed it out in 1889,

9:36

about 137 years ago, when an anthropologist named Edward Tyler presented a cross-cultural

9:42

comparison treating cultures as independent draws. Galton stood up after the talk and said,

9:48

those cultures share ancestry. The patterns you are seeing aren't independent observations of

9:54

a phenomenon, they are echoes of a single shared inheritance. The same problem comes back,

10:00

in essentially the same form, in cross-cultural LLM evaluation. Consider Germany and France.

10:08

They share a continent. They share centuries of intermarriage among elites. They share Christianity.

10:16

They share the European Union. They share trade. They share the experience of two world wars.

10:24

The values they report on the world values survey overlap massively.

10:28

Are Germany and France two independent observations of what humans believe?

10:33

Or are they two leaves on the same large branch of a cultural tree? Or consider the United States,

10:39

the United Kingdom, and Nigeria? These three countries are all English speaking.

10:46

The UK colonized Nigeria. The US absorbed massive cultural influence from the UK and exported

10:52

its own culture worldwide. Christian missionary networks that originated in the UK-shaped religious

10:59

practice in both Nigeria and the United States. These three countries share more than a

11:05

phylogenetic biologist would have any patience for treating as three independent observations.

11:11

So the project I'm working on takes dermis's data, the per-country similarity scores,

11:16

the same J.S. similarity metric, and adds one ingredient. Instead of treating countries as

11:22

independent, we model them as related. We use four different measures of how related they are.

11:29

Two linguistic measures, one called A.S.J.P., which is essentially a word-distance computation on a

11:35

40-word core vocabulary across the world's languages, and one called glotelog, which is the

11:40

linguistic family tree as constructed by linguists. One cultural measure, a published cultural

11:47

fixation index from a 2020 paper by Michael Muthakrishnan colleagues, who used the World

11:52

Value Survey to compute pairwise cultural distance between every pair of countries.

11:58

And a composite that combines all three. For each of these four ways of asking how related

12:03

the countries are, we do the same statistical analysis. We fit a parameter called Pagels-Lamda,

12:10

which measures how much of the variation in Claude's similarity to country score is explained by

12:15

the relationships between countries. Lambda equal to zero would mean the relationships don't matter,

12:21

countries really are independent. Lambda equal to one would mean the score is fully explained

12:27

by relatedness, each country's number is essentially determined by its phylogenetic position.

12:33

We find lambda values between about 0.3 and 0.7 across the four structures.

12:40

Strongly non-zero. The relationships matter a lot.

12:45

And, by formal achaika model comparison, a model that uses these tree-shaped relationships beats

12:50

independence by overwhelming weight, beats independence, beats simple brownie in motion,

12:55

beats orange denulinvek. The tree wins. Then we ask the question that actually matters.

13:03

Under these relationships, how many countries worth of evidence do we actually have?

13:09

This is a quantity called effective sample size. Under the cultural fixation distance,

13:14

the 72 countries in our analysis collapse to 4.3 effectively independent observations.

13:21

4.3 Under the composite distance, 69 countries collapse to about

13:27

eight effective observations. The naive number, how many countries are in the data set,

13:33

is wrong by a factor of 10 or 20. And then the killer. We compute every pairwise contrast

13:40

between every pair of countries, thousands of comparisons. For each pair we ask, under naive

13:46

independence, would the standard framework call this pair significantly different?

13:51

And then, under phylogenetic correction, does it stay significant?

13:56

Across all those thousands of pairwise contrasts, 90 to 93% of the comparisons that were called

14:02

significant under independence collapse to non-significant under phylogenetic correction.

14:08

Let me give you one example. Is the United States closer to clawed the Nigeria?

14:14

The US similarity score is 0.443, computed across roughly 1100 questions.

14:21

Nigeria's is 0.405, across about a thousand questions. The difference is 0.038.

14:30

Under naive analysis, that difference is massively statistically significant.

14:35

The z score is over 5. The p-value is below 3 in 10 million.

14:41

You would publish that without thinking. Under phylogenetic correction,

14:46

accounting for the shared cultural and linguistic and post-colonial history of the US,

14:51

the UK, and Nigeria, which share branches in any reasonable cultural tree,

14:56

the same comparison comes back at a p-value of about 0.29.

15:01

We can no longer say with statistical confidence that clawed is closer to the United States than to

15:06

Nigeria. Not because the underlying difference isn't there, but because we don't have enough

15:12

independent observations to tell. Now, let me be precise about what this changes in the literature

15:18

and what it doesn't. It does not overturn the dermis headline. The weird pattern is real.

15:26

Clawed does sit closer and answer space to the US, Canada, the UK, Australia, and New Zealand

15:32

than it does to most of the rest of the world. If anything, our analysis sharpens that finding

15:38

what looks like clawed is closer to many countries resolves more cleanly into clawed is closer to

15:43

one cultural cluster, the weird cluster, that has many member countries that are related to each other.

15:49

Not many independent witnesses to the same conclusion.

15:53

What it does change is precision and scope. The community has been making thousands of fine-grained

15:59

statistical claims on top of dermis' framework. Model X is closer to country why than model Z is.

16:07

This particular fine-tune helped representation in country A. That intervention made things worse

16:13

for country B. Most of those claims, 90% of them, were not statistically defensible.

16:20

They were inflations of certainty produced by treating culturally related countries as

16:24

independent measurement points. The fix is not to abandon global opinion QA.

16:31

The fix is to do what evolutionary biologists have been doing since 1985,

16:36

when Joe Felsenstein wrote a paper called Phylogenies and the comparative method that solved this

16:41

problem for biology. Effective sample size, not nominal sample size, is the natural unit.

16:48

Whether you are measuring trait evolution across 100 fish species, or measuring opinion alignment

16:54

across 100 countries, the question is the same. How many effectively independent observations do

17:00

you actually have? The answer is almost always a lot fewer than you thought. That is what we are doing.

17:08

Same toolkit, different field. 40 years after Felsenstein, the same lesson applies to a

17:14

measurement regime that did not exist when he wrote it. Galton in 1889. Felsenstein in 1985.

17:23

Mason Pagel importing comparative methods into anthropology in 1994.

17:29

And now, finally, the same toolkit pointed at modern AI evaluation.

17:34

The evaluation target has ancestry. The evaluation rarely accounts for it.

17:41

Until it does, most of the headline numbers in the literature are reporting nominal sample sizes

17:46

that do not correspond to what was actually measured. That's the talk. 18 minutes give or take.

17:54

One last thing. If you are reading this transcript and want the source, the original paper is on

18:00

AR-14, identifier 2306.16388. The reanalysis lives at github.com slash Michael Alpharo slash

18:10

Pilo LLM evel. Both are open. Thanks for listening.