Can LLMs develop social bias? Not like this

Another case of "I don't think that experiment means what you think it does"

Yet another paper showing bias in AI.

In the experiment, they task an LLM to select people from different population groups for four jobs. The only information the LLM has is which village the people come from, plus their performance once they’re in the job. The experiment shows that LLMs consistently bias people from one village over another. The twist is that all people are equally competent.

This proves that large language models develop novel social biases.

No. No it doesn’t


I love this kind of simulations. I used to write them when I was reading for my PhD.

You create a controlled environment of agents, stipulate some causal mechanism, and show how our inferences about said mechanism are all wrong.

Here, the authors stipulated that all candidates had equal competence, and watched an LLM that only has access to limited information repeatedly favour one type of background over another.

The result is certainly terrible for anyone thinking about letting Claude run HR but one thing it doesn’t show is that LLMs develop novel social biases.

Not unless you think an if statement can develop social biases too.


Here’s the rub. You can recreate the exact same results without AI. All you need is an if and perhaps an else — although the latter isn’t even necessary.

hire candidate if candidate.village == last_hire.village && last_hire.kicked_ass?

What it shows is an information cascade.

Whenever agents or people don’t have first order knowledge, they must rely on second order knowledge to fill the gaps. Unfortunately, using that second order knowledge can create unexpected and unwanted results.

Think of it this way. Imagine you’re in a new city and you’re looking to go to a restaurant. You have no idea about the quality of the individual restaurants. You stumble upon two restaurants — one that’s empty and one that looks pretty busy. Which do you choose?

The answer is the busy one. You may not know anything about either restaurant but you do know that good restaurants tend to be busy.

Now step back to earlier in the evening and imagine that both restaurants are empty. Also imagine both restaurants are equally good. You don’t know that though — you lack this first-order evidence. So you toss a coin and pick the first.

You are now second order evidence. Another person comes along who has no idea about the quality of the restaurants but sees you dining at one. They infer that the busier restaurant is probably better so they select the one you’re at.

Now they are second-order evidence.

This process continues. The more that people pick restaurants based on popularity the more popular that restaurant becomes. This is an information cascade. Both restaurants can be as good as each other, the popular restaurant could even be worse, but second order information plus initial conditions can lead to runaway effects.

This is the exact experimental design of the paper.

So no, this paper doesn’t show that LLMs developed bias. LLMs are full of bias, but not for this reason.

What this paper really shows is the effects are not always caused by what we think they are.