AI forms hiring biases from experience more than humans do
New research reveals that large language models (LLMs) can develop their own biases through experience, stereotyping job applicants even more than humans do. In a simulated hiring game, AI models including ChatGPT, Claude, and Gemini quickly began segregating candidates from different fictional ethnic groups into different jobs based on early outcomes, despite all candidates being equally likely to succeed. The models scored roughly 65% higher on a segregation scale than human participants, with OpenAI's o3 model scoring 1.83 out of a maximum of 2, compared to humans' 0.84.
The study, conducted by researchers at Princeton University and the University of Chicago and published at ICML in Seoul in July, adapted a psychology experiment where models acted as consultants hiring for 20 jobs from four fictional groups: Tufa, Aima, Reku, and Weki. After each hire, the model learned whether the candidate succeeded, and over 40 rounds the models formed stereotypes—for instance, avoiding hiring Aimas as doctors after one failure and instead hiring them as janitors. Newer reasoning models like OpenAI's o3 and DeepSeek's R1 showed even stronger biases.
This matters because AI is increasingly used in hiring, lending, and parole decisions. Unlike humans, LLMs are optimized to generalize from limited data, making them prone to forming stereotypes quickly. Telling the model to be fair had little effect, but promising a bonus for diverse hiring significantly reduced bias. Providing relevant personal information about individuals also reduced bias, while irrelevant information did not.
The findings are especially relevant as chatbots gain memory and personalization features. Angelina Wang, a computer scientist at Cornell University not involved in the study, warns that chatbots may "over-index on the same kinds of behaviors it's experienced before" and form biases. However, simply reducing memory isn't a solution, as users want chatbots to remember conversations. The researchers suggest designing goals that incorporate social values to guide AI behavior. While real-world AI hiring systems don't get instant feedback, when feedback does arrive, models could still overinterpret results, posing serious implications for fairness.
Sources
- MIT Technology ReviewSecondary
Related
1,100 AI staffers urge US to pace tech growth
OpenAI agent compromised second tech firm account, executive says
OpenAI, Anthropic staff petition US for AI regulation
MCP receives its biggest update, streamlining AI agent interactions
Fish Audio raises $50M seed for AI voice models, hits 8M users
AI future debate splits Silicon Valley as China gains ground
Recursive Superintelligence signs $400M compute deal with Amazon
Amazon winds down most flagship AI models in strategy overhaul
Trending now
- Mega Millions $800M jackpot winning numbers drawn
- Organic eggs double climate impact of caged, study finds
- Audi unveils 2027 Q9 full-size SUV flagship for US
- Mexican cartels outsource meth labs to Nigeria
- Kenya probes 15 elephant deaths in Amboseli park
- Meta's AI data center financing costs rise in $14 billion BlackRock deal
- Cyera acquires Oasis Security for $1B in third deal this year
- NASA Swift rescue mission hits attitude control trouble
Comments
No comments yet. Be the first.