AI is more likely than humans to form biases when hiring


The models were even more likely to stereotype people by demographic group than the human participants in the original study. On the study’s segregation scale, where 2 means every group has been completely confined to its own job niche, human participants scored 0.84. The models scored roughly 65% higher, with OpenAI’s reasoning model o3 scoring 1.83, close to the maximum possible.

That’s because LLMs “really are eager to create generalizations from limited data,” says Ryan Liu, a PhD student at Princeton University and a coauthor of the study, which was published in a paper at ICML in Seoul in July. “That’s literally a lot of what they’re optimized for.” Every decision-maker, human or machine, faces a trade-off between sticking with what worked before and trying something new that might work better—a phenomenon psychologists call the “exploration-exploitation dilemma.” It’s like choosing between a new restaurant and your reliable favorite. 

Because LLMs are trained on math, coding, and science problems—tasks that reward generalizing from just a few examples—they can settle on a hunch too early. And the same instinct that helps LLMs crack logic puzzles also makes them quick to stereotype. In the experiment, newer models with higher reasoning capabilities, such as OpenAI’s o3 and DeepSeek’s R1, showed even stronger biases. When LLMs rush to generalize in social settings, “that’s when things tend to go wrong,” says Liu. OpenAI and Anthropic did not respond to requests for comment.

The finding is especially relevant now that chatbots are gaining improved memory and personalization features, says Angelina Wang, a computer scientist at Cornell University who did not work on the study. When a chatbot draws on its previous conversation history, it can “over-index on the same kinds of behaviors it’s experienced before” and form biases, she says. Simply having chatbots remember less isn’t a fix, though, because users want chatbots to remember what they say. “We still are trying to figure out just the right amount that isn’t too much or too little,” says Wang.

Telling the model to be fair didn’t change its behavior much. “Either it can’t put these values into action or that process is being submerged under the tendency to try to optimize for the goal of getting the most correct hires,” says Liu. But promising the models an additional bonus for diverse hiring made them far less biased. The trick, then, is to design goals that “incorporate desirable social values in order to make the large language model act in socially desirable ways,” says Liu.



Source link

  • Related Posts

    The cost of GPUs goes far beyond AI data centers

    HowHow does AI make you feel? Are you excited to “vibe-code” your smart home? Or anxious about all the added pollution and billions of gallons of water used by data…

    The Army Is Burning Through Its AI Tokens

    A little over a month after the Department of Defense (DOD) bragged that nearly half of its 3.5 million employees were using AI at work, members of the Army’s Combat…

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    Residential school counselling pause disrupts care for N.W.T. clients, leaves providers in limbo

    Residential school counselling pause disrupts care for N.W.T. clients, leaves providers in limbo

    The cost of GPUs goes far beyond AI data centers

    The cost of GPUs goes far beyond AI data centers

    The High Cost of Just Cause

    The High Cost of Just Cause

    Man flipped by bison speaks out: “They move faster than you could ever imagine”

    Man flipped by bison speaks out: “They move faster than you could ever imagine”

    Matt Walker: Sussex appoint new head coach to succeed Paul Farbrace

    Matt Walker: Sussex appoint new head coach to succeed Paul Farbrace

    New UK Prime Minister Burnham Promises Hope and Change. The Hurdles Are High.

    New UK Prime Minister Burnham Promises Hope and Change. The Hurdles Are High.