Skip to content
Breaking:

LLM Benchmark Exposes Deep Demographic Bias in AI Expert Recommendations

A study across 22 leading language models reveals persistent trade-offs between factual accuracy and social representation when AI systems suggest technical experts.

By The Company Wire4 min read
Share
Complexity Science Hub — LLM Benchmark Exposes Deep Demographic Bias in AI Expert Recommendations
Complexity Science Hub — LLM Benchmark Exposes Deep Demographic Bias in AI Expert Recommendations. Photo: TechXplore.

Artificial intelligence systems built on large language models routinely perpetuate and amplify social biases when tasked with recommending professional experts, according to research first reported by TechXplore. Researchers evaluating 22 widely used AI models—spanning proprietary and open-weight architectures such as GPT, Gemini, Llama, DeepSeek, Qwen, Grok, Mistral, and Gemma—found that while chatbots frequently cite real professionals, their outputs heavily favor male, senior, highly cited scholars based in the United States.

To measure these tendencies, a research team led by Complexity Science Hub fellow Lisette Espín-Noboa developed LLMScholarBench, an open evaluation framework that measures nine technical and social metrics including factuality, diversity, parity, validity, and consistency. The researchers audited model outputs against a ground-truth dataset of more than 450,000 scientists who published in American Physical Society (APS) journals between 1893 and 2020.

Initial tests revealed that models mismatched scientists with their correct research subfields roughly 40% of the time on average. Demographic disparities were equally pronounced. Though women represent between 14% and 32% of published authors in the APS historical record depending on the discipline and era, LLMs routinely cited lower proportions of female experts or omitted them entirely. Similarly, while Asian researchers constitute the largest demographic group in the APS database, models consistently overrepresented white scholars while Black and Latino experts were often left out.

Across the 22 evaluated systems, model performance varied across performance dimensions. Technical factuality scores ranged from 0.63 to 0.82, meaning 63% to 82% of recommended experts corresponded to verified records in the baseline database. DeepSeek and Gemini achieved some of the highest factuality ratings, whereas Qwen, Llama, and Mistral scored lower. Diversity scores spanned from 0.44 to 0.69, led by DeepSeek, while Google's Gemma models led in demographic parity, scoring between 54% and 60%.

The research, presented in part at the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2026) and documented in arXiv preprints, evaluated user interventions like prompt engineering and Retrieval-Augmented Generation (RAG). The team discovered a fundamental trade-off: prompt adjustments successfully steered demographic outputs, and RAG improved factual accuracy, but combining interventions failed to resolve both issues at once. Testing newer models like Gemini 2.5 Pro and Flash with web search capabilities increased accuracy but reduced representation diversity.

Testing across six academic disciplines further indicated that geographic context within prompts directly altered expert recommendations, whereas prompt language and user personas—such as acting as a recruiter versus a student—had no notable impact. To enable public auditing, CSH data visualization expert Yi Zhe Ang and the team created an interactive platform called "Whose name comes up," allowing users to compare model family metrics.

Espín-Noboa highlighted that the identified biases stem from systemic training data issues rather than isolated model defects, noting that similar LLM recommendation workflows are increasingly deployed in medicine, law, and corporate hiring. To help address the problem, the research team aims to build a broader, more comprehensive scientific repository to inform future model training.

Sources

  1. TechXplore

Company: Complexity Science Hub

Written by

The Company Wire

Newsroom · San Francisco

Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.