LLM Benchmark Exposes Deep Demographic Bias in AI Expert Recommendations
A new evaluation benchmark testing 22 large language models found widespread technical errors and demographic biases when AI systems generate expert recommendations across scientific disciplines.
