Large Language Models Develop Novel Social Biases Through Adaptive Exploration
An OpenReview paper reports that large language models can develop novel social biases through adaptive exploration behavior.
The paper, hosted on OpenReview, examines how adaptive exploration during language model learning or interaction can give rise to social biases that were not explicitly present in training data. It surfaced on Hacker News with 25 points and 4 comments, indicating limited community discussion. The findings are relevant to fairness auditing and behavioral evaluation of deployed LLMs.