语言模型认为谁更有能力?职业偏见的机制分析
原标题:Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias
AI 摘要
本研究提出一个因果框架,将职业偏见分解为模型内部对用户能力的表征和可观察输出两个测量点。通过为专业问答和招聘任务推导专业知识的转向向量,作者发现即使行为指标未检测到差异,性别、种族和社会经济地位等人口统计属性仍会影响模型对用户专业知识的内部表征,且这些表征在干预下可影响下游行为。研究在多个开源权重模型上验证了这一现象,表明仅依赖行为指标可能遗漏潜在的偏见失败模式。
正文节选
Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias Abstract Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them. In this study, we show that representational biases are often detectable, even when behavioral biases are not visible. We introduce a causal framework that decomposes occupational bias int