FairLens:评估视觉语言模型在高风险决策中的公平性
原标题:FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making
AI 摘要
arXiv 论文提出 FairLens,一个用于评估视觉语言模型在高风险决策(招聘、法律、医疗)中公平性和有效性的基准框架。该框架结合真实人脸图像与封闭/开放问题,从四个视角评估八个 VLM,发现主要失败模式是无根据推断而非不平等对待,且法律和医疗领域问题最严重。研究强调仅靠差异指标不足,需结合合理性评估,并指出自由文本偏见与多项选择准确性关联松散。
正文节选
FairLens: Benchmarking Fairness in Vision–Language Models for High-Stakes Decision-Making Abstract Vision–language models (VLMs) are increasingly used to make decisions from visual inputs. We introduce FairLens, a benchmark and evaluation framework for measuring both the fairness and the validity of VLM responses in three high-stakes domains: hiring, legal, and healthcare. FairLens pairs real face images spanning gender, race, and age groups with closed- and open-ended questions, giving more tha