返回全部动态

安全为谁?拒绝话题中的正确子集而非整个话题

原标题:Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

Hugging Face Blog一手来源研究质量 88

AI 摘要

Hugging Face Blog 介绍了一篇新论文《Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal》,研究如何让模型只拒绝某个话题中与部署政策不兼容的子集,而非整个话题。作者以政治说服为测试场景,指出标准自生成训练流程存在覆盖缺口、误拒副作用和边界度量缺失三个问题,并提出升级重试、分布内良性数据和边界配对等修复方法。实验显示,在 Qwen3-8B 上政治拒绝率从 9.47% 升至 84.75%,但 XSTest 过度拒绝率也从 2.00% 升至 74.00%,说明必须同时报告安全与过度拒绝两个维度。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Real deployments rarely fit the topic-level picture. The same base model may be adapted for a general assistant, an educational product, an enterprise system, or a public-sector service, and each setting needs different boundaries within the same topic. A civics tutor and a public-sector assistant can share a model yet require opposite behaviour on politics: both should answer factual questions about an election, but only one may need to refuse a request to write targeted political manipulation.


发布时间:—
抓取时间:2026-09-21 22:00
来源机构:Hugging Face
阅读原文huggingface.co