返回全部动态
立场:对齐社区无意中构建了审查工具包
原标题:Position: The Alignment Community is Unintentionally Building a Censor's Toolkit
AI 摘要
这篇立场论文认为,现代AI对齐方法虽然旨在防止有害输出,但具有双重用途,可能被恶意行为者滥用于审查和操纵。作者通过将当前对齐技术映射到可能的滥用案例,指出追求“完美对齐”的模型无意中为信息控制提供了工具,并呼吁社区讨论这一风险,提出缓解策略。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
Position: The Alignment Community is Unintentionally Building a Censor’s Toolkit Abstract This position paper argues that modern AI alignment methods – originally designed to prevent harmful output – are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibility and actual cases of misuse, we show that the quest for a “perfectly aligned” model inadvertently also provides malicious actors with a
发布时间:2026-08-15 12:00
抓取时间:2026-08-15 12:00
来源机构:arXiv