返回全部动态

分片可防止LLM监督失败与对抗性利用

原标题:Sharding Prevents LLM Oversight Failures and Adversarial Exploitation

arXiv cs.LG一手来源研究质量 85

AI 摘要

arXiv 上的一项研究指出,给 LLM 裁判更多计算资源并不一定能使其检查更多要求,当单次调用需返回多个裁决时,部分决策的证据基础会变弱。研究发现,将需求分组并分别调用(即分片)能显著提高与专家的一致性,且分片后的较弱裁判可胜过能力更强但整体评估的裁判。此外,分片还能抵御利用过载的对抗性攻击,但对针对每个标准的单独说服攻击无效,需结合辩论式对抗。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Computer Science > Machine Learning Title:Sharding Prevents LLM Oversight Failures and Adversarial Exploitation View PDF HTML (experimental) Abstract:Giving an LLM judge more compute does not necessarily make it check more requirements. When one call must return many verdicts, some decisions become weakly grounded in the evidence, even when that call receives the same token or tool budget as a panel of separate calls. Across expert-graded research replications, legal work, and clinic


发布时间:2026-08-10 12:00
抓取时间:2026-08-10 12:04
来源机构:arXiv
阅读原文arxiv.org