返回全部动态
Anthropic研究展示自我改进AI系统提升对齐性能
原标题:An Anthropic researcher just gave us a peek at self-improving AI
AI 摘要
Anthropic研究员Chen Yueh-Han领导的一项新研究展示了AI系统如何通过自动化研究流程改善模型的对齐性能。该系统在10个对齐基准上均提升了性能,且未降低整体表现,成本仅为每小时4美元,远低于人类研究员的150美元。研究为递归自我改进提供了早期证据,但也指出其依赖基准准确性和文献维护等局限性。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
Training AI models with other AI models has become a very popular goal for neolabs — and now, a researcher in Anthropic’s fellows program has given us an early look at what it might look like in practice. On Friday, Anthropic published a new paper titled “Automated Researchers Can Reliably Mitigate Alignment Failures,” detailing how AI systems could reliably improve a model’s performance on a set of alignment benchmarks. When given 10 benchmarks for specific misaligned behaviors, the automated s
发布时间:2026-08-29 03:30
抓取时间:2026-08-29 04:24
来源机构:TechCrunch