OpenAI 称 Astra 为最危险模型,安全监控面临挑战
原标题:OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder
AI 摘要
OpenAI 将其即将推出的 Astra 模型评为首个具有“严重”网络攻击能力的系统,同时声称这是其构建的最安全模型。内部测试显示 Astra 在漏洞利用基准上表现优异,并发现了两个零日漏洞。为应对风险,OpenAI 计划推出新的安全措施,包括监控模型的思维链,但该技术可能因模型内部推理的不可见性而存在局限。
正文节选
OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder OpenAI is rating its upcoming Astra model as the first system with "critical" cyber capabilities, while promising it's also the safest model the company has built. But a report on Astra's architecture raises questions. It's an unusual way to announce a product: OpenAI says its upcoming Astra model is so dangerous that it hits the highest risk tier for cybersecurity in the company's own Preparedness Fra