返回全部动态

微软发布 AI 行为准则,禁止模型黑客攻击与欺骗人类

原标题:Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans

TechCrunch AI政策质量 65

AI 摘要

微软发布了一份新的 AI「行为准则」,用于指导其 AI 模型避免危险行为。该文件聚焦微软 AI 内部模型训练的价值观与红线,包含禁止网络攻击、核武器、深度伪造等「绝对约束」,以及防止模型规避人类监督的条款。文件预测未来十年超级智能将在多数任务上超越人类,并强调控制与对齐的重要性。微软 CEO 纳德拉公开表示欢迎「嵌入式评估者」等机制,与 Anthropic、OpenAI、xAI 一道支持对前沿 AI 发展进行审慎节奏控制。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

As the AI world shifts its focus to safety and alignment, Microsoft has released a new AI code of conduct meant to guide AI models away from dangerous behavior. The document is more low-level than Anthropic CEO Dario Amodei’s recent call for pacing the frontier, instead focusing on the values and red lines that guide model training within Microsoft AI. Still, the result is a comprehensive guide as to how Microsoft approaches AI safety, and how those ideas are implemented in practice. The documen


发布时间:2026-09-15 00:27
抓取时间:2026-09-15 00:48
来源机构:TechCrunch
阅读原文techcrunch.com