AX-Ray 诊断工具揭示两个公开模型的因果泄漏缺陷
原标题:AX-Ray, Finding Causal-Leakage Defects in Two General-Purpose Public Models
AI 摘要
VIDRAFT 推出了 AX-Ray,一个用于 AI 和 AX 部署的安全诊断层,基于 FINAL-Bench Diagnostics 构建。AX-Ray 在公开诊断中识别出 Zyphra/Zamba2-1.2B 和 nvidia/Nemotron-H-8B-Base-8K 两个模型存在因果泄漏缺陷,即前缀行为依赖未来 token,属于结构性正确性失败。该缺陷可能影响前缀不变性、隐藏状态正确性、缓存可靠性等,AX-Ray 将其视为部署阻断问题,并通过 F-gate 处理。AX-Ray 围绕 MODEL-SCAN、AX-SCAN 和 AGENT-SCAN 三个诊断轴,包含 117 个公开诊断条目,强调高能力不等于部署安全。
正文节选
AI models should no longer be evaluated only by how well they answer benchmark questions. Capability matters, but deployment safety depends on a wider set of properties: causal correctness, serving consistency, robustness under adversarial or long-context conditions, data integrity, infrastructure security, regulatory readiness, and agentic risk. VIDRAFT built AX-Ray to address that gap. AX-Ray is a safety-diagnostics layer for AI and AX deployment, powered by FINAL-Bench Diagnostics. It evaluat