返回全部动态

AX-Ray 诊断工具揭示两个公开模型的因果泄漏缺陷

原标题:AX-Ray, Finding Causal-Leakage Defects in Two General-Purpose Public Models

Hugging Face Blog一手来源研究质量 83

AI 摘要

VIDRAFT 推出了 AX-Ray,一个用于 AI 和 AX 部署的安全诊断层,基于 FINAL-Bench Diagnostics 构建。AX-Ray 在公开诊断中识别出 Zyphra/Zamba2-1.2B 和 nvidia/Nemotron-H-8B-Base-8K 两个模型存在因果泄漏缺陷,即前缀行为依赖未来 token,属于结构性正确性失败。该缺陷可能影响前缀不变性、隐藏状态正确性、缓存可靠性等,AX-Ray 将其视为部署阻断问题,并通过 F-gate 处理。AX-Ray 围绕 MODEL-SCAN、AX-SCAN 和 AGENT-SCAN 三个诊断轴,包含 117 个公开诊断条目,强调高能力不等于部署安全。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

AI models should no longer be evaluated only by how well they answer benchmark questions. Capability matters, but deployment safety depends on a wider set of properties: causal correctness, serving consistency, robustness under adversarial or long-context conditions, data integrity, infrastructure security, regulatory readiness, and agentic risk. VIDRAFT built AX-Ray to address that gap. AX-Ray is a safety-diagnostics layer for AI and AX deployment, powered by FINAL-Bench Diagnostics. It evaluat


发布时间:—
抓取时间:2026-08-14 10:04
来源机构:Hugging Face
阅读原文huggingface.co