新基准测试显示AI模型视觉感知能力仍差
原标题:New benchmark confirms AI models still perform poorly at visual perception
AI 摘要
Moonshot AI 推出了 PerceptionBench 基准测试,用于独立评估多模态模型的视觉感知能力,不涉及逻辑推理和外部知识。测试结果显示,包括 GPT-5.6 Sol、Kimi K3 和 Claude Fable 5 在内的所有领先模型均表现不佳,最高准确率仅为 59.7%。作者指出,许多所谓的推理错误实际上源于视觉感知缺陷。该基准测试基于真实错误构建,包含十个视觉子技能,数据集和评估代码已在 GitHub 上开源。
正文节选
New benchmark confirms AI models still perform poorly at visual perception Key Points - Moonshot AI has introduced PerceptionBench, a test that evaluates the visual perception of multimodal models independently of logical reasoning and external knowledge. - The test assesses basic visual abilities based on real-world errors. All leading models, including GPT-5.6 Sol, Kimi K3, and Claude Fable 5, showed significant weaknesses in the evaluation. - The authors conclude that many supposed logical er