SceneBench:3D场景视觉语言理解的分层基准
原标题:SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes
AI 摘要
研究者提出 SceneBench,一个面向 3D 场景视觉语言理解的分层基准,包含 966 个用 Gaussian Splatting 重建的逼真 3D 场景,并通过人机协同流程(约 1500 人工小时)标注了超过 18.3 万个带文本描述和 3D 边界框的层级语义节点。基准定义了存在性问答、空间智能问答和需跨语义层级多步推理的 Grounded QRA 三类任务。实验显示当前最先进视觉语言模型在基础识别上表现较好(检测最高约 85%),但在分层与组合推理上明显下降(计数低至约 60%),暴露出既有基准未覆盖的局限。
正文节选
SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes Abstract Vision-language models have achieved strong performance on 2D image understanding, but their ability to reason about 3D environments remains limited. Progress in spatial intelligence is hindered by limitations in existing benchmarks. First, many 3D datasets rely on point clouds, which capture geometry but discard rich visual appearance such as texture, text, and materials. Second, annotations typically t