GEB-Bench:多声部讲述的抽象结构基准
原标题:GEB-Bench: Abstract Structures Told in Many Voices
AI 摘要
GEB-Bench 是一个新基准,用于评估模型识别抽象结构(如自指、怪圈、莫比乌斯扭转)的能力,这些结构通过自然场景、民间故事、数学定理和程序骨架等多种形式呈现。对十二个开源和专有模型的评估发现,模型在单一形式内识别结构的能力远强于跨形式映射,且所有模型都存在这一差距,只有前沿模型能缩小差距。错误模式与设计的正式几何结构高度一致,表面复杂度对所有模型都有影响,但更大容量只能提供更多余量而非完全免疫。该基准完全生成式,并已发布其流水线。
正文节选
Computer Science > Computer Vision and Pattern Recognition Title:GEB-Bench: Abstract Structures Told in Many Voices View PDF HTML (experimental) Abstract:Can a model look at a river delta and a lightning bolt and see that they share a structure? We introduce GEB-Bench, a benchmark whose unit is an abstract structural motif--self-reference, a strange loop, a Mobius twist--in the spirit of Godel, Escher, Bach. Each motif is told in several voices: a natural scene whose composition is t