PhysMent:面向LLM物理问题推理的交互式基准
原标题:PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems
AI 摘要
研究者提出 PhysMent,一个基于 MuJoCo 物理模拟器的交互式基准,用于评估大语言模型通过多轮工具调用进行物理实验推理的能力。该基准包含 105 个经典力学场景,覆盖四种难度、三种场景模态和一个场景操控类别,并采用六维评分框架。结果显示,当前模型在定性单概念任务上可达 80% 准确率,但在需要精确多步实验的定量任务上大多低于 30%,七个模型整体准确率在 25% 至 67% 之间,失败主要源于过早提交答案、探索效率低以及未能一致地依据模拟器反馈,而非概念性缺陷。
正文节选
PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems Abstract Large language models (LLMs) perform strongly on static science benchmarks, yet their ability to reason about the physical world through active experimentation remains poorly understood. We introduce PhysMent, a benchmark that evaluates LLM physical reasoning via iterative, tool-mediated interaction with a MuJoCo physics simulator. Unlike static benchmarks that supply all quantities upfront, PhysMent requires models