返回全部动态

PhysMent:面向LLM物理问题推理的交互式基准

原标题:PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems

arXiv cs.CL一手来源研究质量 88

AI 摘要

研究者提出 PhysMent,一个基于 MuJoCo 物理模拟器的交互式基准,用于评估大语言模型通过多轮工具调用进行物理实验推理的能力。该基准包含 105 个经典力学场景,覆盖四种难度、三种场景模态和一个场景操控类别,并采用六维评分框架。结果显示,当前模型在定性单概念任务上可达 80% 准确率,但在需要精确多步实验的定量任务上大多低于 30%,七个模型整体准确率在 25% 至 67% 之间,失败主要源于过早提交答案、探索效率低以及未能一致地依据模拟器反馈,而非概念性缺陷。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

PhysMent: An Interactive Approach For LLM Reasoning In Physics Problems Abstract Large language models (LLMs) perform strongly on static science benchmarks, yet their ability to reason about the physical world through active experimentation remains poorly understood. We introduce PhysMent, a benchmark that evaluates LLM physical reasoning via iterative, tool-mediated interaction with a MuJoCo physics simulator. Unlike static benchmarks that supply all quantities upfront, PhysMent requires models


发布时间:2026-09-15 12:00
抓取时间:2026-09-15 12:11
来源机构:arXiv
阅读原文arxiv.org