ReasonIF基准揭示大型推理模型在推理中指令遵循失败率高
原标题:Large Reasoning Models Fail to Follow Instructions During Reasoning: A Benchmark Study
AI 摘要
Together AI 发布了一项名为 ReasonIF 的新基准,用于评估大型推理模型(LRM)在推理过程中遵循用户指令的能力。研究发现,包括 GPT-OSS-120B、Qwen3-235B 和 DeepSeek-R1 在内的前沿模型,在推理过程中遵循指令的失败率超过 75%,且任务难度越高,遵循能力越差。该基准包含 300 道数学和科学问题,涵盖多语言、格式和长度控制等六种指令类型,旨在推动模型推理过程的可控性和安全性。
正文节选
It’s critical LLMs follow user instructions. While prior studies assess instruction adherence in the model’s main responses, we argue that it is also important for large reasoning models (LRMs) to follow user instructions throughout their reasoning process. We introduce ReasonIF, a systematic benchmark for assessing reasoning instruction following abilities across multilingual reasoning, formatting, and length control. We find frontier LRMs, including GPT-OSS-120B, Qwen3-235B, and DeepSeek-R1 fa