返回全部动态

The Tasteful Agent:长周期任务中智能体品味的度量与提升

原标题:Paper page - The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

Hugging Face Daily Papers一手来源研究质量 82

AI 摘要

该论文提出「品味」(taste)概念,指 LLM 智能体在长周期任务中于决策分叉点选择更优方向的能力,并构建了 Taste-Bench 基准。Taste-Bench 从 SWE-bench Pro 和 METR AI R&D 轨迹中自动挖掘出 502 个决策分叉,无需人工标注。评测显示 14 个前沿模型中最优者仅答对 59.7%,且决定证据出现越晚准确率越低,增大推理预算也无帮助。作者进一步证明品味可训练:将事后教师模型的判断蒸馏到学生模型后,Qwen3.6-27B 在未见任务上从 30.0% 提升至 47.9%,并使 SWE-bench Pro 成功率从 14.6% 升至 33.7%。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks Abstract LLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run. Making these decisions well is becoming a key capability for both engineering and research agents. We refer to the ability to make good long-horizon decisions as the taste of an agent. While existing benchmarks mea


发布时间:—
抓取时间:2026-09-23 20:27
来源机构:Hugging Face
阅读原文huggingface.co