返回全部动态

DSAgentBench:评估智能体在真实环境中自动化数据科学工作流的能力

原标题:DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

Hugging Face Daily Papers一手来源研究质量 82

AI 摘要

DSAgentBench 是首个在真实计算机环境中评估智能体自动化端到端数据科学工作流能力的基准,包含 275 个覆盖数据科学生命周期的任务。实验表明,最强智能体 Claude-4.6-Sonnet 的任务成功率仅为 56.70%,而所有开源智能体成功率均低于 1%,揭示了当前智能体系统与真实数据科学工作流之间的巨大能力差距。该基准已开源发布。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? Abstract DSAgentBench evaluates autonomous agents on complete, multi-tool data-science workflows in real computing environments and reveals major performance gaps. Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases with


发布时间:—
抓取时间:2026-08-12 12:13
来源机构:Hugging Face
阅读原文huggingface.co