返回全部动态

LoopArena:评估模型作为循环工程运行时控制器的基准

原标题:Paper page - LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering

Hugging Face Daily Papers一手来源研究质量 82

AI 摘要

LoopArena 是一个用于评估模型作为循环工程中运行时控制器能力的新基准。它通过三种设置测试控制器模型指导独立编码代理完成长任务的表现,发现最佳严格成功率为 24.69%,同时推理成本平均降低 64.4%。该基准的代码和数据已开源。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Abstract LoopArena benchmarks how well a controller model guides a separate coding agent through long tasks, revealing low strict success rates and significant cost reductions. Loop Engineering is emerging as a practice for organizing development work around coding agents. Instead of writing each prompt by hand, practitioners design loops that monitor progress, assign work, run checks, and decide what the agent should do


发布时间:—
抓取时间:2026-08-31 19:38
来源机构:Hugging Face
阅读原文huggingface.co