返回全部动态

NeoHorse-1:通过路由框架的智能体后训练实现递归自我改进

原标题:Paper page - NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

Hugging Face Daily Papers一手来源研究质量 83

AI 摘要

NeoHorse-1 通过智能路由、结构化反馈循环和基于课程的知识蒸馏进行智能体后训练,以提升模型在智能体基准上的能力。该系统结合异构模型池与智能路由,将交互记录转化为训练示例,并采用三阶段课程和路由引导的在线蒸馏。在11个基准上,4B模型宏平均从58.94提升至64.87,9B模型从65.60提升至69.04,缩小了4B与9B模型间的差距。该工作为递归自我改进提供了初步原型。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Abstract NeoHorse-1 uses agentic post-training with intelligent routing, structured feedback loops, and curriculum-based distillation to improve model capabilities across agent benchmarks. Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-


发布时间:—
抓取时间:2026-09-09 11:53
来源机构:Hugging Face
阅读原文huggingface.co