广义智能体迭代:统一迭代策略改进与递归自我改进的形式化框架
原标题:Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
AI 摘要
该论文提出广义智能体迭代(GAI)形式化框架,将迭代策略改进与递归自我改进(RSI)统一描述为同一学习范式的两个特例。GAI将智能体定义为系统中可修改组件的配置,并把学习过程建模为智能体评估与智能体改进的循环,通过两个关键维度区分不同实例:改进机制是否属于智能体本身,以及衡量标准是否锚定在智能体外部。作者据此将现有自我改进系统置于同一坐标系中,并指出递归自我改进相对于经典广义策略迭代(GPI)的四种缺陷。
正文节选
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement Abstract Abstract When we speak of recursive self-improvement (RSI), are we speaking of a phenomenon, a mechanism, or a prospect? Towards autonomous and evolving intelligence, RSI is being claimed at many scales, while no single framework that formally describes these emerging instances exists. Its counterpart in the classical realm, iterative policy improvement, is characterized by