返回全部动态

SFT冲突与RL共存:大语言模型多任务学习的理论与实证分析

原标题:SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs

Hugging Face Daily Papers一手来源研究质量 87

AI 摘要

该研究探讨了监督微调(SFT)和强化学习(RL)在提升大语言模型多任务推理能力时的不同表现。实验发现SFT在多阶段训练中面临严重的任务冲突,而RL能使不同任务稳定共存。研究从参数层面和理论角度分析了这一机制,指出SFT的干扰受梯度范数限制,而RL的干扰受梯度方差限制,并据此提出Parallel-RL范式,通过解耦多任务训练提升效率和灵活性。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Abstract Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under multi-stage training, whereas RL enables stable coexistence across diverse tasks. Empirically, we trace this to the parameter l


发布时间:—
抓取时间:2026-08-10 15:34
来源机构:Hugging Face
阅读原文huggingface.co