SFT冲突与RL共存:大语言模型多任务学习的理论与实证分析
原标题:SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
AI 摘要
该研究探讨了监督微调(SFT)和强化学习(RL)在提升大语言模型多任务推理能力时的不同表现。实验发现SFT在多阶段训练中面临严重的任务冲突,而RL能使不同任务稳定共存。研究从参数层面和理论角度分析了这一机制,指出SFT的干扰受梯度范数限制,而RL的干扰受梯度方差限制,并据此提出Parallel-RL范式,通过解耦多任务训练提升效率和灵活性。
正文节选
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Abstract Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under multi-stage training, whereas RL enables stable coexistence across diverse tasks. Empirically, we trace this to the parameter l