Wnuan:面向企业专有知识问答的分阶段后训练方法
原标题:Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge
AI 摘要
Wnuan 提出了一种三阶段后训练流程,用于企业专有知识问答,包括从文档构建任务导向监督、带通用数据回放的监督微调,以及针对残余错误的强化学习。在 WnuanBench 基准上,32B 模型的可接受答案率从 52.76% 提升至 91.51%,但通用基准平均分下降 5.17 分,主要影响指令跟随能力。研究还发现残余错误采样优于全池和随机采样,并指出 RAG 与模型专业化并非必然互补。
正文节选
Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge Abstract Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs task-oriented supervision from documents, performs supervised fine-tuning with general-data replay, and applies reinforcement learning to residual errors. On the 707-question WnuanBench, the primary 32B route raises acceptabl