返回全部动态

在线策略增量蒸馏提升多语言数学推理能力

原标题:On-Policy Delta Distillation for Multilingual Math Reasoning

Hugging Face Daily Papers一手来源研究质量 82

AI 摘要

本文研究了在线策略蒸馏(OPD)及其改进版在线策略增量蒸馏(OPD^2)在多语言数学推理中的应用,涵盖英语、韩语和日语。实验基于Qwen3模型,结果显示OPD^2在韩语和日语上显著优于OPD,并缩小了英语与韩语之间的性能差距。此外,仅使用英语数据的OPD虽能提升韩语和日语性能,但会导致回答偏向英语,凸显了多语言数据的重要性。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

On-Policy Delta Distillation for Multilingual Math Reasoning Abstract On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD^2), for mathematical reasoning in English, Korean, and Japanese. OPD^2 improves OPD by using the probability gap between a post-trained teacher and its base model as the


发布时间:
抓取时间:2026-08-07 10:12
来源机构:Hugging Face
阅读原文huggingface.co