返回全部动态

监督微调何时降低指令敏感性?

原标题:When Does Supervised Fine-Tuning Reduce Instruction Sensitivity?

arXiv cs.IR一手来源研究质量 82

AI 摘要

该研究探讨了监督微调(SFT)对大型语言模型指令敏感性的影响。通过Qwen3(1.7B、4B、8B)及Mistral-7B、Gemma-2-9B模型在MS MARCO和ESCI-English上的实验,发现SFT并非统一降低指令敏感性:小模型(1.7B、4B)显著降低(54-71%),而8B模型个体变化不显著,但训练指令间的对比差异可靠。跨模型结果不一致,且评估协议会影响鲁棒性结论。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

When Does Supervised Fine-Tuning Reduce Instruction Sensitivity? Abstract. Large language models can exhibit substantial performance variation across alternative formulations of the same task instruction, yet it remains unclear how conventional task-specific supervised fine-tuning (SFT) changes this instruction sensitivity. We study this question by evaluating fixed model checkpoints under multiple paraphrased instructions and defining instruction sensitivity as the standard deviation of task pe


发布时间:2026-08-28 12:00
抓取时间:2026-08-28 18:31
来源机构:arXiv
阅读原文arxiv.org