少样本退化并非表示扭曲:12模型跨任务的行为与表示分析
原标题:Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 Architectures
AI 摘要
该研究挑战了少样本提示导致模型性能下降的「表示扭曲假说」,指出少样本提示比零样本提示长5-8倍,长度差异本身就能解释40-79%的表示偏移。作者提出「随机文本对照」方法,用长度匹配的随机token替换示例,分离出由示例内容引起的「内容增量」。结果发现内容增量越大,少样本提示反而越有益,而退化最严重的Llama 3.3 70B其偏移几乎全部来自长度伪影,将其对示例的注意力置零可恢复至零样本基线以上。
正文节选
Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 Architectures Abstract When few-shot prompting degrades a language model, the intuitive explanation is distortion: demonstrations shift the model’s internal representations away from the correct answer, and more shift means more harm. We show this explanation is wrong. Evaluating 12 open-weight models on two tasks (news classification and Ukrainian l