自监督视觉在策略蒸馏:无需特权信息提升小模型性能
原标题:Paper page - Self-Supervised Visual On-Policy Distillation
AI 摘要
本文提出了一种名为S2VOPD的自监督视觉在策略蒸馏方法,通过从原始图像蒸馏到强增强的学生视图,无需特权标注或更大教师模型即可提升小视觉语言模型性能。在六个细粒度感知基准上,该方法将Qwen3.5-4B的准确率从70.7%提升至77.4%,超越包括Qwen3-VL 235B在内的所有开源模型,并超过GPT-5.4。在相同训练数据下,该方法恢复了使用特权信息方法所获改进的96%。
正文节选
Abstract Self-supervised visual on-policy distillation improves small vision-language models by distilling from original images into strongly augmented student views without privileged annotations or larger teachers. Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest. This raises a fundamental question: where can informative asy