返回全部动态

自监督视觉在策略蒸馏:无需特权信息提升小模型性能

原标题:Paper page - Self-Supervised Visual On-Policy Distillation

Hugging Face Daily Papers一手来源研究质量 84

AI 摘要

本文提出了一种名为S2VOPD的自监督视觉在策略蒸馏方法,通过从原始图像蒸馏到强增强的学生视图,无需特权标注或更大教师模型即可提升小视觉语言模型性能。在六个细粒度感知基准上,该方法将Qwen3.5-4B的准确率从70.7%提升至77.4%,超越包括Qwen3-VL 235B在内的所有开源模型,并超过GPT-5.4。在相同训练数据下,该方法恢复了使用特权信息方法所获改进的96%。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Abstract Self-supervised visual on-policy distillation improves small vision-language models by distilling from original images into strongly augmented student views without privileged annotations or larger teachers. Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest. This raises a fundamental question: where can informative asy


发布时间:—
抓取时间:2026-08-17 21:01
来源机构:Hugging Face
阅读原文huggingface.co