多模态大语言模型安全格局演变:新兴威胁与防护综述
原标题:Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards
AI 摘要
这篇arXiv论文对多模态大语言模型(MLLMs)的安全格局进行了系统性综述,提出了基于多模态的威胁分类法,涵盖对抗攻击、数据投毒、越狱和幻觉等威胁,并分析了威胁模型的变化。文章还总结了更新的安全假设和近期安全策略进展,最后讨论了开放挑战和未来方向,以推动多模态系统更规范、可扩展的安全机制发展。
正文节选
Computer Science > Machine Learning Title:Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards View PDF HTML (experimental) Abstract:Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger understanding and reasoning. However, this architectural shift reshapes the safety landscape of machine learning. Increased model complexity and cross-modal interacti