多模态大语言模型中的不确定性感知决策:来源、信号、校准与行动综述
原标题:Uncertainty-Aware Decision Making in Multimodal Large Language Models
AI 摘要
这篇综述围绕多模态大语言模型(MLLM)中的不确定性感知决策展开,提出了一个以决策为中心的框架,将不确定性来源、信号、校准和行动联系起来。文章指出,MLLM 的不确定性不仅源于语言,还涉及视觉、文本、时间、声学等多模态证据的质量和冲突,并强调不确定性评估应关注其能否改善系统在证据不足或冲突时的行为。综述涵盖了校准、选择性回答、弃权、澄清、检索等行动策略,并总结了开放问题。
正文节选
Uncertainty-Aware Decision Making in Multimodal Large Language Models: A Survey of Sources, Signals, Calibration, and ActionsJournal: Neurocomputing Abstract Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied evidence. Their failures are therefore not only linguistic. A fluent answer may conceal poor input quality, a perceptual error, weak grounding, conflict between modalities, uns