C4 框架:评估多模态大语言模型的跨概念创造力
原标题:Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
AI 摘要
Hugging Face 每日论文介绍了一项名为 C4 的评估框架,用于测试多模态大语言模型(MLLMs)的跨概念创造力,该框架基于中文成语设计。研究构建了包含 184 个合成项目和 37 个人类创建项目的 C4-Eval 数据集,并在十个 MLLM 上进行了评估。结果显示,最强的闭源模型准确率约为 50%,而开源模型显著较低,揭示了当前 MLLMs 在解码创造性编码意义方面的不足。
正文节选
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding Abstract Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented tasks. Cross-concept understanding is a core cognitive capacity underlying receptive creativity. It enables a perceiver to recover intended meaning from non-obvious but meaningful conc