MemeCULT-1K:评估多模态模型的南亚文化语境与幽默理解
原标题:MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models
AI 摘要
研究人员推出MemeCULT-1K基准,包含1000个南亚多语言梗图(孟加拉语、英语、印地语),并附带文化背景注释和人工解释。评估13个视觉语言模型发现,提供文化上下文能显著提升解释质量,平均SBERT相似度从44.6升至56.4。错误分析显示闭源模型主要失败于实体和指代识别,开源模型受限于文化知识缺口,语言和音韵错误最难通过上下文修正。数据集和代码已公开。
正文节选
MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models Abstract Meme understanding goes beyond recognizing visual content or literal text; it requires implicit cultural knowledge and pragmatic inference that most vision-language models still lack. We introduce MemeCULT-1K, a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, where each meme is paired with a cultural context note and three human-written explanations, alon