识别却无法生成:文化特定亲属称谓生成基准
原标题:Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms
AI 摘要
该研究提出一个面向文化特定亲属称谓的生成基准,用印地语、泰米尔语和韩语测试五个开源权重模型(GPT-OSS-120B、Sarvam-M、Llama-3.3-70B、GLM-5.1、Kimi-K2.6)在生日祝福和婚礼邀请中生成亲属称谓的能力,并与同一关系-语言单元的四选一选择基线对比。结果显示模型选择正确称谓的比例远高于生成比例,例如GPT-OSS-120B为90.67%对36.00%,Llama-3.3-70B为77.92%对24.24%,作者将此差异解释为评估格式差距而非词汇知识完好的直接证据。研究还发现父系优势具有语言特异性,在印地语中明显,在韩语中较弱或反转,表明即使关系被明确说明,文化特定亲属称谓生成仍然困难,呼吁在多项选择测试之外采用基于生成的评估。
正文节选
Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms Abstract Current literature evaluates large language models (LLMs) on multilingual kinship understanding using multiple-choice benchmarks, treating it as a recognition problem. We instead prompt five open-weight LLMs to generate kinship terms in three non-Western languages (Hindi, Tamil, and Korean) across two communicative tasks, and pair this with a matched option-supported selection baseline. On identica