返回全部动态

识别却无法生成:文化特定亲属称谓生成基准

原标题:Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms

arXiv cs.CL一手来源研究质量 81

AI 摘要

该研究提出一个面向文化特定亲属称谓的生成基准,用印地语、泰米尔语和韩语测试五个开源权重模型(GPT-OSS-120B、Sarvam-M、Llama-3.3-70B、GLM-5.1、Kimi-K2.6)在生日祝福和婚礼邀请中生成亲属称谓的能力,并与同一关系-语言单元的四选一选择基线对比。结果显示模型选择正确称谓的比例远高于生成比例,例如GPT-OSS-120B为90.67%对36.00%,Llama-3.3-70B为77.92%对24.24%,作者将此差异解释为评估格式差距而非词汇知识完好的直接证据。研究还发现父系优势具有语言特异性,在印地语中明显,在韩语中较弱或反转,表明即使关系被明确说明,文化特定亲属称谓生成仍然困难,呼吁在多项选择测试之外采用基于生成的评估。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms Abstract Current literature evaluates large language models (LLMs) on multilingual kinship understanding using multiple-choice benchmarks, treating it as a recognition problem. We instead prompt five open-weight LLMs to generate kinship terms in three non-Western languages (Hindi, Tamil, and Korean) across two communicative tasks, and pair this with a matched option-supported selection baseline. On identica


发布时间:2026-09-24 12:00
抓取时间:2026-09-24 12:04
来源机构:arXiv
阅读原文arxiv.org