返回全部动态

评估大语言模型推荐中的品牌检索与排序

原标题:Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations

arXiv cs.IR一手来源研究质量 81

AI 摘要

该研究提出一个评估大语言模型品牌推荐的新框架,将竞争品牌集合独立于模型输出定义,并通过重复采样估计推荐概率(BRP@)和平均倒数排名(MRR@)。作者对六个LLM在五个产品类别上测试发现,仅给类别查询时成熟品牌被大量遗漏,推荐突出度更多与搜索兴趣和在线品牌讨论等市场可见度信号相关,而非传统品牌知名度;基于需求的查询会改变被检索的品牌,诊断性定位提示可使原本被遗漏的品牌在特定线索下被检索到。研究强调应将LLM推荐视为随机检索与排序过程来评估,并开源了软件与数据。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations Abstract. Large language models (LLMs) are increasingly used for product recommendation, but evaluating their recommendations presents challenges that differ from conventional information retrieval and recommender systems. LLMs can generate recommendations without an explicit candidate set, and repeated responses to the same query can produce different brands and rankings. We introduce a framework for evaluating open-


发布时间:2026-09-16 12:00
抓取时间:2026-09-16 12:30
来源机构:arXiv
阅读原文arxiv.org