DS@GT ARC:大语言模型用于检索增强辩论
原标题:DS@GT ARC at Touch\'e: Large Language Models for Retrieval-Augmented Debate
AI 摘要
DS@GT ARC 团队提交了 Touché 2025 检索增强辩论任务的论文,该任务包括生成辩论中的下一句话和根据格赖斯准则评估辩论回复。他们使用来自三家提供商的六种领先大语言模型,通过检索增强提示流水线进行生成和评估。分析发现,前沿 LLM 作为生成器表现强劲,作为评估器在模型家族内部高度一致,但这种共识并不能可靠地预测官方评估结果,尤其在质量准则上差距最大。
正文节选
Computer Science > Information Retrieval Title:DS@GT ARC at Touché: Large Language Models for Retrieval-Augmented Debate View PDF HTML (experimental) Abstract:We extend the DS@GT ARC working-note submission to the Touché 2025 Retrieval-Augmented Debate task. The task has two subtasks: generating the next utterance in a simulated debate, and evaluating debate responses according to the Gricean maxims of Quantity, Quality, Relation, and Manner. The DS@GT ARC submission consisted of six