返回全部动态

TRACE Bench:任务驱动的角色扮演智能体检查表评估框架

原标题:TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation

arXiv cs.CL一手来源研究质量 83

AI 摘要

TRACE Bench 是一个任务驱动的角色扮演智能体检查表评估框架,将角色档案分解为固定检查表,并通过用户智能体自然对话动态更新检查表状态,使评分可追溯到具体检查项和对话证据。实验显示,其覆盖率在更少轮次内达到 99.91%,而现有 M2 自由对话转录仅覆盖 73.74%。该框架支持 26 个模型的排名、能力分解和闭环基准演化,以提升评估的可审计性和可靠性。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

GameMind TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic checklist evaluation framework. It decomposes each role profile offline into a fixed checklist, then uses a User Agent to converse naturally with the target roleplay model while privately updating che


发布时间:2026-08-13 12:00
抓取时间:2026-08-13 12:05
来源机构:arXiv
阅读原文arxiv.org