从过程损失到组装红利:多智能体LLM协作的人类基准诊断
原标题:From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration
AI 摘要
该研究对比了人类群组聊天与匹配的LLM审议轨迹,在Wason演绎推理任务上发现人类和LLM群组都表现出相同的“组装红利不对称”:讨论更常提升平均成员而非保留最佳初始成员。主要差异在过程层面:LLM群组更常跟随多数、更少呈现独特信息、更早收敛,正确的少数信号只有在早期被重新表达时才有效。基于人类群体决策研究的干预仅带来适度改善,未能消除协调瓶颈。
正文节选
From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration Abstract LLM agents are increasingly used for collaborative problem solving and human-group simulation. This makes outcome-only evaluation insufficient: if LLM groups are used as models of human groups, we need to know whether they succeed or fail through human-like deliberative mechanisms. We compare human group chats with matched LLM deliberation traces on Wason-style deductive reasoning, then test