多LLM代理系统的动态治理:实现协作对话结果
原标题:Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes
AI 摘要
该论文提出了一种名为Experience Orchestrator(EO)的框架,通过控制理论(PID控制器、POMDP信念跟踪和上下文赌博机)来治理多LLM代理系统,以解决代理间缺乏共享目标函数导致的协作失败问题。在金融服务的模拟环境中,EO将高意向顾问联系率从46.1%提升至78.1%,提升32个百分点,且治理策略占结果差异的97%。但研究基于LLM模拟,尚未在真实人类交互中验证。
正文节选
Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes Abstract Classical multi-agent reinforcement learning composes a shared policy through joint reward optimization. LLM agents lack this foundation: deployed in multi-agent settings with structurally opposed objectives, they drift toward attractor states rather than converging to cooperative equilibria. This paper asks whether a control theory-informed governance layer can substitute for the missing goal functi