同行投票压力测试发现LLM智能体词汇趋同但无分布式来源优势
原标题:Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources
AI 摘要
该研究提出PV-SST测试平台,通过448次试验和112个完整块,评估LLM智能体群体在社交反馈下的行为。结果显示,基于同行点赞的排名信息流显著提高了最终帖子的词汇相似度(核心面板平均差异+0.0082 TF-IDF余弦单位),但未发现分布式来源相比单一来源在改变立场上的可靠优势。研究强调其结论仅适用于合成LLM智能体群体,不直接推断人类行为。
正文节选
Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources Abstract Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We introduce PV-SST, a peer-voted social-platform testbed, and report a separately frozen, preregistered matched-exposure experiment spanning four topics, four unused seeds, four open-weight model families, and three prespecified larger