审议赤字:LLM在民主话语中的实证批评
原标题:The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse
AI 摘要
该研究对大型语言模型在民主审议中的推理能力提出实证批评,指出其能力不能从可验证任务基准推断。通过应用政治学中的审议理性指数(DRI),对1980次五智能体LLM运行进行分析,发现LLM群体的程序性话语质量接近人类,但视角多样性仅约为人类的三分之一,且审议过程呈现发散而非收敛。研究结论认为,LLM可作为支持人类推理的工具,但当前证据不支持将其视为自主审议代理。
正文节选
The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse Abstract Large language models (LLMs) are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer exists and where decision quality instead depends on integrating plu