跳过正文
  1. 政治学英文顶刊消息速递/
  2. 论文时间线/

Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment

PNAS 123/37 pp. e2600126123-e2600126123 2026-09-09

Original

Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment

Francesco Veri, Gustavo Umbelino

Abstract

Large Language Models (LLMs) are entering democratic contexts as instruments of governance, where the challenges at hand are ill-structured, marked by ambiguity and contestation. Ill-structured democratic problems demand more than factual precision; they call for intersubjective reasoning: context-sensitive judgments that others can understand and publicly accept. Using the Deliberative Reason Index (DRI), this study evaluates 60 LLMs against human deliberation across nine policy scenarios. Only four models consistently exceed the permutation-based null benchmark for alignment with human patterns of reason-giving. Most models fall short: their reason-preference structures rarely clear this threshold, even though their outputs can still appear coherent and persuasive. Yet outputs can appear reasonable even when this alignment is absent. The observed gap between surface plausibility and deliberative coherence urges caution: deploying LLMs in governance contexts requires prior assessment of their deliberative reasoning capacity, not just their surface outputs.

中文

看似合理的无稽之谈与审议推理:将大语言模型与人类判断进行基准比较

Francesco Veri, Gustavo Umbelino

摘要

大语言模型(LLMs)正作为治理工具进入民主情境,而其中的挑战往往是结构不良的,充满模糊性与争议。结构不良的民主问题不仅要求事实准确,还要求主体间推理:即能够为他人所理解并公开接受的、情境敏感的判断。本研究使用审议理性指数(DRI),在九个政策情境中将60个大语言模型与人类审议进行对比评估。只有四个模型持续超过基于置换的零基准,即在与人类给出理由模式的一致性方面。大多数模型未能达标:它们对理由的偏好结构很少越过这一阈值,尽管其输出仍可能显得连贯且有说服力。然而,即使这种一致性并不存在,输出也可能显得合理。表面合理性与审议连贯性之间观察到的差距警示我们:在治理情境中部署大语言模型,需要事先评估其审议推理能力,而不仅仅是其表面输出。

关键词

大语言模型、审议民主、审议推理、基准测试、人类判断、治理