跳过正文

Large Language Models

Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment

PNAS 123/37 pp. e2600126123-e2600126123 2026-09-09 Original Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment Francesco Veri, Gustavo Umbelino Abstract Large Language Models (LLMs) are entering democratic contexts as instruments of governance, where the challenges at hand are ill-structured, marked by ambiguity and contestation. Ill-structured democratic problems demand more than factual precision; they call for intersubjective reasoning: context-sensitive judgments that others can understand and publicly accept. Using the Deliberative Reason Index (DRI), this study evaluates 60 LLMs against human deliberation across nine policy scenarios. Only four models consistently exceed the permutation-based null benchmark for alignment with human patterns of reason-giving. Most models fall short: their reason-preference structures rarely clear this threshold, even though their outputs can still appear coherent and persuasive. Yet outputs can appear reasonable even when this alignment is absent. The observed gap between surface plausibility and deliberative coherence urges caution: deploying LLMs in governance contexts requires prior assessment of their deliberative reasoning capacity, not just their surface outputs.

Large language models and conversational counter-arguments to antipublic sector bias

JPART — 2026-08-05 Original Large language models and conversational counter-arguments to antipublic sector bias John D Marvel, Sheeling Neo, Rachel Cho, Sangwon Ju Abstract Abstract Can a good argument change an individual’s mind? In three preregistered experiments, we explore this question in the domain of public sector organizational performance. We observe human subjects as they engage in conversations with a generative artificial intelligence programmed to argue in one of seven distinct “styles,” including a confrontational challenger style, a didactic style, and a sycophantic style. We develop a theory of effective argumentation predicting that conversational styles which are pleasant and engaging will be more persuasive than styles which are unpleasant or unstimulating. Contrary to this prediction, we find that conversational styles which challenge subjects’ negative views of government agencies produce significant positive attitude change, while sycophantic styles that indulge those views do not. Troublingly, subjects find the sycophantic styles more enjoyable, less frustrating, and more credible than the challenger styles. This dissociation between user experience and persuasive outcome—what we call “grudging persuasion”—suggests that attitude change does not require a pleasant conversational experience, and that the styles subjects enjoy most may be precisely the ones least likely to move them. Our findings point to a potentially dark side of large language model-based persuasion: sycophantic styles that users find most appealing are the least effective at correcting misinformed views.

Running With Scissors? Integrating GPT Models Into Public Policy Research

PSJ — 2026-06-23 Original Running With Scissors? Integrating GPT Models Into Public Policy Research Giulia Mariani, Allegra H. Fullerton Abstract ABSTRACT The integration of large language models (LLMs) into public policy research presents both exciting opportunities and methodological challenges. This research note explores how OpenAI's GPT can be used to semi‐automate the annotation of legislative testimony within the Advocacy Coalition Framework, focusing on emotion‐belief dyads. Building on Emotion‐Belief Analysis, we demonstrate how GPT can assist in identifying these complex constructs under human supervision. Our contributions are threefold: (1) we provide practical guidance for applying LLMs to publicly available textual data, (2) we propose a semiautomated workflow that strengthens conceptual clarity, transparency, consistency, replicability, and accessibility, and (3) we reflect on the ethical and methodological implications of LLM‐assisted research. As LLMs continue to advance, this research note aims to help scholars balance innovation with rigor and integrate these tools responsibly into policy research, offering lessons that extend to the study of frames, discourses, narratives, and other ideational dimensions of policymaking.