PSJ 54/4 2026-09-14
Original Evaluating the Effectiveness of Open‐Source LLMs for Automated Analysis of Multilingual Consultation Feedback: A Swiss Case Study
PNAS 123/37 pp. e2600126123-e2600126123 2026-09-09 Original Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment
Francesco Veri, Gustavo Umbelino
Abstract
Large Language Models (LLMs) are entering democratic contexts as instruments of governance, where the challenges at hand are ill-structured, marked by ambiguity and contestation. Ill-structured democratic problems demand more than factual precision; they call for intersubjective reasoning: context-sensitive judgments that others can understand and publicly accept. Using the Deliberative Reason Index (DRI), this study evaluates 60 LLMs against human deliberation across nine policy scenarios. Only four models consistently exceed the permutation-based null benchmark for alignment with human patterns of reason-giving. Most models fall short: their reason-preference structures rarely clear this threshold, even though their outputs can still appear coherent and persuasive. Yet outputs can appear reasonable even when this alignment is absent. The observed gap between surface plausibility and deliberative coherence urges caution: deploying LLMs in governance contexts requires prior assessment of their deliberative reasoning capacity, not just their surface outputs.
PA — 2026-09-07
Original Strategies for Public Administration Research: Insights From the Scholarship of Christopher Hood
Ruth Dixon, Rozana Himaz, Maia King, Barbara Maria Piotrowska
AJPS — 2026-09-07
Original Designing multi‐site studies for external validity: Site selection via synthetic purposive sampling
PAR — 2026-09-03
Original The Practice and Origins of Performance Measurement: Assessing the Sunnyvale Model, 1984–2023
PNAS 123/36 pp. e2530532123-e2530532123 2026-09-03 Original Leveraging generative AI for causal inference with unstructured data
Kosuke Imai, Kentaro Nakamura
Abstract
We introduce GenAI-Powered Inference (GPI), a statistical framework for causal inference using unstructured data, including text and images. GPI leverages open-source pretrained Generative AI (GenAI) models-such as large language models and diffusion models-not only to generate unstructured data at scale but also to extract low-dimensional representations that are guaranteed to capture their underlying structure. Applying machine learning to these representations, GPI enables estimation of causal effects while quantifying estimation uncertainty. Unlike existing approaches to representation learning, GPI does not require fine-tuning of GenAI models, making it computationally efficient and broadly accessible. We illustrate the versatility of the GPI framework through three applications: 1) estimating the effects of Chinese social media censorship while adjusting for textual confounders, 2) isolating the impact of specific image features from that of other correlated features in the same image, and 3) assessing the persuasiveness of political rhetoric. An open-source software package is available for implementing GPI.
APSR — pp. 1-18 2026-09-01 Original Measuring Partisanship and Representation in Online Congressional Communications
MICHAEL KISTNER, MICHAEL HESELTINE, ROBERT ALVAREZ, MAYA FITCH, LUCAS LOTHAMER, ELIZABETH SIMAS
JOP — 2026-08-31
Original Politicians and Daughters: A Large-Sample Preregistered Multiverse Test
Anna Dreber Almenberg, Magnus Johannesson, Jaakko Pekka Meriläinen, Janne Tukiainen
PNAS 123/35 pp. e2530627123-e2530627123 2026-08-28 Original Personality pairing improves human–AI collaboration
Harang Ju, Sinan Aral
Abstract
Here we examine how AI agent "personalities" interact with human personalities to shape human-AI collaboration and performance. In a large-scale, preregistered randomized experiment, we paired 1,258 participants with AI agents prompted to exhibit varying levels of the Big Five personality traits. These human-AI teams produced 7,266 display ads for a real think tank, which we evaluated using 1,168 independent human raters, and a field experiment on X that generated nearly 5 million impressions. We found that human and AI personalities individually shaped ad quality and teamwork and that human-AI personality pairings directly influenced ad quality. For example, extraverted humans paired with conscientious AI produced the lowest quality ads, followed by conscientious humans paired with agreeable AI and neurotic humans paired with conscientious AI. In the field experiment, ad quality significantly influenced ad performance, measured by click-through rates and cost-per-click. Together, these results demonstrate that personality pairing can improve human-AI collaboration and performance. They also motivate future research on the complex implications of AI personalization for human-AI collaboration, teamwork, and performance.
PNAS 123/34 pp. e2610388123-e2610388123 2026-08-17 Original Value misalignment in X’s feed algorithm is a reflection of value tensions in engagement
Ziv Epstein, Farnaz Jahanbakhsh, Tiziano Piccardi, Axel Peytavin, Isabel Gallegos, Shardul Sapkota, Dora Zhao, Johan Ugander, Michael S. Bernstein