PSJ 54/4 2026-09-14
Original Evaluating the Effectiveness of Open‐Source LLMs for Automated Analysis of Multilingual Consultation Feedback: A Swiss Case Study
PNAS 123/37 pp. e2600126123-e2600126123 2026-09-09 Original Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment
Francesco Veri, Gustavo Umbelino
Abstract
Large Language Models (LLMs) are entering democratic contexts as instruments of governance, where the challenges at hand are ill-structured, marked by ambiguity and contestation. Ill-structured democratic problems demand more than factual precision; they call for intersubjective reasoning: context-sensitive judgments that others can understand and publicly accept. Using the Deliberative Reason Index (DRI), this study evaluates 60 LLMs against human deliberation across nine policy scenarios. Only four models consistently exceed the permutation-based null benchmark for alignment with human patterns of reason-giving. Most models fall short: their reason-preference structures rarely clear this threshold, even though their outputs can still appear coherent and persuasive. Yet outputs can appear reasonable even when this alignment is absent. The observed gap between surface plausibility and deliberative coherence urges caution: deploying LLMs in governance contexts requires prior assessment of their deliberative reasoning capacity, not just their surface outputs.
PNAS 123/36 pp. e2530532123-e2530532123 2026-09-03 Original Leveraging generative AI for causal inference with unstructured data
Kosuke Imai, Kentaro Nakamura
Abstract
We introduce GenAI-Powered Inference (GPI), a statistical framework for causal inference using unstructured data, including text and images. GPI leverages open-source pretrained Generative AI (GenAI) models-such as large language models and diffusion models-not only to generate unstructured data at scale but also to extract low-dimensional representations that are guaranteed to capture their underlying structure. Applying machine learning to these representations, GPI enables estimation of causal effects while quantifying estimation uncertainty. Unlike existing approaches to representation learning, GPI does not require fine-tuning of GenAI models, making it computationally efficient and broadly accessible. We illustrate the versatility of the GPI framework through three applications: 1) estimating the effects of Chinese social media censorship while adjusting for textual confounders, 2) isolating the impact of specific image features from that of other correlated features in the same image, and 3) assessing the persuasiveness of political rhetoric. An open-source software package is available for implementing GPI.
PNAS 123/35 pp. e2530627123-e2530627123 2026-08-28 Original Personality pairing improves human–AI collaboration
Harang Ju, Sinan Aral
Abstract
Here we examine how AI agent "personalities" interact with human personalities to shape human-AI collaboration and performance. In a large-scale, preregistered randomized experiment, we paired 1,258 participants with AI agents prompted to exhibit varying levels of the Big Five personality traits. These human-AI teams produced 7,266 display ads for a real think tank, which we evaluated using 1,168 independent human raters, and a field experiment on X that generated nearly 5 million impressions. We found that human and AI personalities individually shaped ad quality and teamwork and that human-AI personality pairings directly influenced ad quality. For example, extraverted humans paired with conscientious AI produced the lowest quality ads, followed by conscientious humans paired with agreeable AI and neurotic humans paired with conscientious AI. In the field experiment, ad quality significantly influenced ad performance, measured by click-through rates and cost-per-click. Together, these results demonstrate that personality pairing can improve human-AI collaboration and performance. They also motivate future research on the complex implications of AI personalization for human-AI collaboration, teamwork, and performance.
PSJ 54/3 2026-08-09
Original Power Through Visibility: How Policy Narratives Matter in Climate Change Debate on X
Simon Bulian, Anne‐Marie Parth, Maria Becker, Lars Tapken
JPART — 2026-08-05
Original Large language models and conversational counter-arguments to antipublic sector bias
John D Marvel, Sheeling Neo, Rachel Cho, Sangwon Ju
Abstract
Abstract Can a good argument change an individual’s mind? In three preregistered experiments, we explore this question in the domain of public sector organizational performance. We observe human subjects as they engage in conversations with a generative artificial intelligence programmed to argue in one of seven distinct “styles,” including a confrontational challenger style, a didactic style, and a sycophantic style. We develop a theory of effective argumentation predicting that conversational styles which are pleasant and engaging will be more persuasive than styles which are unpleasant or unstimulating. Contrary to this prediction, we find that conversational styles which challenge subjects’ negative views of government agencies produce significant positive attitude change, while sycophantic styles that indulge those views do not. Troublingly, subjects find the sycophantic styles more enjoyable, less frustrating, and more credible than the challenger styles. This dissociation between user experience and persuasive outcome—what we call “grudging persuasion”—suggests that attitude change does not require a pleasant conversational experience, and that the styles subjects enjoy most may be precisely the ones least likely to move them. Our findings point to a potentially dark side of large language model-based persuasion: sycophantic styles that users find most appealing are the least effective at correcting misinformed views.
PNAS 123/31 pp. e2606267123-e2606267123 2026-07-27 Original When coordination is avoidable: A monotonicity analysis of organizational tasks
Harang Ju
Abstract
Organizations devote substantial resources to coordination, yet which tasks actually require it for correctness remains unclear. The problem is acute in multiagent AI systems, where coordination cost is directly measurable and can exceed the cost of the work itself. Distributed systems theory provides a precise criterion: Coordination is required when a task specification is nonmonotonic, meaning that as histories grow, new information can invalidate prior conclusions. Here we show that Thompson's classic taxonomy of interdependence maps to that criterion, yielding a decision rule for when coordination is required for correctness. We formalize the correspondence in a bridge theorem, apply the rule to 65 workflows from the American Productivity & Quality Center (APQC), and (with a calibrated large language model (LLM), 13,417 Occupational Information Network (O*NET tasks), and illustrate it in multiagent AI simulations. Under our decompositions, 74% of workflows and 42% of O*NET tasks are monotonic, implying that up to 24 to 57% of coordination spending is unnecessary for correctness.
PSJ 54/3 2026-07-23
Original Corporate Quasi‐Sovereignty: Big Tech and the Politics of Sovereign Authority in the Digital Age
PNAS 123/30 pp. e2509757122-e2509757122 2026-07-20 Original There is no free benchmark: An institutional view of legal AI benchmarking
Neel Guha, Andy K. Zhang, Christine Tsang, Christopher D. Manning, Julian Nyarko, Daniel E. Ho
PNAS 123/30 pp. e2509768123-e2509768123 2026-07-20 Original The backfiring effect of weak AI safety regulation
Benjamin Laufer, Jon Kleinberg, Hoda Heidari
Abstract
Recent policy proposals aim to improve the safety of general-purpose AI, but there is little understanding of the efficacy of different regulatory approaches. We present a strategic model that explores interactions between safety regulation, general-purpose AI technology creators, and domain specialists-those who adapt the technology for specific applications. Our analysis examines how regulatory measures targeting different parts of the AI development chain affect the outcome of this game. Our model assumes AI technology is characterized by two key attributes: safety and performance. The regulator first sets a minimum safety requirement that applies to one or both players. The general-purpose creator then invests in the technology, establishing its initial safety and performance levels. Next, domain specialists refine the AI for their use cases, updating the safety and performance levels and taking the product to market. Resulting revenue is shared between the specialist and generalist. Our analysis reveals two insights: first, weak safety regulation imposed predominantly on domain specialists can backfire. While it might seem logical to regulate AI use cases, our analysis shows that weak regulations targeting domain specialists alone can reduce safety in a large class of parameterizations. Second, in contrast to the previous finding, we observe that stronger, well-placed regulation can mutually benefit all players. When regulators impose appropriate safety standards on both general-purpose AI creators and domain specialists, the regulation can function as a commitment device, leading to safety and performance gains, surpassing what is achievable under no regulation or regulating only one player.