跳过正文

研究方法

Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment

PNAS 123/37 pp. e2600126123-e2600126123 2026-09-09 Original Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment Francesco Veri, Gustavo Umbelino Abstract Large Language Models (LLMs) are entering democratic contexts as instruments of governance, where the challenges at hand are ill-structured, marked by ambiguity and contestation. Ill-structured democratic problems demand more than factual precision; they call for intersubjective reasoning: context-sensitive judgments that others can understand and publicly accept. Using the Deliberative Reason Index (DRI), this study evaluates 60 LLMs against human deliberation across nine policy scenarios. Only four models consistently exceed the permutation-based null benchmark for alignment with human patterns of reason-giving. Most models fall short: their reason-preference structures rarely clear this threshold, even though their outputs can still appear coherent and persuasive. Yet outputs can appear reasonable even when this alignment is absent. The observed gap between surface plausibility and deliberative coherence urges caution: deploying LLMs in governance contexts requires prior assessment of their deliberative reasoning capacity, not just their surface outputs.

Leveraging generative AI for causal inference with unstructured data

PNAS 123/36 pp. e2530532123-e2530532123 2026-09-03 Original Leveraging generative AI for causal inference with unstructured data Kosuke Imai, Kentaro Nakamura Abstract We introduce GenAI-Powered Inference (GPI), a statistical framework for causal inference using unstructured data, including text and images. GPI leverages open-source pretrained Generative AI (GenAI) models-such as large language models and diffusion models-not only to generate unstructured data at scale but also to extract low-dimensional representations that are guaranteed to capture their underlying structure. Applying machine learning to these representations, GPI enables estimation of causal effects while quantifying estimation uncertainty. Unlike existing approaches to representation learning, GPI does not require fine-tuning of GenAI models, making it computationally efficient and broadly accessible. We illustrate the versatility of the GPI framework through three applications: 1) estimating the effects of Chinese social media censorship while adjusting for textual confounders, 2) isolating the impact of specific image features from that of other correlated features in the same image, and 3) assessing the persuasiveness of political rhetoric. An open-source software package is available for implementing GPI.

Personality pairing improves human–AI collaboration

PNAS 123/35 pp. e2530627123-e2530627123 2026-08-28 Original Personality pairing improves human–AI collaboration Harang Ju, Sinan Aral Abstract Here we examine how AI agent "personalities" interact with human personalities to shape human-AI collaboration and performance. In a large-scale, preregistered randomized experiment, we paired 1,258 participants with AI agents prompted to exhibit varying levels of the Big Five personality traits. These human-AI teams produced 7,266 display ads for a real think tank, which we evaluated using 1,168 independent human raters, and a field experiment on X that generated nearly 5 million impressions. We found that human and AI personalities individually shaped ad quality and teamwork and that human-AI personality pairings directly influenced ad quality. For example, extraverted humans paired with conscientious AI produced the lowest quality ads, followed by conscientious humans paired with agreeable AI and neurotic humans paired with conscientious AI. In the field experiment, ad quality significantly influenced ad performance, measured by click-through rates and cost-per-click. Together, these results demonstrate that personality pairing can improve human-AI collaboration and performance. They also motivate future research on the complex implications of AI personalization for human-AI collaboration, teamwork, and performance.

Value misalignment in X’s feed algorithm is a reflection of value tensions in engagement

PNAS 123/34 pp. e2610388123-e2610388123 2026-08-17 Original Value misalignment in X’s feed algorithm is a reflection of value tensions in engagement Ziv Epstein, Farnaz Jahanbakhsh, Tiziano Piccardi, Axel Peytavin, Isabel Gallegos, Shardul Sapkota, Dora Zhao, Johan Ugander, Michael S. Bernstein