Original
Large language models and conversational counter-arguments to antipublic sector bias
Abstract
Abstract Can a good argument change an individual’s mind? In three preregistered experiments, we explore this question in the domain of public sector organizational performance. We observe human subjects as they engage in conversations with a generative artificial intelligence programmed to argue in one of seven distinct “styles,” including a confrontational challenger style, a didactic style, and a sycophantic style. We develop a theory of effective argumentation predicting that conversational styles which are pleasant and engaging will be more persuasive than styles which are unpleasant or unstimulating. Contrary to this prediction, we find that conversational styles which challenge subjects’ negative views of government agencies produce significant positive attitude change, while sycophantic styles that indulge those views do not. Troublingly, subjects find the sycophantic styles more enjoyable, less frustrating, and more credible than the challenger styles. This dissociation between user experience and persuasive outcome—what we call “grudging persuasion”—suggests that attitude change does not require a pleasant conversational experience, and that the styles subjects enjoy most may be precisely the ones least likely to move them. Our findings point to a potentially dark side of large language model-based persuasion: sycophantic styles that users find most appealing are the least effective at correcting misinformed views.
中文
大语言模型与针对反公共部门偏见的对话式反驳
摘要
摘要 一个好的论证能改变个人的想法吗?在三个预注册实验中,我们围绕公共部门组织绩效这一领域探讨该问题。我们观察人类受试者与一个生成式人工智能进行对话,该人工智能被编程为以七种不同“风格”之一进行论证,包括对抗性挑战者风格、说教式风格和谄媚式风格。我们发展了一种有效论证理论,预测令人愉快且引人投入的对话风格会比令人不快或缺乏激励的风格更具说服力。与这一预测相反,我们发现,挑战受试者对政府机构负面看法的对话风格能产生显著的正向态度改变,而迎合这些看法的谄媚式风格则不能。令人不安的是,受试者认为谄媚式风格比挑战者风格更令人愉快、更少令人沮丧且更可信。这种用户体验与说服结果之间的分离——我们称之为“不情愿的说服”——表明态度改变并不需要愉快的对话体验,而受试者最喜欢的风格可能恰恰是最不可能改变他们的风格。我们的发现指向基于大语言模型的说服的一个潜在黑暗面:用户觉得最吸引人的谄媚式风格,在纠正错误看法方面效果最差。
关键词
大语言模型、对话式说服、反公共部门偏见、态度改变、谄媚式风格、预注册实验、公共部门绩效