跳过正文
  1. 政治学英文顶刊消息速递/
  2. 论文时间线/

The backfiring effect of weak AI safety regulation

PNAS 123/30 pp. e2509768123-e2509768123 2026-07-20

Original

The backfiring effect of weak AI safety regulation

Benjamin Laufer, Jon Kleinberg, Hoda Heidari

Abstract

Recent policy proposals aim to improve the safety of general-purpose AI, but there is little understanding of the efficacy of different regulatory approaches. We present a strategic model that explores interactions between safety regulation, general-purpose AI technology creators, and domain specialists-those who adapt the technology for specific applications. Our analysis examines how regulatory measures targeting different parts of the AI development chain affect the outcome of this game. Our model assumes AI technology is characterized by two key attributes: safety and performance. The regulator first sets a minimum safety requirement that applies to one or both players. The general-purpose creator then invests in the technology, establishing its initial safety and performance levels. Next, domain specialists refine the AI for their use cases, updating the safety and performance levels and taking the product to market. Resulting revenue is shared between the specialist and generalist. Our analysis reveals two insights: first, weak safety regulation imposed predominantly on domain specialists can backfire. While it might seem logical to regulate AI use cases, our analysis shows that weak regulations targeting domain specialists alone can reduce safety in a large class of parameterizations. Second, in contrast to the previous finding, we observe that stronger, well-placed regulation can mutually benefit all players. When regulators impose appropriate safety standards on both general-purpose AI creators and domain specialists, the regulation can function as a commitment device, leading to safety and performance gains, surpassing what is achievable under no regulation or regulating only one player.

中文

弱人工智能安全监管的反噬效应

Benjamin Laufer, Jon Kleinberg, Hoda Heidari

摘要

近期政策提案旨在提升通用人工智能的安全性,但对于不同监管方法的有效性仍缺乏了解。我们提出一个策略模型,探讨安全监管、通用人工智能技术创造者与领域专家——即针对具体应用调整该技术者——之间的互动。我们的分析考察针对人工智能开发链不同环节的监管措施如何影响这一博弈的结果。我们的模型假定人工智能技术具有两个关键属性:安全性与性能。监管者首先设定适用于一方或双方参与者的最低安全要求。通用技术创造者随后投资于该技术,确定其初始安全与性能水平。接下来,领域专家为其使用场景改进人工智能,更新安全与性能水平并将产品推向市场。由此产生的收入在领域专家与通用技术创造者之间分配。我们的分析揭示了两点洞见:第一,主要施加于领域专家的弱安全监管可能产生反效果。尽管监管人工智能使用场景看似合理,但我们的分析表明,仅针对领域专家的弱监管在一大类参数设定下会降低安全性。第二,与前一发现相反,我们观察到更强且位置适当的监管可使所有参与者共同受益。当监管者对通用人工智能创造者和领域专家都施加适当的安全标准时,监管可以发挥承诺装置的作用,带来安全与性能的提升,超过无监管或仅监管一方参与者所能达到的水平。

关键词

人工智能安全监管、通用人工智能、监管反效果、博弈论、人工智能治理、承诺装置