Original
Statutory construction and interpretation for AI
Abstract
AI systems are increasingly governed by natural language rules, yet a key challenge arising from reliance on language remains underexplored: interpretive ambiguity. As in legal systems, ambiguity arises both from how these rules are written and how they are applied. But while legal systems use institutional safeguards to manage such ambiguity, such as transparent appellate review policing interpretive constraints, rule-based AI alignment pipelines lack comparable protections. Different interpretations of the same rule can lead to inconsistent model behavior. Drawing on U.S. legal theory, we identify key gaps in current rule-based alignment pipelines by examining how legal systems constrain ambiguity at both the rule creation and rule application steps. We then propose a computational framework that formalizes interpretive ambiguity as constrained entropy minimization and introduces two law-inspired mechanisms: 1) a rule refinement pipeline that minimizes interpretive disagreement by revising ambiguous rules (analogous to agency rulemaking or iterative legislative action), and 2) prompt-based interpretive constraints that reduce disagreement during rule application (analogous to legal canons that guide judicial discretion). We evaluate our framework on a 5,000-scenario subset of the WildChat dataset and show that both interventions significantly improve judgment consistency across a panel of reasonable interpreters. Our approach offers a step toward systematically managing interpretive ambiguity, an essential step for building more robust, law-following AI systems.
中文
人工智能的成文法建构与解释
摘要
人工智能系统日益受到自然语言规则的治理,但由依赖语言所引发的一个关键挑战仍未得到充分探讨:解释歧义。与法律体系一样,歧义既源于这些规则被书写的方式,也源于它们被适用的方式。然而,法律体系使用制度性保障来管理此类歧义,例如透明的上诉审查对解释约束进行监督,而基于规则的人工智能对齐流程则缺乏类似保护。对同一规则的不同解释可能导致模型行为不一致。借鉴美国法律理论,我们通过考察法律体系如何在规则制定与规则适用两个环节约束歧义,识别了当前基于规则的对齐流程中的关键缺口。随后,我们提出一个计算框架,将解释歧义形式化为受约束的熵最小化,并引入两项受法律启发的机制:1)规则精炼流程,通过修订歧义规则来最小化解释分歧(类似于行政机关规则制定或迭代性立法行动);2)基于提示的解释约束,在规则适用过程中减少分歧(类似于指导司法裁量的法律解释准则)。我们在WildChat数据集的5,000个情景子集上评估该框架,并表明这两项干预措施显著提高了由一组理性解释者所作判断的一致性。我们的方法为系统性地管理解释歧义迈出了一步,而这是构建更稳健、遵循法律的人工智能系统的重要一步。
关键词
解释歧义、人工智能对齐、规则精炼、法律解释、熵最小化、判断一致性