跳过正文
  1. 政治学英文顶刊消息速递/
  2. 论文时间线/

There is no free benchmark: An institutional view of legal AI benchmarking

PNAS 123/30 pp. e2509757122-e2509757122 2026-07-20

Original

There is no free benchmark: An institutional view of legal AI benchmarking

Neel Guha, Andy K. Zhang, Christine Tsang, Christopher D. Manning, Julian Nyarko, Daniel E. Ho

Abstract

Despite substantial excitement around the use of AI in law, little information exists on the performance and associated risks of the domain's widely marketed tools. Recent work, for instance, has demonstrated the significant potential for "hallucinations"-wherein models make up facts, law, and precedent-leading Chief Justice Roberts to spotlight this risk in his annual report on the judiciary. We argue that there is a need for public AI benchmarking in law. First, relative to other AI application domains, the legal AI ecosystem lacks legibility-there is little information about the design and performance of many commercial legal AI systems. Legal AI has not benefited from the types of benchmarking that have catalyzed, measured, and informed AI innovation and responsible use in other domains. Second, we articulate the challenges of the institutional design of benchmarking. We illustrate how benchmarks can be captured, watered down, and abused. Careful institutional design around the why, who, what, and how of benchmarking will be critical to navigate difficult tradeoffs of transparency, objectivity, expertise, and resources. Third, addressing legal AI's illegibility requires matching institutional models to available resources and constraints. Rather than advocating for a single "best" approach to benchmarking, we show how benchmarking strategies depend on available resources.

中文

没有免费的基准:法律人工智能基准测试的制度视角

Neel Guha, Andy K. Zhang, Christine Tsang, Christopher D. Manning, Julian Nyarko, Daniel E. Ho

摘要

尽管人们对人工智能在法律领域的使用抱有极大热情,但关于该领域广泛营销工具的性能及相关风险的信息却很少。例如,近期研究已表明“幻觉”存在重大潜在风险——即模型会编造事实、法律和先例——这促使首席大法官罗伯茨在其关于司法机构的年度报告中特别强调这一风险。我们主张,法律领域需要公共人工智能基准测试。第一,相对于其他人工智能应用领域,法律人工智能生态系统缺乏可识读性——关于许多商业法律人工智能系统的设计与性能信息甚少。法律人工智能未能受益于那些在其他领域催生、衡量并指导人工智能创新与负责任使用的基准测试。第二,我们阐明了基准测试制度设计所面临的挑战。我们展示了基准如何可能被俘获、稀释和滥用。围绕基准测试的“为何、由谁、测试什么以及如何测试”进行审慎的制度设计,对于权衡透明度、客观性、专业知识与资源之间的艰难取舍至关重要。第三,解决法律人工智能的不可读性问题,需要将制度模式与可用资源及约束相匹配。我们并不主张采用单一“最佳”基准测试方法,而是展示基准测试策略如何取决于可用资源。

关键词

法律人工智能、基准测试、制度设计、人工智能治理、透明度、幻觉、公共基准