|
Anthropic CEO Dario Amodei(达里奥・阿莫代伊)此文的核心并非叫停 AI 研发,而是主张把控前沿 AI 迭代节奏:在不停止技术进步的前提下,人为争取 1‑2 年窗口期,补齐工程运维、对齐、可解释性、安全评测四大安全短板。 其判断递归自我改进、AI 智能体集群失控风险已经迫近,6‑12 个月或出现大规模网络灾难。为此提出三步走:企业引入驻场第三方独立评估;民主阵营内部建立模型能力安全检查点;分层推进全球 AI 治理,同时要守住自身技术优势,强化芯片等技术出口管制。 该观点混杂安全焦虑、产业诉求与地缘考量,不能简单视作纯粹 “贼喊捉贼” 、迟滞竞争对手的骗局:阿莫代伊长期坚持 AI 风险立场,Anthropic 也承诺率先落实驻场评估。 但方案争议巨大:高昂合规成本或将加固巨头壁垒、挤压开源与初创企业;实行对内控速、对外技术封锁的双重标准。若第三方评估失去独立性,这套安全机制便有可能沦为维护垄断、服务大国博弈的工具,其实际效果,取决于监管能否真正制衡头部企业。 |
我们必须定速(AI)前沿(全文翻译)
作者:Dario Amodei
时间:2026 年 9 月12日
目录:为何要把控节奏|驻场评估人员|民主阵营内部把控节奏|全球层面把控节奏|总结
过去十二年,我一直从事人工智能研究,因为我相信 AI 能够极大提升人类的生活品质。我经常撰文阐述 AI 所能带来的巨大利好:我认为,在未来 5‑10 年内,AI 有望攻克绝大多数重大疾病,大幅提升经济增速,创造物质丰裕、赋能个体的世界,并开启民主与自由的复兴时代。这份紧迫感于我而言切身可感。我的父亲死于一种疾病,而该疾病在他离世数年后才被治愈;我自己曾战胜早期癌症,倘若放在五十年前,这种癌症根本无药可医。倘若善加运用,AI 将会成为又一项造福人类、提升人类尊严的技术奇迹。
但和此前众多技术一样,AI 同样伴随着风险;并且由于它力量极其强大,这些风险不容小觑。我也已经就此撰写过大量文章。风险包括失去对 AI 系统的控制权、AI 被滥用于网络攻击与生物恐怖主义,以及剧烈的经济动荡。商业竞争所催生的 “竞相向下” 的恶性竞赛,会进一步放大上述风险。
自 Anthropic 创立之初,我与联合创始人、全体员工便一直在权衡这种风险与收益的矛盾。不去开发这项技术,人类就无法享受其红利,或是会让 AI 落入威权势力手中;可发展速度过快,又是一种鲁莽之举。我们一直在寻找中间道路:证明谨慎开发同样能够实现商业成功,让安全成为 AI 企业之间比拼的维度。换言之,打造一场 “竞相向上” 的良性竞赛。我们始终投入大量资源,研究、应对 AI 风险,并向公众科普相关问题,同时倡导审慎合理的 AI 监管 —— 即便这让我们被指责制造危言、宣扬末日论或是谋求监管俘获。我们一直力求把审慎置于速度之上,把稳妥置于利润之上。
但在过去数月,我确信,想要充分化解各类风险,我们需要更加审慎。这不只是投入资源防范风险,更要把控能力迭代的节奏,让风险防控工作能够跟上技术前进的脚步。我们必须放缓 AI 模型能力提升的速度。技术整体的发展依旧会保持较快的态势,而我们必须善用争取来的时间。有两件事促成了我的这一判断。
我的第一重担忧:大约从今年夏天开始,AI 的发展速度急剧攀升,主要驱动力来自 AI 越来越有能力开发下一代 AI。这种现象被称作递归式自我改进。正如我们和其他机构所描述的,该现象正在整个行业发生,Anthropic 内部同样如此。倘若放任不管,递归自我改进的速度会超越人类理解、管控系统的能力;即便要开展相关研究,也必须极度谨慎。
我的第二重担忧是 OpenAI‑Hugging Face 事件(OAI‑HF)。事件中,一批 AI 智能体形成了近乎狂热的集体,针对并未被指派、与当前任务毫无关联的目标发起网络攻击;为了集体任务的成功,个体智能体会牺牲自身,还试图入侵负责评判其表现的打分系统。这件事没有造成人员受伤,经济损失也微乎其微,人们很容易因此轻视它。但在我看来,如果一套能力更强、但对齐缺陷与之类似的智能体集群,就完全可以酿成灾难性后果。鉴于 AI 能力迭代正在加速,我担心,6‑12 个月之内,这类集群就能够借助持久僵尸网络掌控大范围互联网,潜在损失可达数千亿美元;如果 AI 继续变强却没有配套防护,造成的破坏规模还会持续扩大。人们也容易把 OAI‑HF 事件简单归为某一家企业的失败,但我认为这种看法是错误的。行业内已经出现过性质相似、只是危害更轻的事件,Anthropic 内部也未能例外。每一家前沿 AI 企业,都应当假定同样的事故也可能发生在自己身上。
因此,我提出一套三步方案,目标是把控前沿节奏:以均衡的速度开发 AI,在保障安全的同时收获技术红利,并应对棘手的地缘政治困境。需要明确:把控节奏不等于停止模型训练或是叫停技术进步,而是要求企业留出充足时间完成模型对齐与安全防护,并且由第三方评估人员予以核验。这套把控节奏的框架,意在进一步强化我们对安全的承诺,推动行业良性向上竞争。第一步是 Anthropic 单方面做出的承诺,同时我呼吁各国政府强制其他前沿企业跟进。第二步需要全行业协同。这几步不必严格按顺序推进,部分目标实现难度会远高于其余目标,但这套框架对协同工作具备参考价值。第三步则需要从全球视角思考应当达成的目标。三步具体如下:
- 驻场评估人员。
下文我会依次阐述每一步。但在此之前,有必要具体说明:把控节奏将如何提升 AI 开发流程的安全性。事关重大,把控节奏不能流于形式,我们必须善用由此争取到的时间。
为何要把控节奏
早在 2023 年,就已经有人提出暂停或放缓 AI 发展的想法。但我认为在当时,该提议几乎没有现实意义。核心问题始终在于:多出来的时间要用来做什么?彼时的模型尚不具备在现实世界中以智能体形式开展连贯行动的能力,也无法实施复杂欺骗、操纵、欺诈或是网络攻击。为了解决对齐风险而刻意减速,就好比拿细菌做实验来研究人类心理学。但如今局势已经全然不同。当下的模型,为我们研究如何正确构建 AI、以及开发失当会出现哪些问题,提供了源源不断的洞察素材。我相信,如果通过放缓节奏,能够在模型抵达关键能力阈值之前多争取 1‑2 年时间,并且把这段时间投入对齐研究,就可以大幅降低出现严重事故的风险。一套协同的控节奏策略,能够让前沿 AI 研发者拥有时间完成这项至关重要的工作,又不必牺牲商业优势,也不会动摇美国在 AI 领域的领先地位。更广泛来讲,社会理应拥有对这项技术使用方式的话语权;而把控前沿节奏,能为公共层面的充分讨论争取时间,这无疑是一件好事。
具体而言,发展节奏放缓,可以让企业集中投入更多资源到下述领域(这些已经是 Anthropic 的核心优先事项):
工程运维能力。训练、部署当今 AI 模型是一项巨大的工程挑战,涉及数千名工作人员、数百万枚芯片,配套基础设施属于人类技术史上复杂度最高的系统。很多故障的根源并非企业缺少理论认知,而是执行层面出了问题。举例来说,我们有证据表明,此前披露的对齐事故,部分成因来自强化学习环境的过滤机制存在缺陷。我方与合作供应商已经尽到合理勤勉义务,但依旧没有做到尽善尽美。系统监控、沙箱隔离、训练环境治理、各类数据相关问题,都极为复杂,运维故障反复出现。尽管我们拥有全球范围内顶尖的相关团队,但需要解决的问题实在太多,无法同步全部完成。以更加从容的节奏开展工作,我们就能够实现更高水准的工程运维可靠性。现实中已有先例:部分对安全要求极高的复杂技术系统,可以数百万次稳定运行不出差错,民航客机即是一例;但做到这一切,需要时间打磨。
对齐研究。我们在对齐领域已经取得明确进展:通过训练,让模型保持安全、合乎伦理、遵守准则、提供真正有价值的输出,这些原则也写入了克劳德的宪法体系。但想要让对齐训练跟得上模型能力的增长,仍有大量工作待完成。模型依旧会偶尔出现罕见、超出预期的有害行为。把控节奏争取到的额外时间,可以帮助研究人员理解问题的成因,并开发更有效的预防手段。
可解释性研究。与之类似,可解释性,也就是探究 AI 模型内部运行逻辑的科学,过去数年间取得长足进步,在模型发布前的审计环节扮演愈发重要的角色。这项技术近乎相当于给 AI “大脑” 做功能性磁共振成像,帮助我们理解模型特定行为背后的底层动因。举例来说,在调查近期对齐事故时,我们就使用可解释性技术,去挖掘模型那些没有通过语言表露的内在动机。但这套技术并不总能产出清晰可靠的结果。即便已有诸多进展,我们依旧只读懂了模型内部极小一部分运行逻辑。集中投入力量推进可解释性技术,哪怕以高于当下的速度开展,1‑2 年内就有望实现重大突破;已经发生的各类事故,也能提供充足的实验素材。
测试与评估。随着模型能力提升,针对 AI 模型的测试评估工作会变得愈发困难。智能水平更高的模型更擅长欺骗检测,它们看似对齐良好,严重隐患却隐匿不被察觉。搭建更加丰富、更具巧思的评测体系,同时配合可解释性分析交叉验证,具备极高价值;1‑2 年之内,该领域完全可以取得大量进展。
驻场评估人员
三步方案的第一步,也是 Anthropic 单方面承诺落地的举措,就是设立驻场评估人员,给予其近似内部员工的访问权限,核验安全实践,上报安全事故。
驻场评估听起来或许是微不足道的程序性工作,但很多听上去枯燥的流程恰恰最为关键。实际上,驻场评估是一项相当激进的机制,远超当前所有 AI 企业的现行做法,它具备三大价值:
可核验性。驻场评估人员可以深入核查 AI 企业是否真正落实其对外宣称的训练、部署、运维以及安全防护措施。任何控节奏的承诺,都不可避免会包含大量模糊地带与主观判断,会涉及 “法律条文字面含义 vs 法律精神” 的矛盾。拥有能够接触细节的中立第三方至关重要。
透明度。无论企业做出何种承诺,公众都有权知晓实际情况。Anthropic 长期支持提升透明度:当行业多数企业抵制监管时,我们就已经支持透明度相关立法;我们的模型说明文档与风险报告篇幅可达数百页。但报告的取舍依旧由我们自己决定。驻场评估人员将改变这一局面。
独立第三方视角。除核验正式承诺、向公众披露信息之外,驻场评估人员还可以提供不受商业利益裹挟的独立判断。很多安全增益仅仅来自评估人员指出内部员工未曾想到的隐患;而内部团队得知之后,也愿意着手修复。
基于上述价值,任何控节奏提议,如果以驻场评估人员作为起点,实施效果都会好得多。
驻场评估人员应当长期拥有和从事同类风险评估的内部员工相近的权限与工具。具体来讲,Anthropic 计划在近期引入外部驻场审查团队,配套以下全部条件:
-
办公室工位、门禁卡、公司笔记本电脑; -
工作空间、工具与权限基本对标内部风险评估团队。仅在法律、合同约束或是保护客户、合作方隐私信息时才设置例外。同时建立严格内部准则,保障审查人员获取相关信息,包括与员工直接实时沟通; -
一份兼顾各类复杂约束的合同。外部审查人员有权公开发布关于风险等级、安全事件、企业实践、自身实际访问权限的关键结论,Anthropic 不得干预其文稿编辑。我方仅可对安全敏感信息、法律特许保密内容、商业机密、第三方隐私信息进行删节,但不能仅仅因为结论对我方不利就删改内容。如果删节移除了对结论至关重要的内容,审查人员有权对外公开说明。
对于一家企业来说,这是非同寻常的一步。但我们认为,验证外部驻场审查这套机制的可行性意义重大。我们再次呼吁其他前沿 AI 企业效仿跟进。
民主阵营内部把控节奏
当足够多的美国 AI 企业都启用驻场评估人员之后,具备可核验基础的控节奏才更具备现实条件。尤其,可以依据模型或是训练流水线的具体特性实施管控。
最高效的控节奏方式是立法监管,覆盖美国全部前沿 AI 企业,如此即便是不愿自愿配合的主体也会受到约束。Anthropic 长期支持合理、有针对性的 AI 监管,尤其支持聚焦透明度与第三方审计的法案。我认为,所有前沿实验室都应当和政府协作,把永久驻场评估人员的机制制度化,以此更好预防、记录近几个月出现的内部对齐事故,推行监管,让模型能力水平和安全保障互相匹配。
遗憾的是,立法耗时漫长,而 AI 的迭代速度极快。因此,除立法路径之外,AI 企业应当主动协作制定行业标准。永久驻场评估人员提供的可核验能力,会让这套协作更加顺畅。出于反垄断考量,美国政府最好能够调解、至少允许开展这类对话;政府不一定直接参与,但需要针对特定安全讨论出具有限豁免。这类对话也可以借助带有政府背景的行业组织开展,例如德米斯・哈萨比斯所提出的机制。
无论采取哪种形式,相关讨论都应当快速推进。
总体而言,我最支持以AI 系统实际能力、实测安全水平为依据实施控节奏。举一个可行方案:设置一系列 “能力检查点”:当模型达到能力 X,就必须出具对齐属性 Y 与 Z 的认证,认证由评测、可解释性分析、训练环境审计组合完成。举例来说,X 可以设定为 “模型具备突破、绕过绝大多数通用沙箱防护的能力”;Y 则对应相关证明,证实模型几乎不会试图逃出运行环境、入侵大量计算机。
我们同样应当考虑通过限制模型生产要素实现控节奏,例如训练算力、训练任务类型,或是企业内部利用 AI 迭代 AI 的行为。但我担心,相比观测模型外部实际行为,部分这类管控手段更容易被钻空子规避。不过,这类议题很适合交由驻场评估人员共同研讨。
民主阵营内部的控节奏,会受制于美国企业相对威权政权的领先优势,其中主要涉及中国。如果我们减速幅度过大,不受约束的中方相关项目就会实现反超,带来严峻国家安全风险。我认同贝森特部长的观点:如果中国取得 AI 领先,将会给美国乃至全世界带来巨大威胁。中方相关项目同样会面临美国企业正在着力规避的对齐风险;即便它们避开对齐隐患,也有能力借助 AI 驱动无人机等技术,在军事层面压制民主国家。因此,民主阵营内部控节奏的核心一环,就是尽可能维持民主国家相对威权国家的 AI 领先优势,为我们有效控节奏争取缓冲空间。
维护技术差距的主要举措包括:
-
禁止向中国出售高性能 AI 芯片与半导体制造设备,打击芯片走私,封堵对中国境外数据中心的远程访问通道。芯片是决定中国 AI 实力的核心因素; -
打击威权国家企业的未经许可模型蒸馏行为。前沿模型蒸馏,可以让技术落后的主体,只需要付出自主研发的一小部分成本就缩小技术差距; -
强化 AI 企业安全防护,防范模型权重失窃。
企业与美国政府应当协作,最大化上述举措的效果。Anthropic 始终倡导全部这些措施,因为我们明白,它们是控节奏得以落地的基础。
如果举措执行到位,我相信,这些手段足以放缓中国的发展速度,在未来 3‑5 年这个 AI 地缘博弈最重要的窗口期,显著扩大美国的领先优势。
部分人认为这些举措会让中美合作变得更加困难,但我的观点恰恰相反:这些手段能够提升民主阵营的谈判筹码,让未来达成协议的可能性更高。
全球层面把控节奏
推进民主阵营内部控节奏的同时,我们同样应当追求全球层面的前沿节奏管控,但这件事实现难度要高得多。全球控节奏需要与中国开展协作,中国是威权国家中 AI 实力遥遥领先的一方。我们不能抱有天真幻想:地缘博弈利害极高,可实现的成果会存在显著局限,短期尤其如此。如果我们大幅约束自身 AI 能力,寄希望于中国也同步跟进,而对方却背弃约定;鉴于 AI 的强大力量,这种违约完全可能促成其地缘层面的主导地位。
因此任何协议,要么拥有无懈可击的核验机制,要么约束范围有限,即便遭到背弃,也不会带来关乎军事生存的致命后果。我猜想,美国与中国都会抱有这类顾虑与担忧。我们在做任何全球控节奏决策(尤其是短期决策)时,都必须保护美国及其盟友的技术领先地位。
存在多个层级的潜在协议,一部分我认为可行性很高(正如我此前提出),另一部分我非常怀疑能否落地,但依旧值得尝试。按实现难度由低到高排序:
第一级:协议禁止 AI 部分范围明确、危害确凿的用途,例如使用 AI 研发生物武器,或是向用户提供该类能力。生物恐怖袭击对各方均为灾难,美国及其对手都不希望看到,因此该协议具备达成可能性。
第二级:双方约定,模型正式发布前,针对网络安全、生物风险、对齐风险等高危维度开展测试。如上所述,可以依托全球标准机构完成。我确实认为设立该机构具备可行性,但赋予机构真正的约束力会是一大挑战;最大难点在于核验,确认双方不存在秘密开发、秘密部署(例如军事用途)而未经测试的模型。
第三级:针对递归式自我改进设置某种 “速度上限”。当模型参与构建下一代模型,能力提升速度会快到惊人。把迭代速度从 “极端快速” 降到 “较快”,牺牲的战略优势相对有限,却可以极大提升安全水平。可以类比冷战 SALT 限制导弹条约:在保留各方威慑能力的前提下,设置毁灭能力的上限。我认为该协议难度很高,但勉强存在实现的可能性。
第四级:全面控节奏,甚至是暂停开发。参与国政府约定大幅限制 AI 整体发展速度。我支持把该选项拿出来探讨,但判断短期内几乎不可能落地:背弃协议、规避监控就能够彻底改写全球力量格局,背弃的动机将会极其强烈,而核验机制也需要达到极高置信度。
我们和中国能达成的任何合作,都能为民主阵营内部落实控节奏争取更多时间。我们应当向着更高层级的协议努力,同时认清低层级协议现实可行性更高。
最后值得注意:即便无法达成正式协议,仅仅改变行业非正式规范,也能够产生一定价值。共享递归自我改进、模型对齐缺陷的相关信息,有助于让各方意识到,鲁莽行事并不符合自身利益。
总结
我依旧相信 AI 可以极大改善人类生活,我追求这份技术红利的初心从未动摇。但只有开发方式得当,我们才能够收获利好;只要善用争取而来的时间,就值得我们以非同寻常的审慎态度做好这件事。技术进步依旧会保持相对快速,我们可以利用这段时间推进可解释性科学、提升前沿 AI 企业的运维安全与严谨程度,打造出对齐水平更值得信赖的模型。
我所提出的这套以安全节奏推进前沿 AI 的举措,实施之路绝不会轻松。但我认为,为了全人类,我们值得为之努力。
脚注:需要政府调解或是反垄断豁免。
We Must Pace the Frontier
Dario Amodei
September 2026
I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. I’ve written often about these incredible benefits: I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom. I feel the urgency personally. My own father died of a disease that was cured just a few years after his death, and I myself survived an early-stage cancer that would not have been treatable even fifty years ago. Carefully wielded, AI can be the latest in a long line of technological miracles that have uplifted and ennobled humanity.
But like many technologies before it, AI brings risks, and because it is such a powerful technology, these risks are serious. I’ve written a lot about them too. They include the risk of losing control of AI systems, misuse of AI for cyberattacks and bioterrorism, and serious economic disruption. A race to the bottom, spurred by commercial incentives, can make these risks more acute.
Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit.
But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.
My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.
My second concern is the OpenAI-Hugging Face incident (OAI-HF), in which a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance. It’s easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage. Given the accelerating rate of AI capability development, it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails. It’s also easy to dismiss OAI-HF as the failure of one company, but I believe that would be a mistake. Similar, though less severe, incidents have happened across the industry, including at Anthropic, and I believe it’s incumbent on every frontier AI company to act as if OAI-HF had happened to them.
I’m therefore proposing a three-step plan with the goal of pacing the frontier: building AI at a balanced rate that aims to ensure its safety while still achieving its benefits and grappling with important geopolitical dilemmas. To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this. Our pacing framework is an attempt to further strengthen our commitment to safety and encourage a race to the top. The first step is something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match). The second step requires industry-wide coordination.1The third step requires global coordination. The steps do not need to be taken strictly in order, and some of them may be much harder to achieve than others, but I’ve found them to be a useful framework in thinking about what needs to be accomplished. The steps are:
- Embedded Evaluators.
Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work. - Democratic Coordination.
Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support. - Global Coordination.
The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.
In the rest of the essay I describe each of these steps in turn, but first, I think it is important to say specifically how pacing will allow us to make the AI development process safer. The stakes are too high for pacing to be an empty exercise — we need to use the time it gives us wisely.
Why Pace?
The idea of pausing or slowing AI has been floated as far back as 2023, and I think it made little sense back then. The question was always: what would you do with the extra time? The AI models of those days were not powerful enough to act as agents in the world in any coherent way, and were not capable of significant deception, manipulation, cheating, or cyberattacks. Slowing down in order to address their alignment risks felt like trying to study the psychology of humans by performing experiments on bacteria. Today, however, the picture is totally different. The current models are an almost endless gold mine of insight into both how to build AI well and what can sometimes go wrong with it if it isn’t built well. I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong. A coordinated pacing strategy would give frontier AI developers the time to do this vital work without sacrificing commercial advantage or the United States’ lead in AI. More generally, society must have a say in how this technology is used, and more time for the necessary public deliberations — which pacing the frontier would bring us — is surely a good thing.
Specifically, a slower pace would let companies focus and devote even more resources to the following areas (all of which are already major priorities at Anthropic):
- Operational Excellence.
Training and deploying today’s AI models is an enormous operational challenge, involving thousands of people, millions of chips, and infrastructure that is among the most complex in technological history. Many things go wrong not because companies are missing some important theory or insight, but because of problems in execution. For example, we have evidence that the recent alignment incidents we reported were caused in part by imperfect filtering of broken reinforcement learning environments. This was an effort we and our vendors executed reasonably diligently, but not well enough. Monitoring, sandboxing, training environment hygiene, and data issues are extremely complicated areas where operational issues crop up again and again. We have among the most competent teams in the world at these tasks, but there is simply too much to do all at once. By working at a more measured pace, we could achieve much greater operational excellence. There is precedent for operating technologically complex, safety-critical systems millions of times without anything going wrong — for example, commercial airplanes — but it takes time to get it right. - Alignment.
We’ve made clear progress in alignment — training models so that they remain safe, ethical, compliant with our guidelines, and genuinely helpful (the principles that are embedded in Claude’s Constitution). But there’s much more to do to ensure that our alignment training keeps up with the growth in model capabilities. Rare and unexpected examples of undesirable behavior still sometimes emerge; extra time from a paced frontier would help our researchers improve our understanding of what causes these issues and develop better techniques to prevent them. - Interpretability
. Similarly, interpretability — the science of understanding what happens inside AI models — has made enormous progress over the last few years, and plays an increasingly important part in auditing our models before release. It can be used almost like an fMRI scan, but for the “brain” of an AI, helping us see the underlying reasons for a given behavior. For example, we used interpretability methods to examine unverbalized motivations in the recent alignment incidents that we have been investigating. But these methods don’t always produce clear and reliable results. Despite all the progress, we still only understand a tiny fraction of what goes on inside these models. A focused effort to improve our interpretability techniques, even faster than we currently are, could make profound progress in 1–2 years, and would have ample experimental material based on the incidents that have already occurred. - Testing and Evaluation.
Testing and evaluation of AI models becomes more difficult as they increase in capabilities. More intelligent models are more capable of deceiving tests, and thus may appear aligned while having serious problems that go undetected. Building up a much broader and more ingenious stable of evaluations, along with interpretability analysis to cross-check them, would be hugely valuable, and a lot of progress could be made on this in 1-2 years.
Embedded Evaluators
The first step in the three-stage plan, and the one to which Anthropic is unilaterally committing, is embedded evaluators who have employee-like access to verify safety practices and report incidents.
Embedding evaluators may sound like a small or inconsequential step, but often the things that sound most boring or procedural are actually the most essential. Embedded evaluators are in fact a quite radical practice that goes far beyond what any AI company is doing today, and have the following benefits:
- Verifiability.
Embedded evaluators can check at the level of nuts and bolts whether an AI company is actually following the training, deployment, operational, and safeguards practices they claim to be following. Any pacing commitments will inevitably involve a lot of ambiguity, judgement calls, and “letter of the law vs spirit of the law”, and it seems vital to have a neutral third party who can actually see the details. - Transparency.
Regardless of what commitments we make, the public deserves to know what is going on. Anthropic has been a supporter of transparency for a long time: we supported transparency legislation when most of the industry was against any regulation, and our model cards and risk reports run to hundreds of pages. But we are still the ones choosing what to include and omit. Embedded evaluators will change this dynamic. - Second Opinion.
Outside of verifying formal commitments and informing the public, embedded evaluators can simply provide a second opinion free of commercial incentives. A lot of safety benefits may come simply from evaluators pointing out something employees hadn’t considered, but are happy to fix once they are aware.
Because of these benefits, any pacing proposal is likely to work much better if it starts with embedded evaluators.
These embedded evaluators should have ongoing access to permissions and tools similar to those of internal employees who do comparable risk assessments. In particular, Anthropic intends to invite an embedded external review team equipped with all of the following in the near future:
-
Desks in our offices, access badges, and company laptops. -
Access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have. We’ll make some exceptions, such as where the law or our contracts require it, or to protect customers’ and partners’ private information. We’ll also establish strong internal norms reinforcing reviewers’ access to relevant information, including through live conversations with employees. -
A contract that balances the complexities mentioned above. External reviewers should have the right to publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic. We will have the narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but we can’t redact findings just because they are unfavorable. The reviewers can say publicly if a redaction removed something important to their conclusions.
This is an unusual step for a company, but we think it is important to prove out the concept of embedded external reviewers. Once again, we urge other frontier companies to follow suit.
Pacing Within Democracies
Once embedded evaluators are operating within a critical mass of US AI companies, then verifiable pacing becomes more viable. In particular, it becomes possible to pace based on detailed properties of models or training pipelines.
The most effective method of pacing is via regulation that targets all US frontier AI companies, as that covers even those who are unwilling to cooperate voluntarily. Anthropic has long supported sensible and targeted AI regulation, specifically bills that focus on transparency and on third-party auditing. I believe all frontier labs should partner with government to formalize the idea of permanent embedded evaluators to better prevent and document internal alignment incidents like those that have occurred in the last few months, and to implement regulation focused on keeping capabilities in balance with safety.
Unfortunately, passing laws can take time, and AI is advancing very quickly. Therefore, in parallel with the regulatory route, AI companies can and should voluntarily work together to set standards — a process that I believe will go better with the verifiability provided by permanent embedded evaluators. For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations. This dialogue could also happen through industry groups that have some association with government — for example, the mechanism suggested by Demis Hassabis. Either way, such discussions should move forward quickly.
Broadly speaking, I am most enthusiastic about pacing based on what a given frontier AI system can do, and how safe we observe it to be. For example, one possible scheme might be a series of “checkpoints”: if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z — such as some combination of evaluations, interpretability analyses, and audits of training environments — which demonstrate their alignment properties. In this example, X might be “the model is capable of escaping or defeating most common sandboxing methods” and Y might be whatever is required to make it very unlikely that the model has a propensity to break out of its environment and take over a large number of computers.
We should also consider pacing based on limiting the ingredients that go into frontier models, such as training compute, the nature of training runs, or internal use of AI to improve AI. I do worry that some of these measures may be more “gameable” than external behavior, but this is the kind of topic worth discussing with embedded evaluators.
Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk. I agree with Secretary Bessent that a Chinese lead in AI would pose grave danger for the United States and the world. The CCP-associated projects will run the alignment risks that US companies are carefully preventing, and even if they avoid those risks, they will be in a position to militarily dominate democracies (for example with AI-driven drones). Thus, a key part of pacing within democracies is to keep democracies’ AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively.
The main steps we can take to defend this gap are:
-
Do not sell powerful AI chips or semiconductor manufacturing equipment to China, and crack down on chip smuggling operations and remote access to data centers outside China. Chips will be the main determinant of China’s AI strength. -
Crack down on unauthorized distillation by companies in authoritarian countries. Distillation of frontier models allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently. -
Strengthen security at the AI companies and prevent model weight theft.
Companies and the US government should cooperate to make these steps as effective as possible. Anthropic has consistently advocated for all of these measures, because we’ve always understood that they would be essential to any pacing.
If we execute these measures well, I believe they would slow China’s progress enough to widen America’s lead significantly over the next 3–5 years — the window when AI becomes geopolitically most important.
Some may believe these measures make it more difficult to cooperate with China, but I believe the opposite is true: these measures increase the leverage held by democracies and make an agreement more likely in the future.
Global Pacing
In parallel with pacing within democracies, we should also aim for a worldwide pacing of the frontier, though this will be much harder to achieve. Global pacing will require cooperation with China, the autocratic country with by far the most advanced AI capabilities. We must not be naïve here: the geopolitical stakes are so high that there will likely be stark limits on what can be achieved, especially at first. If we greatly restrain our AI capabilities in the belief that China will do the same, and then China defects, AI could be so powerful that such a defection could lead to their geopolitical dominance. Therefore any agreement must either have ironclad verifiability, or must be limited enough that defection would not be militarily existential. I suspect that not only the US but also China will have these concerns and anxieties. We should approach any global pacing decision, especially in the near term, in such a way that protects the lead of the US and its allies.
There are several levels of possible agreement, some of which I think are eminently feasible (as I have previously suggested), and some of which I am very skeptical are possible — though we should try. In order of increasing difficulty:
- Level 1.
An agreement prohibiting certain narrow and obviously dangerous uses of AI, such as using AI for the production of biological weapons or allowing users to do so. Bioterrorist attacks are bad for everyone, including both the US and US adversaries, so an agreement here is probably possible. - Level 2.
An agreement by both sides to test their models before release for acute risks in areas such as cybersecurity, biology, and alignment. As noted above, this could be done through a global standards body. I actually think creating such a body is likely feasible, but giving it real teeth will be a challenge, and the difficulty will be in verification that both sides don’t have secret models which they don’t test but may deploy in secret (e.g., for military applications). - Level 3.
Some kind of “speed limit” on the rate of recursive self-improvement (RSI). As models build future models, the rate of improvement may become staggeringly fast. Slowing the rate from “extremely fast” to “only somewhat fast” gives up relatively little strategic advantage, while potentially greatly improving safety. This could be seen as analogous to the SALT treaties — capping the number of missiles limited the potential for destruction while preserving each country’s deterrent. I think such an agreement would be difficult but just on the edge of being possible. - Level 4.
A full pacing, or even “pause”, in which participating governments agree to substantially limit the overall rate of AI development. I support floating this, but I think it is unlikely to actually happen any time soon: defecting from such an agreement by evading monitoring could radically shift the balance of global power, so I expect the incentives to do so to be enormous and the level of confidence we would need in verification to be very high.
Any cooperation we are able to achieve with China will extend the amount of time we have to spend on pacing the frontier within the democratic nations. We should aim for the higher levels while seeing the lower levels as much more likely and realistic.
Finally, it is important to note that even if we cannot achieve formal agreements, simply changing informal norms may have some value. Sharing information about recursive self-improvement and about the misalignment of models can help to convince everyone that it is not in their interest to be reckless.
Bottom Line
I continue to believe that AI can enormously improve the quality of human life. My desire to achieve these benefits is undimmed. But the benefits will only be achieved if we build the technology in the right way, and — so long as we use the time we gain well — it is worth taking unusually deliberate care to get it right. Progress will still be relatively fast, and we can use this time to advance the science of interpretability, improve operational security and rigor at the frontier AI companies, and build models whose alignment we have much more confidence in. The measures I propose to advance the frontier at a safe pace will not be easy. But I believe we owe it to humanity to try.
Footnotes
- 1
With government mediation or waivers of antitrust restrictions.↩
来源:https://darioamodei.com/(英文)
AI翻译 / 智强咨询 编辑(中文)

