IIIb. Lock Down the Labs

IIIb. 封锁实验室

本章图录
Leo Szilard.
1. Leo Szilard.
A billboard at the Oak Ridge, Tennessee uranium enrichment facilities, 1943.
2. A billboard at the Oak Ridge, Tennessee uranium enrichment facilities, 1943.
English

**The nation’s leading AI labs treat security as an afterthought. Currently, they’re basically handing the key secrets for AGI to the CCP on a silver platter. Securing the AGI secrets and weights against the state-actor threat will be an immense effort, and we’re not on track. **

In this piece:

Toggle

Toggle



They met in the evening in Wigner’s office. “Szilard outlined the Columbia data,” Wheeler reports, “and the preliminary indications from it that at least two secondary neutrons emerge from each neutron-induced fission. Did this not mean that a nuclear explosive was certainly possible?” Not necessarily, Bohr countered.

“We tried to convince him,” Teller writes, “that we should go ahead with fission research but we should not publish the results. We should keep the results secret, lest the Nazis learn of them and produce nuclear explosions first.”

“Bohr insisted that we would never succeed in producing nuclear energy and he also insisted that secrecy must never be introduced into physics.”

Richard Rhodes, The Making of the Atomic Bomb (p. 430)

On the current course, the leading Chinese AGI labs won’t be in Beijing or Shanghai—they’ll be in San Francisco and London. In a few years, it will be clear that the AGI secrets are the United States’ most important national defense secrets—deserving treatment on par with B-21 bomber or Columbia-class submarine blueprints, let alone the proverbial “nuclear secrets”—but today, we are treating them the way we would random SaaS software. At this rate, we’re basically just handing superintelligence to the CCP.

All the trillions we will invest, the mobilization of American industrial might, the efforts of our brightest minds—none of that matters if China or others can simply steal the model weights (all a finished AI model is, all AGI will be, is a large file on a computer) or key algorithmic secrets (the key technical breakthroughs necessary to build AGI).

America’s leading AI labs self-proclaim to be building AGI: they believe that the technology they are building will, before the decade is out, be the most powerful weapon America has ever built. But they do not treat it as such. They measure their security efforts against “random tech startups,” not “key national defense projects.” As the AGI race intensifies—as it becomes clear that superintelligence will be utterly decisive in international military competition—we will have to face the full force of foreign espionage. Currently, labs are barely able to defend against scriptkiddies, let alone have “North Korea-proof security,” let alone be ready to face the Chinese Ministry of State Security bringing its full force to bear.

And this won’t just matter years in the future. Sure, who cares if GPT-4 weights are stolen—what really matters in terms of weight security is that we can secure the AGI weights down the line, so we have a few years, you might say. (Though if we’re building AGI in 2027, we really have to get moving!) But the AI labs are developing the *algorithmic secrets—*the key technical breakthroughs, the blueprints so to speak—for the AGI *right now *(in particular, the RL/self-play/synthetic data/etc “next paradigm” after LLMs to get past the data wall). AGI-level security for algorithmic secrets is necessary years before AGI-level security for weights. These algorithmic breakthroughs will matter more than a 10x or 100x larger cluster in a few years—this is a much bigger deal than export controls on compute, which the USG has been (presciently!) intensely pursuing. Right now, you needn’t even mount a dramatic espionage operation to steal these secrets: just go to any SF party or look through the office windows.

Our failure today will be irreversible soon: in the next 12-24 months, we will leak key AGI breakthroughs to the CCP. It will be the national security establishment’s single greatest regret before the decade is out.

The preservation of the free world against the authoritarian states is on the line—and a healthy lead will be the necessary buffer that gives us margin to get AI safety right, too. The United States has an advantage in the AGI race. But we will give up this lead if we don’t get serious about security very soon. Getting on this, now, is maybe even the single most important thing we need to do today to ensure AGI goes well.

Underrate state actors at your peril

Too many smart people underrate espionage.

The capabilities of states and their intelligence agencies are extremely formidable. Even in normal, non-all-out-AGI-race times (and from the little we know publicly), nation-states (or less advanced actors) have been able to:

For a further taste of what we’re dealing with when facing down intelligence agencies, I highly recommend Inside the Aquarium, a book by a Soviet GRU (military intelligence) defector.1

Already, China engages in widespread industrial espionage; the FBI director stated the PRC has a hacking operation greater than “every major nation combined.” And just a couple months ago, the Attorney General announced the arrest of a Chinese national who had stolen key AI code from Google to take back with him to the PRC (back in 2022/23, and probably just the tip of the iceberg).2

But that’s just the beginning. We must be prepared for our adversaries to “wake up to AGI” in the next few years. AI will become the #1 priority of every intelligence agency in the world. In that situation, they would be willing to employ extraordinary means and pay any cost to infiltrate the AI labs.

The threat model

There are two key assets we must protect: model weights (especially as we get close to AGI, but which takes years of preparation and practice to get right) and algorithmic secrets (starting yesterday).

Model weights

An AI model is just a large file of numbers on a server. This can be stolen. All it takes an adversary to match your trillions of dollars and your smartest minds and your decades of work is to steal this file. (Imagine if the Nazis had gotten an exact duplicate of every atomic bomb made in Los Alamos.)

If we can’t keep model weights secure, we’re just building AGI for the CCP (and, given the current trajectory of AI lab security, even North Korea).

Even besides national competition, securing model weights is critical for preventing AI catastrophes as well. All of our handwringing and protective measures are for naught if a bad actor (say, a terrorist or rogue state) can just steal the model and do whatever they want with it, circumventing any safety layers. Whatever novel WMDs superintelligence could invent would rapidly proliferate to dozens of rogue states. Moreover, security is the first line of defense against uncontrolled or misaligned AI systems, too (how stupid would we feel if we failed to contain the rogue superintelligence because we didn’t build and test it in an air-gapped cluster first?).

Securing model weights doesn’t matter that much right now: stealing GPT-4 without the underlying recipe doesn’t do that much for the CCP. But it will *really *matter in a few years, once we have AGI, systems that are genuinely incredibly powerful.

Perhaps the single scenario that most keeps me up at night is if China or another adversary is able to steal the automated-AI-researcher-model-weights on the cusp of an intelligence explosion. China could immediately use these to automate AI research themselves (even if they had previously been way behind)—and launch their own intelligence explosion. That’d be all they need to automate AI research, and build superintelligence. Any lead the US had would vanish.

Moreover, this would immediately put us in an existential race; any margin for ensuring superintelligence is safe would disappear. The CCP may well try to race through an intelligence explosion as fast as possible—even months of lead on superintelligence could mean a decisive military advantage—in the process skipping all the safety precautions any responsible US AGI effort would hope to take. We would also have to race through the intelligence explosion to avoid complete CCP dominance. Even if the US still manages to barely pull out ahead in the end, the loss of margin would mean having to run enormous risks on AI safety.

We’re miles away for sufficient security to protect weights today. Google DeepMind (perhaps the AI lab that has the best security of any of them, given Google infrastructure) at least straight-up admits this. Their Frontier Safety Framework outlines security levels 0, 1, 2, 3, and 4 (~1.5 being what you’d need to defend against well-resourced terrorist groups or cybercriminals, 3 being what you’d need to defend against the North Koreas of the world, and 4 being what you’d need to have even a shot of defending against priority efforts by the most capable state actors).3 They admit to being at level 0 (only the most banal and basic measures). If we got AGI and superintelligence soon, we’d literally deliver it to terrorist groups and every crazy dictator out there!

Critically, developing the infrastructure for weight security probably takes many years of lead times—if we think AGI in ~3-4 years is a real possibility and we need state-proof weight security then, we need to be launching the crash effort *now. *Securing weights will require innovations in hardware and radically different cluster design; and security at this level can’t be reached overnight, but requires cycles of iteration.

If we fail to prepare in time, our situation will be dire. We will be on the cusp of superintelligence, but years away from the security necessary. Our choice will be to press ahead, but directly deliver superintelligence to the CCP—with the existential race through the intelligence explosion that implies—or wait until the security crash program is complete, risking losing our lead.

Algorithmic secrets

While people are starting to appreciate (though not necessarily implement) the need for weight security, arguably even more important right now—and vastly underrated—is securing algorithmic secrets.

One way to think about this is that stealing the algorithmic secrets will be worth having a 10x or more larger cluster to the PRC:

Put simply, I think failing to protect algorithmic secrets is probably the most likely way in which China is able to stay competitive in the AGI race. (I discuss this more later.)

It’s hard to overstate how bad algorithmic secrets security is right now. Between the labs, there are thousands of people with access to the most important secrets; there is basically no background-checking, silo’ing, controls, basic infosec, etc. Things are stored on easily hackable SaaS services. People gabber at parties in SF. Anyone, with all the secrets in their head, could be offered $100M and recruited to a Chinese lab at any point.5 You can … just look through office windows. And so on. There are many articles, and rumors flying around SF, purporting to have extensive details of various lab algorithmic advances.

AI lab security isn’t much better than “random startup security.” Directly selling the AGI secrets to the CCP would at least be more honest.

… Is this what we see at OpenAI or any other American AI lab? No. In fact, what we see is the opposite — the security equivalent of swiss cheese. Chinese penetration of these labs would be trivially easy using any number of industrial espionage methods, such as simply bribing the cleaning crew to stick USB dongles into laptops. My own assumption is that all such American AI labs are fully penetrated and that China is getting nightly downloads of all American AI research and code RIGHT NOW…

Marc Andreessen

While it will be tough, I think these secrets are defensible. There are probably only dozens of people who truly “need to know” the key implementation details for a given algorithmic breakthrough at a given lab (even if a larger number need to know the basic high-level idea)—you can vet, silo, and intensively monitor these people, in addition to radically upgraded infosec.

What “supersecurity” will require

There’s a lot of low-hanging fruit on security at AI labs. Merely adopting best practices from, say, secretive hedge funds or Google-customer-data-level security, would put us in a much better position with respect to “regular” economic espionage from the CCP. Indeed, there are notable examples of private sector firms doing remarkably well at preserving secrets. Take quantitative trading firms (the Jane Streets of the world) for example. A number of people have told me that in an hour of conversation they could relay enough information to a competitor such that their firm’s alpha would go to ~zero—similar to how many key AI algorithmic secrets could be relayed in a short conversation—and yet these firms manage to keep these secrets and retain their edge.

While most of America’s leading AI labs have refused to put the national interest first—rejecting even basic security measures in this tier, if they have any cost or require any prioritization of security—picking this low-hanging fruit would be well within their abilities.

But let’s look out just a bit further. Once China begins to truly understand the import of AGI, we should expect the full force of their espionage efforts to come to bear; think billions of dollars invested, thousands of employees, and extreme measures (e.g., special operations strike teams) dedicated to infiltrating American AGI efforts. What will security for AGI and superintelligence require?

In short, this will only be possible with government help. Microsoft, for example, is regularly hacked by state actors (e.g., Russian hackers recently stole Microsoft executives’ emails, as well as government emails Microsoft hosts). A high-level security expert working in the field estimated that even with a complete private crash course, China would still likely be able to exfiltrate the AGI weights if it was their #1 priority—the only way to get this probability to the single digits would require, more or less, a government project.

While the government does not have a perfect track record on security themselves, they’re the only ones who have the infrastructure, know-how, and competencies to protect national-defense-level secrets. Basic stuff like the authority to subject employees to intense vetting; threaten imprisonment for leaking secrets; physical security for datacenters; and the vast know-how of places like the NSA and the people behind the security clearances (private companies simply don’t have the expertise on state-actor attacks).

I’m not one of the people behind the security clearances, so I can’t give a proper accounting of what security for AGI will truly require. The best public resource on this is RAND’s report on weight security. To give a taste of what this state-actor proof security will actually mean:

The giant AGI clusters are being mapped out, right now; the corresponding security effort must be too. If we are building AGI in just a few years, we have very little time.

Still, this immense effort shouldn’t lead to fatalism. The saving grace on security is that the CCP probably isn’t fully AGI-pilled yet, and so not yet investing in the most extreme efforts. American AI lab security “only” has to stay ahead of the curve compared to the intensity of Chinese espionage efforts. That means immediately upgrading security to stay ahead of “more normal” economic espionage (which we are far from resistant to, but private companies probably could be); and that means over the next couple years, as Chinese and other foreign espionage ramps up, rapidly upgrading to much more intense measures in cooperation with the government.



Some argue that strict security measures and their associated friction aren’t worth it because they would slow down American AI labs too much. But I think that’s mistaken:

Others argue that even if our secrets or weights leak, we will still manage to eke out just ahead by being faster in other ways (so we shouldn’t worry about these security measures). That, too, is mistaken, or at least running *way *too much risk:

We are not on track

When it first became clear to a few that an atomic bomb was possible, secrecy, too, was perhaps the most contentious issue. In 1939 and 1940, Leo Szilard became known “throughout the American physics community as the leading apostle of secrecy in fission matters.”10 But he was rebuffed by most; secrecy was not at all something scientists were used to, and it ran counter to many of their basic instincts of open science. But it slowly became clear what had to be done: the military potential of this research was too great for it to simply be freely shared with the Nazis. And secrecy was finally imposed, just in time.

Leo Szilard.

In the fall of 1940, Fermi had finished new carbon absorption measurements on graphite, suggesting graphite was a viable moderator for a bomb. Szilard assaulted Fermi with yet another secrecy appeal. “At this time Fermi really lost his temper; he really thought this was absurd,” Szilard recounted. Luckily, further appeals were eventually successful, and Fermi reluctantly refrained from publishing his graphite results.11

At the same time, the German project had narrowed down on two possible moderator materials: graphite and heavy water. In early 1941 at Heidelberg, Walther Bothe made an incorrect measurement on the absorption cross-section of graphite, and concluded that graphite would absorb too many neutrons to sustain a chain reaction. Since Fermi had kept his result secret, the Germans did not have Fermi’s measurements to check against, and to correct the error. This was crucial: it led the German project to pursue heavy water instead—a decisive wrong path that ultimately doomed the German nuclear weapons effort.

If not for that last-minute secrecy appeal, the German bomb project may have been a much more formidable competitor—and history might have turned out very differently.



There’s a real mental dissonance on security at the leading AI labs. They full-throatedly claim to be building AGI this decade. They emphasize that American leadership on AGI will be decisive for US national security. They are reportedly planning 7T chip buildouts that only make sense if you *really *believe in AGI. And indeed, when you bring up security, they nod and acknowledge “of course, we’ll all be in a bunker” and smirk.

And yet the reality on security could not be more divorced from that. Whenever it comes time to make hard choices to prioritize security, startup attitudes and commercial interests prevail over the national interest. The national security advisor would have a mental breakdown if he understood the level of security at the nation’s leading AI labs.

There are secrets being developed right now, that can be used for every training run in the future and will be the key unlocks to AGI, that are protected by the security of a startup and will be worth hundreds of billions of dollars to the CCP.12 The reality is that, a) in the next 12-24 months, we will develop the key algorithmic breakthroughs for AGI, and promptly leak them to the CCP, and b) we are not even on track for our weights to be secure against rogue actors like North Korea, let alone an all-out effort by China, by the time we build AGI. “Good security for a startup” simply is not even close to good enough, and we have very little time before the egregious damage to the national security of the United States becomes irreversible.

We’re developing the most powerful weapon mankind has ever created. The algorithmic secrets we are developing, right now, are literally the nation’s most important national defense secrets—the secrets that will be at the foundation of the US and her allies’ economic and military predominance by the end of the decade, the secrets that will determine whether we have the requisite lead to get AI safety right, the secrets that will determine the outcome of WWIII, the secrets that will determine the future of the free world. And yet AI lab security is probably worse than a random defense contractor making bolts.

It’s madness.

Basically nothing else we do—on national competition, and on AI safety—will matter if we don’t fix this, soon.

Next post in series*:* *** IIIc. Superalignment***

A billboard at the Oak Ridge, Tennessee uranium enrichment facilities, 1943.



One spoiler, as a taste: on graduation from their ~spy academy, before being sent overseas, aspiring spies had to prove their skills domestically: they had to acquire secret information from a Soviet scientist. The penalty for revealing state secrets, of course, was death. That is: graduating from the ~spy academy meant picking a countryman to condemn to death. HT Ilya Sutskever for the book recommendation.

By the way, the indictment provides a great illustration of how easy it is to evade security at even Google, which likely has the best security of all the AI labs (given their ability to lean on Google’s decades-long investment in security infrastructure). All it took to steal the code, without detection, was pasting code into Apple Notes, then exporting to pdf!

“DING exfiltrated these files by copying data from the Google source files into the Apple Notes application on his Google-issued MacBook laptop. DING then converted the Apple Notes into PDF files. and uploaded them from the Google network into DING Account 1. This method helped DING evade immediate detection.” *(From the **indictment*.)

He only got caught because he did a bunch of other stupid things, like immediately start prominent startups in China which got people suspicious (and later even came back to the US).

Based off of their claimed correspondence of their security levels to RAND’s weight security report’s L1-L5.

I sometimes joke that AI lab algorithmic advances are not shared with the American research community, but they are being shared with the Chinese research community↩

Indeed, I’ve heard from friends that ByteDance emailed basically every person who was on the Google Gemini paper to recruit them, offering them L8 (a very senior position, with presumably similarly high pay), and pitching them by saying they’d report directly to the CTO in America of ByteDance.

Inference fleets will likely be much larger than training clusters, and so there will be overwhelming pressure to use these inference clusters to run automated AI researchers during the intelligence explosion (and run billions of superintelligences more broadly in the immediate aftermath). The AGI/superintelligence weights could thus be exfiltrated from these clusters as well. (I worry that this is underrated, and inference clusters will be much less protected.)

But you can’t rely only on this! Hardware encryption regularly gets side-channeled. Defense-in-depth is key, of course.

For example, space to take an extra 6 months during the intelligence explosion for alignment research to make sure superintelligence doesn’t go awry, time to stabilize the situation after the invention of some novel WMDs by directing these systems to focus on defensive applications, or simply time for human decision-makers to make the right decisions given an extraordinarily rapid pace of technological change with the advent of superintelligence.

We try really hard to prevent nuclear proliferation to rogue states, even if we’d still be “ahead” on nuclear technology compared to their more limited arsenal, given the mayhem proliferation can cause.

The Making of the Atomic Bomb, p. 509

*The Making of the Atomic Bomb, *p. 507

100x+ compute efficiencies, when clusters worth $10s or $100s of billions are being built.

中文

美国顶尖人工智能实验室把安全当作事后考虑的问题。目前,它们基本上是在把 AGI(通用人工智能)的关键机密用银托盘奉送给中共。要防范国家行为体威胁、保护 AGI 的机密与权重,将是一项巨大的工程,而我们并没有走上正轨。

本文内容:

他们在傍晚时分相聚于维格纳(Wigner)的办公室。“西拉德(Szilard)概述了哥伦比亚大学的数据,”惠勒(Wheeler)转述道,“初步迹象表明,每次中子诱发的裂变至少会释放出两个次级中子。这难道不意味着核爆炸物必定能够造出来吗?”未必如此,玻尔(Bohr)反驳道。

“我们试图说服他,”泰勒(Teller)写道,“我们应该继续推进裂变研究,但不应发表研究结果。我们应该保守秘密,以免纳粹得知这些成果并抢先制造出核爆炸物。”

“玻尔坚称我们永远无法实现核能的利用,他也坚称绝不能让保密进入物理学。”

理查德·罗兹(Richard Rhodes),《原子弹秘史》(The Making of the Atomic Bomb),第 430 页

按照目前的路线,领先的中国 AGI 实验室将不会在北京或上海——它们将在旧金山和伦敦。几年之后人们就会明白,AGI 机密是美国最重要的国防机密——应当获得与 B-21 轰炸机或哥伦比亚级潜艇图纸同等的对待,更不用说那些传说中“无价”的“核机密”了——但今天,我们对待它们的方式却像对待普通的 SaaS 软件一样。照此下去,我们基本上就是在把超级智能拱手送给中共。

我们将投入的所有数万亿美元、美国工业实力的动员、我们最聪明头脑的努力——如果中国或其他国家可以直接窃取模型的权重(一个成品 AI 模型的全部,也就是未来 AGI 的全部,不过是计算机上的一个大文件)或关键的算法机密(构建 AGI 所必需的关键技术突破),那么这一切都将毫无意义。

美国顶尖 AI 实验室自称在构建 AGI:他们相信,在本十年结束之前,他们正在构建的技术将成为美国有史以来最强大的武器。但他们并没有这样对待它。他们用“普通科技初创公司”的标准来衡量自己的安全工作,而不是“国家重点国防项目”。随着 AGI 竞赛的加剧——随着超级智能在国际军事竞争中具有决定性作用这一点变得清晰——我们将不得不直面外国间谍活动的全部力量。目前,实验室几乎连脚本小子都防不住,更不用说拥有“可防朝鲜”的安全措施,更不用说做好应对中国国家安全部倾尽全力的准备。

而且这不仅仅是几年之后才会重要的事。当然,GPT-4 的权重被偷了谁在乎呢——就权重安全而言,真正重要的是我们最终能保护好 AGI 的权重,所以你可能觉得我们还有几年时间。(不过如果我们 2027 年就要建成 AGI,那可真得抓紧了!)但 AI 实验室现在就在为 AGI 开发算法机密——即关键的技术突破,可以说是蓝图(尤其是跨越数据墙所需的 LLM 之后那个“下一个范式”:强化学习/自博弈/合成数据等)。算法机密需要 AGI 级的安全,这比权重需要 AGI 级安全要早好几年。这些算法突破几年后将比一个大 10 倍或 100 倍的算力集群更重要——这比美国政府(一直在[有先见之明地]大力推行)对算力的出口管制重大得多。现在,你甚至不需要上演一出惊天动地的间谍行动就能窃取这些机密:随便去一场旧金山的派对,或者透过办公室窗户往里看就行。

我们今天的失败很快将变得不可逆转:在未来 12-24 个月内,我们将把关键的 AGI 突破泄露给中共。这将成为本十年结束前美国国家安全机构最大的遗憾。

自由世界的存续以对抗威权国家,如今岌岌可危——而一个健康的领先优势将是必要的缓冲,让我们也有余裕把 AI 安全做对。美国在 AGI 竞赛中占据优势。但如果我们不尽快认真对待安全,我们将放弃这一领先。现在着手此事,也许正是我们今天要确保 AGI 顺利发展所做的最重要的一件事。

低估国家行为体,后果自负

太多聪明人低估了间谍活动。

国家及其情报机构的能力极其强大。即便在平常时期、在并非 AGI 竞赛全面开打的年代(仅从我们公开所知的一鳞半爪来看),国家行为体(或较不先进的行为体)也已经能够:

若想进一步了解我们与情报机构交手时面对的是什么,我强烈推荐《Inside the Aquarium》(水族馆内幕),这是苏联格鲁乌(GRU,军事情报机构)叛逃者所著的一本书。1

中国已经在大规模从事工业间谍活动;FBI 局长表示,中国的黑客行动规模超过“所有主要国家加起来”。就在几个月前,美国司法部长宣布逮捕一名中国公民,此人窃取了谷歌的关键 AI 代码,打算带回中国(事情发生在 2022/23 年,而且很可能只是冰山一角)。2

但这只是开始。我们必须做好准备,让对手在未来几年“对 AGI 如梦初醒”。AI 将成为全世界每个情报机构的第一优先事项。在这种情况下,他们愿意动用非常手段、不惜任何代价渗透 AI 实验室。

威胁模型

我们必须保护两项关键资产:模型权重(尤其当我们逼近 AGI 时,但要做到位需要多年的准备和演练)和算法机密(从昨天开始就该保护了)。

模型权重

AI 模型只是服务器上一个巨大的数字文件。它可以被窃取。对手想要匹敌你数万亿美元的投入、你最聪明的人才和你数十年的工作,只需要窃取这个文件。(想象一下,如果纳粹得到了洛斯阿拉莫斯(Los Alamos)制造的每一颗原子弹的精确副本。)

如果我们无法保证模型权重的安全,那我们就是在为中共建设 AGI(而且,以 AI 实验室目前的安全走势来看,甚至是在为朝鲜建设)。

即便抛开国家间的竞争不谈,保护模型权重对于防范 AI 灾难同样至关重要。如果一个坏行为体(比如恐怖分子或流氓国家)能直接窃取模型、绕过一切安全层为所欲为,那我们所有的忧心忡忡和防护措施都将化为乌有。无论超级智能能发明出什么样的大规模杀伤性武器(WMD),它们都会迅速扩散到几十个流氓国家。此外,安全也是抵御失控或目标错位的人工智能系统的第一道防线(如果我们因为没先在隔离集群中构建和测试,就未能控制住失控的超级智能,那我们会觉得自己多愚蠢?)。

眼下保护模型权重并没有那么重要:窃取没有底层配方的 GPT-4 对中共来说用处不大。但几年后,一旦我们有了 AGI——那些真正强大得不可思议的系统——它就真的至关重要了。

*也许最让我夜不能寐的单一情景是:中国或某个对手在智能爆炸前夕偷走了自动化 AI 研究员的模型权重。*中国可以立刻用这些权重将自己的 AI 研究自动化(即便他们此前远远落后)——并启动他们自己的智能爆炸。这就是他们自动化 AI 研究、构建超级智能所需的全部。美国拥有的任何领先优势都将荡然无存。

此外,这将立即让我们陷入一场生死存亡的竞赛;任何确保超级智能安全的余裕都将消失。中共很可能会以最快速度冲刺智能爆炸——即使在超级智能上领先几个月也可能意味着决定性的军事优势——并在这一过程中跳过任何负责任的美国 AGI 努力都希望采取的所有安全预防措施。我们也必须冲刺智能爆炸,以避免被中共完全主导。即便美国最终勉强胜出,余裕的丧失也意味着我们不得不在 AI 安全上承担巨大风险。

今天,距离足以保护权重的安全水平我们还差得很远。Google DeepMind(考虑到谷歌的基础设施,它可能是所有 AI 实验室中安全做得最好的)至少直率地承认了这一点。他们的Frontier Safety Framework(前沿安全框架)概述了安全等级 0、1、2、3 和 4(约 1.5 级是抵御资源充足的恐怖组织或网络犯罪分子所需,3 级是抵御世界上这类“朝鲜”所需,4 级则是哪怕有机会抵御最有能力的国家行为体的优先攻势所需)。3 他们承认自己处于 0 级(只有最平庸、最基础的措施)。如果我们很快实现 AGI 和超级智能,我们实际上等于把它交付给恐怖组织和全世界每个疯狂的独裁者!

关键的是,建设权重安全基础设施可能需要多年的提前量——如果我们认为大约 3-4 年后实现 AGI 是现实可能,并且届时需要能抵御国家行为的权重安全,那我们现在就得启动这项突击工程。保护权重将需要硬件创新和彻底不同的集群设计;而且这种级别的安全不可能一夜达成,而是需要一轮又一轮的迭代。

如果我们没能及时做好准备,处境将十分严峻。我们将处在超级智能的门槛上,却距离所需的安全水平还有数年之遥。我们的选择将是:继续推进,但直接把超级智能交付给中共——随之而来的就是穿越智能爆炸的生死存亡竞赛——或者等待安全突击工程完成,冒着失去领先优势的风险。

算法机密

虽然人们开始认识到(尽管未必付诸实施)权重安全的必要性,但可以说目前更重要——而且被大大低估的——是算法机密的安全。

换个角度看:对中国而言,窃取算法机密将抵得上拥有一个 10 倍甚至更大的集群:

  • As discussed in Counting the OOMs, algorithmic progress is probably similarly as important as scaling up compute to AI progress. Given the baseline trend of ~0.5 OOMs of compute efficiency a year (+ additional algorithmic “unhobbling” gains on top), we should expect multiple OOMs-worth of algorithmic secrets between now and AGI. By default, I expect American labs to be years ahead; if they can defend their secrets, this could easily be worth 10x-100x compute. 正如在 Counting the OOMs 中讨论的,对 AI 进展而言,算法进步的重要性可能与扩大算力相当。鉴于算力效率每年约 0.5 个数量级(OOM)的基线趋势(再加上额外的算法“松绑”收益),从目前到 AGI 之间,我们应该预期有相当于多个数量级的算法机密。按照默认预期,我认为美国实验室会领先数年;如果他们能守住自己的机密,这很容易抵得上 10 倍到 100 倍的算力。
  • (Note that we’re willing to incur American investors 100s of billions of dollars of costs by export controlling Nvidia chips—perhaps a 3x increase in compute cost for Chinese labs—but we’re leaking 3x algorithmic secrets all over the place!) (请注意,我们愿意通过对 Nvidia 芯片实施出口管制,让美国投资者承担数千亿美元的成本——这也许会让中国实验室的算力成本上涨 3 倍——但我们却在到处泄露 3 倍价值的算法机密!)
  • Maybe even more importantly, *we may be developing the key paradigm breakthroughs for AGI right now. *As discussed previously, simply scaling up current models will hit a wall: the data wall. Even with way more compute, it won’t be possible to make a better model. The frontier AI labs are furiously at work at what comes next, from RL to synthetic data. They will probably figure out some crazy stuff—essentially, the “AlphaGo self-play”-equivalent for general intelligence. Their inventions will be as key as the invention of the LLM paradigm originally was a number of years ago, and they will be the key to building systems that go far beyond human-level. We still have an opportunity to deny China these key algorithmic breakthroughs, without which they’d be stuck at the data wall. But without better security in the next 12-24 months, we may well irreversibly supply China with these key AGI breakthroughs. 也许更重要的是,*我们现在可能正在开发 AGI 的关键范式突破。*正如此前讨论的那样,单纯扩大当前模型将撞上一堵墙:数据墙。即使算力大幅增加,也无法造出更好的模型。前沿 AI 实验室正在拼命攻关下一步,从强化学习(RL)到合成数据。他们很可能会想出一些疯狂的东西——本质上是通用智能领域“AlphaGo 自博弈”的对应物。他们的发明将像几年前 LLM 范式最初诞生时一样关键,也将是构建远超人类水平的系统的关键。我们仍有机会不让中国获得这些关键算法突破——没有它们,中国将被困在数据墙前。但如果未来 12-24 个月安全没有改善,我们很可能将不可逆转地把这些关键 AGI 突破提供给中国。
  • It’s easy to underrate how important an edge algorithmic secrets will be—because up until ~a couple years ago, everything was published. The basic idea was out there: scale up Transformers on internet text. Many algorithmic details and efficiencies were out there: Chinchilla scaling laws, MoE, etc. Thus, open source models today are pretty good, and a bunch of companies have pretty good models (mostly depending on how much $$$ they raised and how big their clusters are). But this will likely change fairly dramatically in the next couple years. Basically all of frontier algorithmic progress happens at labs these days (academia is surprisingly irrelevant), and the leading labs have stopped publishing their advances. We should expect far more divergence ahead: between labs, between countries, and between the proprietary frontier and open source models. A few American labs will be way ahead—a moat worth 10x, 100x, or more, way more than, say, 7nm vs. 3nm chips—unless they instantly leak the algorithmic secrets.4 算法机密带来的优势有多重要,人们很容易低估——因为直到大约几年前,一切都公开发表。基本思路早已公开:在互联网文本上扩展 Transformer。许多算法细节和效率手段也已公开:Chinchilla 缩放定律、MoE(混合专家)等等。因此,如今的开源模型相当不错,不少公司也拥有不错的模型(主要取决于它们融了多少钱、集群有多大)。但这在未来几年很可能会发生相当大的变化。如今,前沿算法进展基本上都发生在实验室(学术界出人意料地无关紧要),而领先实验室已经停止发表自己的进展。我们应当预期未来会出现更大的分化:实验室之间、国家之间、专有前沿与开源模型之间。少数美国实验室将遥遥领先——一个价值 10 倍、100 倍甚至更多的护城河,远超比如 7nm 与 3nm 芯片的差距——除非它们把算法机密立刻泄露出去。4

简言之,我认为无法保护好算法机密,很可能是中国得以在 AGI 竞赛中保持竞争力的最可能途径。(我将在后文进一步讨论。)

算法机密的安全状况有多糟糕,再怎么强调也不为过。各实验室之间,有成千上万的人能接触到最重要的机密;基本上没有任何背景审查、信息隔离、控制措施、基础信息安全等等。机密存放在极易被黑的 SaaS 服务上。人们在旧金山的派对上大聊特聊。任何人,只要脑子里装着全部机密,随时都可能被开出 1 亿美元($100M)的价码招募到中国实验室。5 你甚至……只需透过办公室窗户看一眼就行。诸如此类。旧金山流传着许多文章和传言,声称掌握各实验室算法进展的大量细节。

AI 实验室的安全并不比“普通初创公司的安全”好多少。直接把 AGI 机密卖给中共,至少还更诚实一些。

……这是我们在 OpenAI 或其他任何美国 AI 实验室看到的情况吗?不。事实上,我们看到的是相反的情况——在安全上就像瑞士奶酪一样千疮百孔。中国要用任何一种工业间谍手段渗透这些实验室都轻而易举,比如直接收买清洁工,往笔记本电脑上插 USB 设备。我个人的假设是,所有这类美国 AI 实验室都已被完全渗透,中国此刻正在夜夜下载所有美国 AI 研究和代码……

马克·安德森(Marc Andreessen)

虽然很艰难,但我认为这些机密是可以守住的。在某个实验室,真正“需要知道”某一算法突破关键实现细节的人可能只有几十个(即使有更多的人需要知道大致的高层思路)——你可以对这些人大规模审查、隔离和严密监控,再加上彻底升级的信息安全措施。

“超级安全”需要什么

AI 实验室的安全方面有很多唾手可得的改进空间。仅仅采纳比如行事隐秘的对冲基金的最佳实践,或达到谷歌客户数据级别的安全,就能让我们在应对中共“常规”经济间谍活动时处于好得多的位置。事实上,私营部门公司里有不少保密做得特别出色的显著例子。以量化交易公司(这个世界的 Jane Street 们)为例。不少人告诉我,在一小时的交谈中,他们能把足够多的信息传递给竞争对手,让公司的 alpha(超额收益)趋近于零——就像许多关键的 AI 算法机密也能在一次简短交谈中被传递出去一样——然而这些公司却能守住这些机密、保住自己的优势。

尽管美国大多数顶尖 AI 实验室拒绝把国家利益放在首位——只要这一层级的基本安全措施有任何成本、或需要对安全进行任何优先安排,它们连这些基本措施都拒绝——但摘取这些唾手可得的果实,完全在它们的能力范围之内。

但让我们再往前看一点。一旦中国开始真正理解 AGI 的分量,我们就应当预期他们的间谍活动会倾巢而出;想想投入的数十亿美元、数千名雇员,以及专门用于渗透美国 AGI 事业的极端手段(例如特种作战突击队)。那么,AGI 和超级智能的安全需要什么?

简言之,这只有借助政府的力量才可能实现。例如,微软经常被国家行为体入侵(例如,俄罗斯黑客最近窃取了微软高管的邮件,以及微软托管的政府邮件)。一位在该领域工作的高级安全专家估计,即使私营部门全力突击,如果窃取 AGI 权重是中国的第一优先事项,他们仍很可能能够将其偷走——要把这种可能性降到个位数,或多或少都需要一个政府项目。

虽然政府自身在安全方面并非履历完美,但他们是唯一拥有保护国防级机密所需的基础设施、专门知识和能力的一方。比如这些基本要素:对雇员进行严格审查的权限;以监禁威胁泄密者;数据中心的物理安全;以及像 NSA(美国国家安全局)这类机构和负责安全审查的人员所掌握的庞大专门知识(私营公司根本没有应对国家行为体攻击的专业能力)。

我不是负责安全审查的人,所以我无法准确说明 AGI 的安全究竟需要什么。关于这个问题最好的公开资料是兰德公司(RAND)的权重安全报告。为了让你感受一下这种“可防国家行为体”的安全实际上意味着什么:

  • Fully airgapped datacenters, with physical security on par with most secure military bases (cleared personnel, physical fortifications, onsite response team, extensive surveillance and extreme access control). 完全物理隔离(airgapped)的数据中心,物理安全与最安全的军事基地同级(通过审查的人员、实体防御工事、现场应急队伍、广泛监控和极其严格的门禁控制)。
  • And not just for training clusters—inference clusters need the same intense security6 而且不只是训练集群——推理集群也需要同样严密的安全!6
  • Novel technical advances on confidential compute / hardware encryption7 and extreme scrutiny on the entire hardware supply chain. 机密计算/硬件加密方面的全新技术进步,7 以及对整个硬件供应链的极度审查。
  • All research personnel working from a SCIF (Sensitive Compartmented Information Facility, pronounced “skiff”, see this visualization). 所有研究人员都在 SCIF(敏感隔离信息设施)中工作(英文全称 Sensitive Compartmented Information Facility,发音为“skiff”,参见此可视化)。
  • Extreme personnel vetting and security clearances (including regular employee integrity testing and the like), constant monitoring and substantially reduced freedoms to leave, and rigid information siloing. 极端的人员审查和安全许可(包括定期的员工诚信测试等)、持续监控、大幅缩小的离职自由,以及严格的信息隔离。
  • Strong internal controls, e.g. multi-key signoff to run any code. 严格的内部控制,例如运行任何代码都需要多方密钥签署批准。
  • Strict limitations on any external dependencies, and satisfying general requirements of TS/SCI networks. 对外部依赖的严格限制,以及满足 TS/SCI(绝密/敏感隔离信息)网络的一般要求。
  • Ongoing intense pen-testing by the NSA or similar. 由 NSA 或类似机构持续进行高强度渗透测试。
  • And so on… 诸如此类……

巨大的 AGI 集群此刻正在被规划出来;相应的安全工作也必须是。如果我们再过几年就要构建 AGI,那我们几乎没有多少时间。

不过,这项艰巨的努力不应该导致宿命论。安全方面的一线希望是,中共可能还没有完全“对 AGI 上头”,因此尚未投入最极端的努力。美国 AI 实验室的安全“只需要”相对中国间谍活动的强度保持领先。这意味着要立即升级安全,以领先于“更常规的”经济间谍活动(对这类间谍活动我们远谈不上有抵抗力,但私营公司或许能够做到);也意味着在未来几年,随着中国和其他外国间谍活动的升级,要与政府合作,迅速升级到强度高得多的措施。

有些人认为,严格的安全措施及其带来的摩擦不值得,因为它们会大大拖慢美国 AI 实验室的进度。但我认为这是错误的:

  • This is a tragedy of the commons problem. For a given lab’s commercial interests, security measures that cause a 10% slowdown might be deleterious in competition with other labs. But the national interest is clearly better served if every lab were willing to accept the additional friction: American AI research is way ahead of Chinese and other foreign algorithmic progress, and America retaining 90%-speed algorithmic progress as our national edge is clearly better than retaining 0% as a national edge (with everything instantly stolen)! 这是一个公地悲剧问题。对某个实验室的商业利益而言,造成 10% 减速的安全措施在与其他实验室的竞争中可能是有害的。但如果每个实验室都愿意接受额外的摩擦,国家利益显然会得到更好的维护:美国 AI 研究远远领先于中国和其他国家的算法进展,美国以 90% 速度的算法进步作为国家优势,显然比以 0% 作为国家优势(一切都被瞬间窃取)要好!
  • Moreover, ramping security now will be the less painful path in terms of research productivity in the long run. Eventually, inevitably, if only on the cusp of superintelligence, in the extraordinary arms race to come, the USG will realize the situation is unbearable and demand a security crackdown. It will be so much more painful, and cause much more of a slowdown, to have to implement extreme, state-actor-proof security measures from a standing start, rather than iteratively. 此外,从长远看,在科研生产力方面,现在就开始加强安全将是痛苦更少的路径。最终,不可避免的——哪怕只是在超级智能的门槛上——在即将到来的非凡军备竞赛中,美国政府会意识到局面难以忍受,并要求进行安全整顿。如果从零开始(而不是循序渐进地)实施极端的、可防国家行为体的安全措施,那将会痛苦得多,造成的减速也大得多。

另一些人认为,即使我们的机密或权重泄露,我们仍能通过在其他方面更快来勉强领先(所以不必担心这些安全措施)。这也是错误的,或者说至少是在承担过大的风险:

  • As I discuss in a later piece, I think the CCP may well be able to brutely outbuild the US (a 100GW cluster will be much easier for them). More generally, China might not have the same caution slowing it down that the US will (both reasonable and unreasonable caution!). Even if stealing the algorithms or weights “only” puts them on par with the US model-wise, that might be enough for them to win the race to superintelligence. 正如我在后续文章中讨论的,我认为中共很可能能在规模上以蛮力压过美国(一个 100GW 的集群对他们来说要容易得多)。更一般地说,中国可能没有美国那种会拖慢自己的谨慎(无论合理的还是不合理的谨慎!)。即使窃取算法或权重“只是”让它们在模型层面与美国平起平坐,那也可能足以让它们赢得通往超级智能的竞赛。
  • Moreover, even if the US squeaks out ahead in the end, the difference between a 1-2 year and 1-2 month lead will really matter for navigating the perils of superintelligence. A 1-2 year lead means at least a reasonable margin to get safety right, and to navigate the extremely volatile period around the intelligence explosion and post-superintelligence.8 A mere 1-2 month lead means a breakneck international arms race with extreme pressures, racing through the intelligence explosion, and no room at all to get safety right. It is that neck-and-neck, existential race in which we face the greatest risks of self-destruction. 此外,即使美国最终惊险胜出,领先 1-2 年与领先 1-2 个月之间的差别,对驾驭超级智能的危险而言将真的至关重要。领先 1-2 年意味着至少有合理的余裕把安全做对,并驾驭智能爆炸及超级智能之后那段极度动荡的时期。8 仅仅领先 1-2 个月则意味着一场压力极大的疯狂国际军备竞赛,冲刺穿过智能爆炸,完全没有把安全做对的空间。正是在那种并驾齐驱、生死存亡的竞赛中,我们面临最大的自我毁灭风险
  • Don’t forget about Russia, Iran, North Korea, and so on. Their hacking capabilities are no slouch. On the current course, we’re freely sharing superintelligence with them too! Without much better security, we’re proliferating what will be our most powerful weapon to a plethora of incredibly dangerous, reckless, and unpredictable actors.9 别忘了俄罗斯、伊朗、朝鲜等等。它们的黑客能力可不容小觑。按照目前的路线,我们也在免费与它们分享超级智能!如果没有大幅改进的安全措施,我们就是在把我们最强大的武器扩散给一大批极其危险、鲁莽、不可预测的行为体。9

我们并未走上正轨

当少数人最初意识到原子弹是可能造出来的时候,保密也许同样是最具争议的问题。在 1939 年和 1940 年,利奥·西拉德(Leo Szilard)“在整个美国物理学界”成为众所周知的“裂变事务保密运动的首席倡导者”。10 但他被大多数人拒绝了;保密根本不是科学家们习惯的东西,而且与开放科学这一科学家的许多基本本能相悖。但该做什么逐渐变得清楚:这项研究的军事潜力太过巨大,不能简单地与纳粹自由分享。保密最终得以实施,而且恰逢其时。

利奥·西拉德(Leo Szilard)

1940 年秋,费米(Fermi)完成了对石墨的新碳吸收测量,结果表明石墨是制造炸弹时一种可行的慢化剂。西拉德又用一次保密呼吁缠上了费米。“这时费米真的发了脾气;他真的认为这很荒谬,”西拉德回忆道。幸运的是,随后的呼吁最终奏效了,费米不情愿地克制住自己没有发表他的石墨测量结果。11

与此同时,德国项目把候选慢化剂缩小到了两种:石墨和重水。1941 年初在海德堡,瓦尔特·博特(Walther Bothe)对石墨的吸收截面做了一次错误的测量,并得出结论:石墨会吸收过多中子,无法维持链式反应。由于费米对自己的结果保密,德国人没有费米的测量数据可供对照、无从纠正这个错误。这一点至关重要:它导致德国项目转而采用重水——一条决定性的错误道路,最终断送了德国的核武器计划。

如果没有那次最后关头的保密呼吁,德国的原子弹项目本可能成为一个强大得多的竞争者——而历史也许会走向完全不同的方向。

在顶尖 AI 实验室,安全问题上存在着一种真实的认知失调。他们信誓旦旦地宣称要在本十年内构建 AGI。他们强调美国在 AGI 上的领导地位将对美国国家安全具有决定性意义。据报道,他们正在规划7 万亿美元的芯片建设计划,这种计划只有在你真的相信 AGI 时才有意义。而且的确,当你提起安全问题时,他们会点点头承认“当然,我们最后都会躲进地堡里”,然后露出一丝得意。

然而,安全方面的现实与这种说法再脱节不过了。每当需要做出优先考虑安全的艰难抉择时,初创公司心态和商业利益总是压过国家利益。如果国家安全顾问了解了国家顶尖 AI 实验室的安全水平,他会精神崩溃的。

有些机密此刻正在被开发出来,它们未来可用于每一次训练运行,将成为打开 AGI 大门的关键钥匙,却只受到初创公司级安全的保护,而且对中共而言价值数千亿美元。12 现实是:a) 在未来 12-24 个月内,我们将开发出 AGI 的关键算法突破,并很快把它们泄露给中共;b) 到我们构建 AGI 之时,我们甚至都没有走上让权重能抵御朝鲜这样的流氓行为体的轨道,更不用说抵御中国的全力行动。“对初创公司来说良好的安全”距离足够好还差得远,而在对美国的国家安全造成不可逆转的严重损害之前,我们几乎没有什么时间了。

我们正在开发人类有史以来最强大的武器。我们现在正在开发的算法机密,字面意义上就是这个国家最重要的国防机密——这些机密将在本十年结束时成为美国及其盟友经济和军事优势的基础,这些机密将决定我们是否拥有把 AI 安全做对所需的领先优势,这些机密将决定第三次世界大战(WWIII)的结局,这些机密将决定自由世界的未来。然而,AI 实验室的安全很可能还不如一个造螺栓的普通国防承包商。

这简直是疯了。

如果我们不尽快解决这个问题,我们做的其他一切——无论是在国家竞争方面,还是在 AI 安全方面——基本上都无关紧要了。

本系列下一篇*:* *** IIIc. Superalignment(超级对齐)***

田纳西州橡树岭(Oak Ridge)铀浓缩设施的广告牌,1943 年。

剧透一个细节,尝个鲜:在从“间谍学院”毕业、被派往海外之前,未来的间谍们必须在国内证明自己的本领:他们必须从一位苏联科学家那里获取秘密情报。泄露国家机密的惩罚,当然是死刑。也就是说:从“间谍学院”毕业意味着要挑选一个本国同胞,把他送上死路。感谢 Ilya Sutskever 推荐这本书。

顺便说一句,这份起诉书极好地说明了即使是在谷歌——很可能是所有 AI 实验室中安全做得最好的(因为它可以依托谷歌数十年在安全基础设施上的投入)——规避安全措施是多么容易。想神不知鬼不觉地偷走代码,只需要把代码粘贴进 Apple Notes,再导出成 PDF 就行!

“丁某(DING)将谷歌源文件中的数据复制到他谷歌发放的 MacBook 笔记本电脑上的 Apple Notes 应用中,从而窃取了这些文件。丁某随后将 Apple Notes 转换为 PDF 文件,并从谷歌网络上传到丁某账户 1。这种方法帮助丁某躲过了即时的检测。”(摘自起诉书。)

他之所以落网,只是因为他干了一堆别的蠢事,比如立即在中国创办显眼的新公司,引起了人们的怀疑(后来又甚至回到了美国)。

依据他们声称的、其安全等级与RAND 权重安全报告中 L1-L5 级别的对应关系。

我有时开玩笑说,AI 实验室的算法进展没有与美国科研界分享,却在与中国科研界分享!

确实,我从朋友那里听说,字节跳动(ByteDance)给谷歌 Gemini 论文的几乎每一位作者发了邮件招募他们,开出 L8 级别(一个非常高级的职位,薪资大概也相应很高),并承诺他们直接向字节跳动美国 CTO 汇报。

推理集群很可能会比训练集群大得多,因此在智能爆炸期间,将会有巨大压力用这些推理集群来运行自动化 AI 研究员(并在紧随其后的阶段更广泛地运行数十亿个超级智能)。AGI/超级智能的权重因此也可能从这些集群中被窃取。(我担心这一点被低估了,推理集群得到的保护会少得多。)

但你不能只依赖这一层!硬件加密经常遭到侧信道攻击。当然,纵深防御才是关键。

例如,在智能爆炸期间腾出额外 6 个月的时间进行对齐研究,以确保超级智能不会失控;在某种新型大规模杀伤性武器被发明出来后,有时间通过引导这些系统专注于防御性应用来稳定局势;或者仅仅是让人类决策者在超级智能带来极端快速的技术变革之际,有时间做出正确的决策。

我们会非常努力地防止核武器扩散到流氓国家,即使与它们有限的武库相比,我们在核技术上仍然“领先”,因为扩散可能造成的破坏太可怕了。

《原子弹秘史》(The Making of the Atomic Bomb),第 509 页

《原子弹秘史》(The Making of the Atomic Bomb),第 507 页

100 倍以上的算力效率,而正在建设的集群价值高达数百亿乃至数千亿美元。