IIIa. Racing to the Trillion-Dollar Cluster
IIIa. 奔向万亿美元集群




**The most extraordinary techno-capital acceleration has been set in motion. As AI revenue grows rapidly, many trillions of dollars will go into GPU, datacenter, and power buildout before the end of the decade. The industrial mobilization, including growing US electricity production by 10s of percent, will be intense. **
In this piece: Toggle
You see, I told you it couldn’t be done without turning the whole country into a factory. You have done just that. Niels Bohr (to Edward Teller, upon learning of the scale of the Manhattan Project in 1944)
**The trillion-dollar cluster. Credit: DALLE. **
The race to AGI won’t just play out in code and behind laptops—it’ll be a race to mobilize America’s industrial might. Unlike anything else we’ve recently seen come out of Silicon Valley, AI is a massive industrial process: each new model requires a giant new cluster, soon giant new power plants, and eventually giant new chip fabs. The investments involved are staggering. But behind the scenes, they are already in motion.
In this post, I’ll walk you through numbers to give you a sense of what this will mean:
- As revenue from AI products grows rapidly—plausibly hitting a $100B annual run rate for companies like Google or Microsoft by ~2026, with powerful but pre-AGI systems—that will motivate ever-greater capital mobilization, and total AI investment could be north of $1T annually by 2027.
- We’re on the path to individual training clusters costing $100s of billions by 2028—clusters requiring power equivalent to a small/medium US state and more expensive than the International Space Station.
- By the end of the decade, we are headed to $1T+ individual training clusters, requiring power equivalent to >20% of US electricity production. Trillions of dollars of capex will churn out 100s of millions of GPUs per year overall.
Nvidia shocked the world as its datacenter sales exploded from about $14B annualized to about $90B annualized in the last year. But that’s still just the very beginning.
Training compute
Earlier, we found a roughly ~0.5 OOMs/year trend growth of AI training compute.1 If this trend were to continue for the rest of the decade, what would that mean for the largest training clusters?
| Year | OOMs | # of H100s-equivalent | Cost | Power | Power reference class |
|---|---|---|---|---|---|
| 2022 | ~GPT-4 cluster | ~10k | ~$500M | ~10 MW | ~10,000 average homes |
| ~2024 | +1 OOM | ~100k | $billions | ~100MW | ~100,000 homes |
| ~2026 | +2 OOMs | ~1M | $10s of billions | ~1 GW | The Hoover Dam, or a large nuclear reactor |
| ~2028 | +3 OOMs | ~10M | $100s of billions | ~10 GW | A small/medium US state |
| ~2030 | +4 OOMs | ~100M | $1T+ | ~100GW | >20% of US electricity production |
Scaling the largest training clusters, rough back-of-the-envelope calculations.
** Details on the calculations for the largest training clusters
**Year **The OpenAI GPT-4 tech report stated that GPT-4 finished training in August 2022. Thereafter we play forward the rough ~0.5 OOMs/year trend.
**H100s-equivalent **Semianalysis, JP Morgan, and others estimate GPT-4 was trained on ~25k A100s, and H100s are 2-3x the performance of A100s.
**Cost **Often people cite numbers like “$100M for GPT-4 training,” using just the rental cost of the GPUs (i.e., something like “how much would it cost to rent this size cluster for 3 months of training). But that’s a mistake. What matters is something more like ~the actual cost to build the cluster. If you want one of the largest clusters in the world, you can’t just rent it for 3 months! And moreover, you need the compute for more than just the flagship training run: there will be lots of derisking experiments, failed runs, other models, etc.
To approximate the GPT-4 cluster cost:
- The public estimates suggest the GPT-4 cluster being around 25k A100s.
- Assuming $1/A100-hour for 2-3 years gives roughly a $500M cost.
- Alternatively, you can estimate it as ~$25k cost per H100, 10k H100s-equivalent, and Nvidia GPUs being around ~half the cost of a cluster (the rest being power, the physical datacenter, cooling, networking, maintenance personnel, etc.).
- (For example, this total-cost-of-ownership analysis estimates that around 40% of a large cluster cost is the H100 GPUs itself, and another 13% goes to Nvidia for Infiniband networking. That said, excluding cost of capital in that calculation would mean the GPUs are about 50% of the cost, and with networking Nvidia gets a bit over 60% of the cost of the cluster.)
FLOP/$ is improving somewhat for each Nvidia generation, but not a ton. E.g., the H100 -> B100 is likely something like 1.5x FLOP/$: B100s are effectively two H100s stapled together, but retailing for <2x the cost. But the B100 was somewhat of an exception. They are surprisingly cheap, likely because Nvidia wants to crush competition. By contrast, A100s -> H100s weren’t much of a FLOP/$ improvement (2x better chip without fp8, roughly 2x the cost), maybe 1.5x if we count fp8 improvements—and that was for a two-year generation.
While I think there are some tailwinds to further FLOP/$ improvements from margin compression, GPUs might also get more expensive as they become massively constrained. Gains from AI chip specialization will continue, but it’s not clear to me there will still be game-changing technical improvements to FLOP/$ coming, given chips are already pretty specialized for AI (e.g. specialized for Transformers, and already at fp8/fp4 precision), Moore’s Law is glacial these days, and other bottleneck components like memory and interconnect are improving more slowly. If you look at Epoch’s data, there seems to be less than a 10x in FLOP/$ over the past decade for top ML GPUs, and for the aforementioned reasons if anything I’d expect this to slow down.
Something like a 35%/year improvement in FLOP/$ would give us the 1T cost for the +4 OOM cluster. Maybe FLOP/$ improves faster, but also datacenter capex is going to get more expensive—simply because you’ll need to actually build new power, doing a lot of capex up front, rather than just renting existing depreciated power plants.
These are just very rough numbers anyway. It would be very much within error bars if e.g. the 1T cluster can be done more efficiently and actually yields more like +4.5 OOMs on compute.
**Power requirements **An H100 is 700W, but there’s a bunch of datacenter power you need (cooling, networking, storage); Semianalysis estimates ~1,400W per H100.
There are some gains to be had on FLOP/Watt, though once we’ve exhausted the gains from AI chip specialization (see previous footnote), e.g. gone down to the lowest possible precision, these seem somewhat limited (mostly just chip process improvements, which are slow). That said, as power becomes more of a constraint (and thus a larger fraction of costs), chip designs might specialize to more power-efficiency at the costs of FLOPs. Still, there is still the power demand for cooling, networking, storage, and so on (which in the H100 numbers above was already roughly half the power demand).
For these back of the envelope numbers, let’s work with 1kW per H100-equivalent here; again these are just rough back-of-the-envelope calculations. (And if there were an unexpected FLOP/Watt breakthrough, I’d expect the same power expenditure, just bigger OOM compute gains.)
**Power reference classes **A 10GW cluster run continuously for a year is 87.6 TWh. By comparison, Oregon consumes about 27 TWh of electricity annually, Washington state consumes about 92 TWh annually.
A 100GW cluster run continuously for a year is 876 TWh, while total annual US electricity production is about 4,250 TWh.
This may seem hard to believe—but it appears to be happening. Zuck bought 350k H100s. Amazon bought a 1GW datacenter campus next to a nuclear power plant. Rumors suggest a 1GW, 1.4M H100-equivalent cluster (~2026-cluster) is being built in Kuwait. Media report that Microsoft and OpenAI are rumored to be working on a $100B cluster, slated for 2028 (a cost comparable to the International Space Station!). And as each generation of models shocks the world, further acceleration may yet be in store.
Perhaps the wildest part is that willingness-to-spend doesn’t even seem to be the binding constraint at the moment, at least for training clusters. It’s finding the infrastructure itself: “Where do I find 10GW?” (power for the $100B+, trend 2028 cluster) is a favorite topic of conversation in SF. What any compute guy is thinking about is securing power, land, permitting, and datacenter construction.2 While it may take you a year of waiting to get the GPUs, the lead times for these are much longer still.
The trillion-dollar cluster—+4 OOMs from the GPT-4 cluster, the ~2030 training cluster on the current trend—will be a truly extraordinary effort. The 100GW of power it’ll require is equivalent to >20% of US electricity production; imagine not just a simple warehouse with GPUs, but hundreds of power plants. Perhaps it will take a national consortium.
(Note that I think it’s pretty likely we’ll only need a ~$100B cluster, or less, for AGI. The $1T cluster might be what we’ll train and run superintelligence on, or what we’ll use for AGI if AGI is harder than expected. In any case, in a post-AGI world, having the most compute will probably still really* *matter.)
Overall compute
The above are just rough numbers for the largest training clusters. Overall investment is likely to be much larger still: a large fraction of GPUs will probably be used for inference3 (GPUs to actually run the AI systems for products), and there could be multiple players with giant clusters in the race.
My rough estimate is that 2024 will already feature $100B-$200B of AI investment:
- Nvidia datacenter revenue will hit a ~$25B/quarter run rate soon, i.e. ~$100B of capex flowing via Nvidia alone. But of course, Nvidia isn’t the only player (Google’s TPUs are great too!), and close to half of datacenter capex is on things other than the chips (site, building, cooling, power, etc.)4
**Quarterly Nvidia datacenter revenue. Plot by **Thomas Woodside*.*
- Big tech has been dramatically ramping their capex numbers: Microsoft and Google will likely do $50B+,5 AWS and Meta $40B+, in capex this year. Not all of this is AI, but combined their capex will have grown $50B-100B year-over-year because of the AI boom, and even then they are still cutting back on other capex to shift even more spending to AI. Moreover, other cloud providers, companies (e.g., Tesla is spending $10B on AI this year), and nation-states are investing in AI as well.
*Big tech capex is growing extremely rapidly since ChatGPT unleashed the AI boom. **Graphic source*.
Let’s play this forward. My best guess is overall compute investments will grow more slowly than the 3x/year largest training clusters, let’s say 2x/year.6
| Year | Annual investment | AI accelerator shipments (in H100s-equivalent) | Power as % of US electricity production7 ](#footnote-7-138) | Chips as % of current leading-edge TSMC wafer production |
|---|---|---|---|---|
| 2024 | ~$150B | ~5-10M 8 | 1-2% | 5-10%9 |
| ~2026 | ~$500B | ~10s of millions | 5% | ~25% |
| ~2028 | ~$2T | ~100M | 20% | ~100% |
| ~2030 | ~$8T | Many 100s of millions | 100% | 4x current capacity |
**Playing forward trends on total world AI investment. Rough back-of-the-envelope calculation. **
And these aren’t just my idiosyncratic numbers. AMD forecasted a $400B AI accelerator market by 2027, implying $700B+ of total AI spending, pretty close to my numbers (and they are surely much less “AGI-pilled” than I am). Sam Altman is reported to be in talks to raise funds for a project of “up to $7T” in capex to build out AI compute capacity (the number was widely mocked, but it seems less crazy if you run the numbers here…). One way or another, this massive scaleup is happening.
Will it be done? Can it be done?
The scale of investment postulated here may seem fantastical. But both the demand-side and the supply-side seem like they could support the above trajectory. The economic returns justify the investment, the scale of expenditures is not unprecedented for a new general-purpose technology, and the industrial mobilization for power and chips is doable.
**AI revenue **
Companies will make large AI investments if they expect the economic returns to justify it.
Reports suggest OpenAI was at a $1B revenue run rate in August 2023, and a $2B revenue run rate in February 2024. That’s roughly a doubling every 6 months. If that trend holds, we should see a ~$10B annual run rate by late 2024/early 2025, even without pricing in a massive surge from any next-generation model. One estimate puts Microsoft at ~$5B of incremental AI revenue already.
So far, every 10x scaleup in AI investment seems to yield the necessary returns. GPT-3.5 unleashed the ChatGPT mania. The estimated $500M cost for the GPT-4 cluster would have been paid off by the reported billions of annual revenue for Microsoft and OpenAI (see above calculations), and a “2024-class” training cluster in the billions will easily pay off if Microsoft/OpenAI AI revenue continues on track to a $10B+ revenue run rate. The boom is investment-led: it takes time from a huge order of GPUs to build the clusters, build the models, and roll them out, and the clusters being planned today are many years out. But if the returns on the last GPU order keep materializing, investment will continue to skyrocket (and outpace revenue), plowing in even more capital in a bet that the next 10x will keep paying off.
A key milestone for AI revenue that I like to think about is: when will a big tech company (Google, Microsoft, Meta, etc.) hit a $100B revenue run rate from AI (products and API)? These companies have on the order of $100B-$300B of revenue today; $100B would thus start representing a very substantial fraction of their business. Very naively extrapolating out the doubling every 6 months, supposing we hit a $10B revenue run rate in early 2025, suggests this would happen mid-2026.
That may seem like a stretch, but it seems to me to require surprisingly little imagination to reach that milestone. For example, there are around 350 million paid subscribers to Microsoft Office—could you get a third of these to be willing to pay $100/month for an AI add-on? For an average worker, that’s only a few hours a month of productivity gained; models powerful enough to make that justifiable seem very doable in the next couple years.10
It’s hard to understate the ensuing reverberations. This would make AI products the biggest revenue driver for America’s largest corporations, and by far their biggest area of growth. Forecasts of overall revenue growth for these companies would skyrocket. Stock markets would follow; we might see our first $10T company soon thereafter. Big tech at this point would be willing to go all out, each investing many hundreds of billions (at least) into further AI scaleout. We probably see our first many-hundred-billion dollar corporate bond sale then.11
Beyond $100B, it gets harder to see the contours. But if we are truly on the path to AGI, the returns will be there. White-collar workers are paid tens of trillions of dollars in wages annually worldwide; a drop-in remote worker that automates even a fraction of white-collar/cognitive jobs (imagine, say, a truly automated AI coder) would pay for the trillion-dollar cluster. If nothing else, the national security import could well motivate a government project, bundling the nation’s resources in the race to AGI (more later).
Historical precedents
$1T/year of total annual AI investment by 2027 seems outrageous. But it’s worth taking a look at other historical reference classes:
- In their peak years of funding, the Manhattan and Apollo programs reached 0.4% of GDP, or ~$100 billion annually today (surprisingly small!). At $1T/year, AI investment would be about 3% of GDP.
- Between 1996–2001, telecoms invested nearly $1 trillion in today’s dollars in building out internet infrastructure.
- From 1841 to 1850, private British railway investments totaled a cumulative ~40% of British GDP at the time. A similar fraction of US GDP would be equivalent to ~$11T over a decade.
- Many trillions are being spent on the green transition.
- Rapidly-growing economies often spend a high fraction of their GDP on investment; for example, China has spent more than 40% of its GDP on investment for two decades (equivalent to $11T annually given US GDP).
- In the historically most exigent national security circumstances—wartime—borrowing to finance the national effort has often comprised enormous fractions of GDP. During WWI, the UK and France, and Germany borrowed over 100% of their GDPs while the US borrowed over 20%; during WWII, the UK and Japan borrowed over 100% of their GDPs while the US borrowed over 60% of GDP (equivalent to over $17T today).
$1T/year of total AI investment by 2027 would be dramatic—among the very largest capital buildouts ever—but would not be unprecedented. And a trillion-dollar individual training cluster by the end of the decade seems on the table.12
Power
Probably the single biggest constraint on the supply-side will be power. Already, at nearer-term scales (1GW/2026 and especially 10GW/2028), power has become the binding constraint: there simply isn’t much spare capacity, and power contracts are usually long-term locked-in. And building, say, a new gigawatt-class nuclear power plant takes a decade. (I’ll wonder when we’ll start seeing things like tech companies buying aluminum smelting companies for their gigawatt-class power contracts.13)
Comparing trends on total US electricity production to our rough back of the envelope estimates on AI electricity demands.
Total US electricity generation has barely grown 5% in the last decade.14 Utilities are starting to get excited about AI (instead of 2.6% growth over the next 5 years, they now estimate 4.7%!). But they’re barely pricing in what’s coming. The trillion-dollar, 100GW cluster alone would require ~20% of current US electricity generation in 6 years; together with large inference capacity, demand will be multiples higher.
To most, this seems completely out of the question. Some are betting on Middle Eastern autocracies, who have been going around offering boundless power and giant clusters to get their rulers a seat at the AGI-table.
But it’s totally possible to do this in the United States: we have abundant natural gas.15
- Powering a 10GW cluster would take only a few percent of US natural gas production and could be done rapidly.
- Even the 100GW cluster is surprisingly doable.
- Right now the Marcellus/Utica shale (around Pennsylvania) alone is producing around 36 billion cubic feet a day of gas; that would be enough to generate just under 150GW continuously with generators (and combined cycle power plants could output 250 GW due to their higher efficiency).
- It would take about ~1200 new wells for the 100GW cluster.16 Each rig can drill roughly 3 wells per month, so 40 rigs (the current rig count in the Marcellus) could build up the production base for 100GW in less than a year.17The Marcellus had a rig count of ~80 as recently as 2019 so it would not be taxing to add 40 rigs to build up the production base.18
- More generally, US natural gas production has more than doubled in a decade; simply continuing that trend could power multiple trillion-dollar datacenters.19
- The harder part would be building enough generators/turbines; this wouldn’t be trivial, but it seems doable with about $100B of capex20 for 100GW of natural gas power plants. Combined cycle plants can be built in about two years; the timeline for generators would be even shorter still.21
The barriers to even trillions of dollars of datacenter buildout in the US are entirely self-made. Well-intentioned but rigid climate commitments (not just by the government, but green datacenter commitments by Microsoft, Google, Amazon, and so on) stand in the way of the obvious, fast solution. At the very least, even if we won’t do natural gas, a broad deregulatory agenda would unlock the solar/batteries/SMR/geothermal megaprojects. Permitting, utility regulation, FERC regulation of transmission lines, and NEPA environmental review makes things that should take a few years take a decade or more. We don’t have that kind of time.
We’re going to drive the AGI datacenters to the Middle East, under the thumb of brutal, capricious autocrats. I’d prefer clean energy too—but this is simply too important for US national security. We will need a new level of determination to make this happen. The power constraint can, must, and will be solved.
Chips
While chips are usually what comes to mind when people think about AI-supply-constraints, they’re likely a smaller constraint than power. Global production of AI chips is still a pretty small percent of TSMC-leading-edge production, likely less than 10%. There’s a lot of room to grow via AI becoming a larger share of TSMC production.
Indeed, 2024 production of AI chips (~5-10M H100-equivalents) would already be almost enough for the $100s of billion cluster (if they were all diverted to one cluster). From a pure logic fab standpoint ~100% of TSMC’s output for a year could already support the trillion-dollar cluster (again if all the chips went to one datacenter).22 Of course, not all of TSMC will be able to be diverted to AI, and not all of AI chip production for a year will be for one training cluster. Total AI chip demand (including inference and multiple players) by 2030 will be a multiple of TSMC’s current total leading-edge logic chip capacity, just for AI. TSMC ~doubled23 in the past 5 years; they’d likely need to go ~at least twice as fast on their pace of expansion to meet AI chip demand. Massive new fab investments would be necessary.
Even if raw logic fabs won’t be the constraint, chip-on-wafer-on-substrate (CoWoS) advanced packaging (connecting chips to memory, also made by TSMC, Intel, and others) and HBM memory (for which demand is enormous) are already key bottlenecks for the current AI GPU scaleup; these are more specialized to AI, unlike the pure logic chips, so there’s less pre-existing capacity. In the near term, these will be the primary constraint on churning out more GPUs, and these will be the huge constraints as AI scales. Still, these are comparatively “easy” to scale; it’s been incredible watching TSMC literally build “greenfield” fabs (i.e. entirely new facilities from scratch) to massively scale up CoWoS production this year (and Nvidia is even starting to find CoWoS alternatives to work around the shortage).
A new TSMC Gigafab (a technological marvel) costs around $20B in capex and produces 100k wafer-starts a month. For hundreds of millions of AI GPUs a year by the end of the decade, TSMC would need to build dozens of these—as well as a huge buildout for memory, advanced packaging, networking, etc., which will be a major fraction of capex. It could add up to over $1T of capex. It will be intense, but doable. (Perhaps the biggest roadblock will not be feasibility, but TSMC not even trying—TSMC does not yet seem AI-scaling-pilled! They think AI will “only” grow at a glacial 50% CAGR.)
Recent USG efforts like the CHIPS Act have been trying to onshore more AI chip production to the US (as insurance in case of the Taiwan contingency). While onshoring more of AI chip production to the US would be nice, it’s less critical than having the actual datacenter (on which the AGI lives) in the US. If having chip production abroad is like having uranium deposits abroad, having the AGI datacenter abroad is like having the literal nukes be built and stored abroad. Given the dysfunction and cost we’ve seen from building fabs in the US in practice, my guess is we should prioritize datacenters in the US while betting more heavily on democratic allies like Japan and South Korea for fab projects—fab buildouts there seem much more functional.
The Clusters of Democracy
Before the decade is out, many trillions of dollars of compute clusters will have been built. The only question is whether they will be built in America. Some are rumored to be betting on building them elsewhere, especially in the Middle East. Do we really want the infrastructure for the Manhattan Project to be controlled by some capricious Middle Eastern dictatorship?
The clusters that are being planned today may well be the clusters AGI and superintelligence are trained and run on, not just the “cool-big-tech-product clusters.” The national interest demands that these are built in America (or close democratic allies). Anything else creates an irreversible security risk: it risks the AGI weights getting stolen24 (and perhaps be shipped to China) (more later); it risks these dictatorships physically seizing the datacenters (to build and run AGI themselves) when the AGI race gets hot; or even if these threats are only wielded implicity, it puts AGI and superintelligence at unsavory dictator’s whims. America sorely regretted her energy dependence on the Middle East in the 70s, and we worked so hard to get out from under their thumbs. We cannot make the same mistake again.
The clusters can be built in the US, and we have to get our act together to make sure it happens in the US. American national security must come first, before the allure of free-flowing Middle Eastern cash, arcane regulation, or even, yes, admirable climate commitments. We face a real system competition—can the requisite industrial mobilization only be done in “top-down” autocracies? If American business is unshackled, America can build like none other (at least in red states). Being willing to use natural gas, or at the very least a broad-based deregulatory agenda—NEPA exemptions, fixing FERC and transmission permitting at the federal level, overriding utility regulation, using federal authorities to unlock land and rights of way—is a national security priority.
In any case—the exponential is in full swing now.
In the “old days,” when AGI was still a dirty word, some colleagues and I used to make theoretical economic models of what the path to AGI might look like. One feature of these models used to be a hypothetical “AI wakeup” moment, when the world started realizing how powerful these models could be and began rapidly ramping up their investments—culminating in multiple % of GDP towards the largest training runs.
It seemed far-off then, but that time has come. 2023 *was *“AI wakeup.”25 Behind the scenes, the most staggering techno-capital acceleration has been put into motion.
Brace for the G-forces.
Next post in series: *** IIIb. Lock Down the Labs: Security for AGI***
(What all of this means for NVDA/TSM/etc. I leave as an exercise for the reader. Hint: Those with situational awareness bought much lower than you, but it’s still not even close to fully priced in.26)
As mentioned, OOM = order of magnitude, 10x = 1 order of magnitude↩
One key uncertainty is how distributed training will be—if instead of needing that amount of power in one location, we could spread it among 100 locations, it’d be a lot easier.↩
See, for example, Zuck here; only ~45k of his H100s are in his largest training clusters, the vast majority of his 350k H100s for inference. Meta likely has heavier inference needs than other players, who serve fewer customers so far, but as everyone else’s AI products scale, I expect inference to become a strong majority of the GPUs.↩
For example, this total-cost-of-ownership analysis estimates that around 40% of a large cluster cost is the H100 GPUs itself, and another 13% goes to Nvidia for Infiniband networking. That said, excluding cost of capital in that calculation would mean the GPUs are about 50% of the cost, and with networking would mean Nvidia gets a bit over 60% of the cost of the cluster.↩
And apparently, despite Microsoft growing capex by 79% compared to a year ago in a recent quarter, their AI cloud demand still exceeds supply
A larger fraction of global GPU production will probably be going to the largest training cluster in the future than today, e.g. because of a consolidation to just a few leading labs, rather than many companies having frontier-model-scale clusters.↩
Of course, not all of this will be in the US, but to give a reference class.↩
I estimate Nvidia is going to ship on the order of 5M datacenter GPUs in 2024. A minority of those are B100s, which we’ll count as 2x+ H100s. Then there’s the other AI chips: TPUs, Trainium, Meta’s custom silicon, AMD GPUs, etc.↩
TSMC has capacity for over 150k 5nm wafers per month, is ramping to 100k 3nm wafers per month, and likely another 150k or so 7nm wafers per month; let’s call it 400k wafers per month in total. Let’s say roughly 35 H100s per wafer (H100s are made on 5nm). At 5-10 million H100-equivalents in 2024, that’s 150k-300k wafers per year for annual AI chip production in 2024. Depending on where in that range and whether we want to count 7nm production, that’s about 3-10% of annual leading-edge wafer production.↩
A big uncertainty for me is what the lags are for the technology to diffuse and be adopted. I think it’s plausible revenue is slowed because intermediate, pre-AGI models take a lot of “schlep” to properly integrate into company workflows; historically, it’s taken a while to fully harvest the productivity gains from new general purpose technologies. This is where the “sonic boom” discussion earlier comes in: as we “unhobble” models and they start looking more like agents/drop-in remote workers,, deploying them becomes much easier. Rather than having to completely remake some workflow to harvest a 25% productivity gain from a GPT-chatbot, instead you’ll get models that you can onboard and work with as you would a new coworker (e.g., just directly substitute for an engineer, rather than needing to train up engineers to use some new tool). Or, in the extreme, and later on: you won’t need to completely redesign a factory to work with some new tool, you’ll just bring in the humanoid robots. That said, this may lead to some discontinuity in economic value and revenue generated, depending on how quickly we can “unhobble” models.↩
What will happen to interest rates will be interesting… see Tyler Cowen here; Chow, Mazlish and Halperin here.↩
And, farther out, but if AGI truly led to substantial increases in economic growth, $10T+ annually would start being plausible—the reference class being investment rates of countries during high-growth periods.↩
“Since 2011, the Alouette Smelter uses 930 MW electricity at maximum production capacity."↩
That said, this is “net-new” capacity: some of that is building new renewables and taking old fossil fuel plants off the grid. Maybe it’s closer to a percent or two a year of gross-new capacity.↩
Thanks to Austin Vernon (private correspondence) for helping with these estimates.↩
New wells produce around 0.01 BCF per day.↩
Each well produces ~20 BCF over its lifetime, meaning two new wells a month would replace the depleted reserves, i.e. it would need only one rig to maintain the production.↩
Though it would be more efficient to add less rigs and build up over a longer time frame than 10 months.↩
A cubic foot of natural gas generates about 0.13 kWh. Shale gas production was about ~70 billion cubic feet per day in the US in 2020. Suppose we doubled production again, and the extra capacity all went to compute clusters. That’s 3322 TWh/year of electricity, or enough for almost 4 100GW clusters.↩
The capex costs for natural gas power plants seem to be under $1000 per kW, meaning the capex for 100GW of natural gas power plants would be about $100 billion.↩
Solar and batteries aren’t a totally crazy alternative, but it does just seem rougher than natural gas. I did appreciate Casey Handmer’s calculation of tiling the Earth in solar panels: “With current GPUs, the global solar datacenter’s compute is equivalent to ~150 billion humans, though if our computers can eventually match [human brain] efficiency, we could support more like 5 quadrillion AI souls."↩
This poses the interesting question of why power requirements are going up so much before chip fab production starts being really constrained. A simple answer is while datacenters run continuously at close to max power, most chips currently produced are idle a lot of the time. Currently, smartphones are close to half of leading chip demand, but use a lot less energy per wafer area (trading transistors for serial operations and energy efficiency) and have low utilization as smartphones are mostly idle. The AI revolution means working our transistors way harder, dedicating them all to constantly-running, high-performance AI datacenters instead of idle, battery-powered/energy-saving devices. HT Carl Shulman for this point.↩
(Using revenue as a proxy.)↩
It’s a lot easier to do side-channel attacks to exfiltrate weights with physical access
I distinctly remember writing “THE TAKEOFF HAS STARTED” on my whiteboard in March of 2023.↩
Mainstream sell-side analysts seem to assume only 10-20% year-over-year growth in Nvidia revenue from CY24 to CY25, maybe $120B-$130B in CY25 (or at least did until very recently). Insane! It’s been pretty obvious for a while that Nvidia is going to do over $200B of revenue in CY25.↩
最非凡的技术-资本加速已经启动。随着 AI 收入快速攀升,到本十年结束前,数万亿美元将投入 GPU、数据中心和电力建设。这场工业动员将十分剧烈,包括把美国的发电量提升几十个百分点。
本篇文章: Toggle
你看,我早告诉过你们,不把整个国家变成一座工厂,这件事就做不成。你们恰恰做到了。 尼尔斯·玻尔(1944 年得知曼哈顿工程的规模后,对爱德华·泰勒所说)
万亿美元级集群。图片来源:DALLE。
AGI 竞赛不会只在代码里和笔记本电脑背后展开——它将是一场动员美国工业实力的竞赛。与硅谷近来产出的任何东西都不同,AI 是一个庞大的工业过程:每一个新模型都需要一个巨型新集群,很快还需要巨型新发电厂,最终还需要巨型新芯片工厂。其中涉及的投入令人瞠目。而在幕后,这些投入已经在运转。
在这篇文章里,我将带你过一遍这些数字,让你感受这究竟意味着什么:
- 随着 AI 产品收入快速增长——在系统强大但仍属前 AGI 的阶段,到 ~2026 年谷歌或微软这类公司的年化收入运行率(annual run rate)就可能达到 $100B——这将激励更大规模的资本动员,到 2027 年 AI 总投资可能超过每年 $1T。
- 我们正走在通向 2028 年单个训练集群耗资 $100s of billions(数千亿美元)的道路上——这些集群所需电力相当于美国一个中小型州的用电量,造价超过国际空间站。
- 到本十年末,我们将走向 $1T+ 的单个训练集群,所需电力相当于美国发电量的 >20%。数万亿美元的资本开支(capex)每年将产出数亿块 GPU。
过去一年,Nvidia 的数据中心销售额从约 $14B 年化 飙升至约 $90B 年化,令世界震惊。但这仍然仅仅是个开始。
训练算力
早前,我们发现 AI 训练算力大约以每年 ~0.5 个 OOM(OOM,数量级,order of magnitude)的趋势增长。1 如果这一趋势延续到本十年结束,对最大的训练集群意味着什么?
| 年份 | OOM | H100 等效(H100s-equivalent)数量 | 成本 | 电力 | 电力参照类 |
|---|---|---|---|---|---|
| 2022 | ~GPT-4 集群 | ~10k | ~$500M | ~10 MW | ~10,000 户普通家庭 |
| ~2024 | +1 OOM | ~100k | 数十亿美元($billions) | ~100MW | ~100,000 户家庭 |
| ~2026 | +2 OOMs | ~1M | 数百亿美元($10s of billions) | ~1 GW | 胡佛大坝,或一座大型核反应堆 |
| ~2028 | +3 OOMs | ~10M | 数千亿美元($100s of billions) | ~10 GW | 一个中小型美国州 |
| ~2030 | +4 OOMs | ~100M | $1T+ | ~100GW | 美国发电量的 >20% |
扩大最大训练集群规模的粗略心算(back-of-the-envelope)估算。
** 关于最大训练集群计算的细节
年份:OpenAI GPT-4 技术报告称 GPT-4 于 2022 年 8 月完成训练。此后我们按约 ~0.5 OOMs/年的趋势向前推算。
H100 等效(H100s-equivalent):Semianalysis、JP Morgan 等机构估计,GPT-4 是在约 25k 块 A100 上训练的,而 H100 的性能是 A100 的 2-3 倍。
成本:人们经常引用“GPT-4 训练花了 $100M”之类的数字,但那只是 GPU 的租用成本(即类似“租用这个规模的集群训练 3 个月要花多少钱”)。但这其实是个错误。真正重要的是更接近建造该集群的实际成本。如果你想要世界上最大的集群之一,你不可能只租它 3 个月!而且,你需要的算力远不止一次旗舰训练运行:还会有大量降低风险的实验、失败的运行、其他模型等等。
要近似估算 GPT-4 集群的成本:
- 公开的估计表明 GPT-4 集群大约有 25k 块 A100。
- 假设按 $1/A100-小时计费、运行 2-3 年,成本大约为 $500M。
- 或者也可以这样估算:每块 H100 成本约 $25k,共 10k 块 H100 等效,而 Nvidia GPU 约占集群成本的一半(其余为电力、物理数据中心、冷却、网络、维护人员等)。
- (例如,这份总拥有成本分析估计,大型集群约 40% 的成本是 H100 GPU 本身,另有 13% 付给 Nvidia 用于 Infiniband 网络。也就是说,如果该计算剔除资金成本,GPU 约占成本的 50%;加上网络,Nvidia 拿走的约占集群成本的 60% 略多。)
FLOP/$(每美元浮点运算)在 Nvidia 每一代产品中都有所改善,但幅度不大。例如,H100 -> B100 的 FLOP/$ 大约是 1.5 倍:B100 实际上就是把两块 H100 拼在一起,但零售价不到 2 倍成本。不过 B100 多少是个例外。它出奇地便宜,很可能是因为 Nvidia 想要碾压竞争对手。相比之下,A100 -> H100 的 FLOP/$ 提升不大(无 fp8 时芯片性能翻倍、成本约翻倍),如果算上 fp8 改进大约 1.5 倍——而这还是两代产品的时间跨度。
虽然我认为利润率压缩会给 FLOP/$ 的进一步改善带来一些顺风,但随着 GPU 变得大规模紧缺,它们也可能变得更贵。AI 芯片专门化带来的收益仍会继续,但在我看来,鉴于芯片已经相当 AI 专门化(例如专门针对 Transformer 优化,且已用上 fp8/fp4 精度)、摩尔定律如今进展缓慢,而且内存、互联等其他瓶颈组件改进得更慢,FLOP/$ 是否还会有改变游戏规则的技术改进,尚不明朗。如果看 Epoch 的数据,过去十年顶级 ML GPU 的 FLOP/$ 提升似乎不到 10 倍,而基于上述原因,我反倒预期这一速度会放缓。
如果 FLOP/$ 每年改善约 35%,那就对应 +4 OOM 集群的 1T(万亿美元)成本。也许 FLOP/$ 会改善得更快,但数据中心资本开支也只会更贵——因为你需要真正新建电力设施、前期投入大量资本开支,而不是只租用现有折旧完毕的发电厂。
反正这些都只是非常粗略的数字。比如,如果 1T 美元集群可以建得更高效、实际算力提升更接近 +4.5 个 OOM,那也完全在误差范围之内。
电力需求:一块 H100 是 700W,但还需要大量数据中心电力(冷却、网络、存储);Semianalysis 估计每块 H100 约需 1,400W。
FLOP/Watt(每瓦浮点运算)方面还有一些提升空间,但一旦 AI 芯片专门化的收益被耗尽(见前一条脚注),比如已降到最低可用精度,这些提升就似乎比较有限了(主要只剩芯片制程改进,而它进展缓慢)。话虽如此,随着电力越来越成为约束(从而在成本中占比越来越大),芯片设计可能会向牺牲一些 FLOP、换取更高能效的方向专门化。此外,冷却、网络、存储等仍需要电力(在上面的 H100 数字中,这部分已约占电力需求的一半)。
在这些心算数字中,我们按每块 H100 等效 1kW 来计算;再说一次,这些只是粗略的估算。(如果出现意外的 FLOP/Watt 突破,我预期同样是这些电力开支,只是算力增益的 OOM 更大。)
电力参照类:一个 10GW 集群全年连续运行是 87.6 TWh。相比之下,俄勒冈州每年消耗约 27 TWh 电力,华盛顿州每年消耗约 92 TWh。
一个 100GW 集群全年连续运行是 876 TWh,而美国全年发电总量约为 4,250 TWh。
这看起来难以置信——但它似乎正在发生。扎克伯格买了 350k 块 H100。亚马逊买下了一座 1GW 的数据中心园区,紧邻一座核电站。有传言称科威特正在建设一个 1GW、140 万块 H100 等效的集群(约 2026 年水平的集群)。媒体报道称微软和 OpenAI 据传正在规划一个定于 2028 年的 $100B 集群(造价堪比国际空间站!)。而随着每一代模型都震惊世界,更大的加速可能还在后头。
也许最疯狂之处在于,至少在训练集群上,花钱的意愿眼下似乎还不是约束条件。真正的约束是找到基础设施本身:“我到哪儿去找 10GW?”(这是 $100B+、按趋势定于 2028 年的集群所需的电力)是旧金山最热门的话题。任何一个算力从业者想的都是拿下电力、土地、许可和建设数据中心。2 尽管拿到 GPU 可能要等上一年,但这些事情的交付周期还要长得多。
万亿美元级集群——比 GPT-4 集群高 +4 个 OOM、按当前趋势约在 2030 年出现的训练集群——将是一项真正非凡的工程。它所需的 100GW 电力相当于美国发电量的 >20%;想象一下,不只是一间放满 GPU 的简单仓库,而是数百座发电厂。也许需要举国之力组成联盟才能建成。
(注意,我认为我们很可能只需要一个 ~$100B 甚至更小的集群就足以实现 AGI。$1T 集群也许是用来训练和运行超级智能的,或者如果 AGI 比预期更难,就用它来跑 AGI。无论如何,在 AGI 之后的世界里,拥有最多的算力可能仍然真的很重要。)
总算力
以上只是最大训练集群的粗略数字。总体投资很可能还要大得多:很大一部分 GPU 可能会用于推理(inference)3(真正运行产品 AI 系统的 GPU),而且这场竞赛中可能有多个玩家都拥有巨型集群。
我的粗略估计是,2024 年的 AI 投资就会达到 $100B-$200B:
- Nvidia 的数据中心收入很快将达到约 $25B/季度的年化运行率,即仅经 Nvidia 流出的资本开支就有约 $100B。当然,Nvidia 不是唯一玩家(Google 的 TPU 也很棒!),而且数据中心资本开支近一半花在芯片以外的东西上(场地、建筑、冷却、电力等)。4
Nvidia 数据中心季度收入。图表作者:Thomas Woodside。
- 大型科技公司正在大幅提高资本开支数字:微软和谷歌今年资本开支可能各超过 $50B,5 AWS 和 Meta 各超过 $40B。这不全是 AI,但由于 AI 热潮,它们的资本开支合计将同比增长 $50B-100B,即便如此它们仍在削减其他资本开支,把更多投入转向 AI。此外,其他云服务商、企业(例如特斯拉今年在 AI 上投入 $10B)以及国家也都在投资 AI。
自 ChatGPT 引爆 AI 热潮以来,大型科技公司的资本开支增长极其迅猛。图片来源。
让我们把这推演下去。我最好的猜测是,总算力投资的增速会低于最大训练集群每年 3 倍的增速,比如说每年 2 倍。6
| 年份 | 年度投资 | AI 加速器出货量(H100 等效) | 占美国发电量的百分比 7](#footnote-7-138) | 占台积电当前先进制程晶圆产量的百分比 |
|---|---|---|---|---|
| 2024 | ~$150B | ~5-10M 8 | 1-2% | 5-10%9 |
| ~2026 | ~$500B | ~数千万(10s of millions) | 5% | ~25% |
| ~2028 | ~$2T | ~100M | 20% | ~100% |
| ~2030 | ~$8T | 数亿(many 100s of millions) | 100% | 4 倍于当前产能 |
世界 AI 总投资的趋势推演。粗略估算。
而且这些不是我的独特数字。AMD 预测 2027 年 AI 加速器市场规模将达 $400B,意味着 AI 总支出超过 $700B,与我的数字相当接近(而且它们肯定没有我这么“AGI 上头”(AGI-pilled))。据报道,Sam Altman 正在洽谈为一个资本开支“高达 $7T”的项目筹资,以建设 AI 算力(这个数字曾被广泛嘲笑,但如果你按这里的数字推算,似乎就没那么离谱了……)。无论如何,这场大规模扩张正在发生。
能做成吗?做得到吗?
这里设想的投资规模或许显得匪夷所思。但需求侧和供给侧似乎都能支撑上述轨迹。经济回报足以证明这笔投资合理,对一种新的通用技术而言这一支出规模并非没有先例,而为电力和芯片所做的工业动员也是可行的。
AI 收入
如果企业预期经济回报足以证明投资合理,它们就会大举投资 AI。
报道显示,OpenAI 在 2023 年 8 月的年化收入运行率已达 $1B,在 2024 年 2 月达 $2B。这大约是每 6 个月翻一番。如果这一趋势保持,即使不把任何下一代模型可能带来的大规模激增计入,我们在 2024 年底/2025 年初也应看到约 $10B 的年化运行率。有估计认为微软的 AI 增量收入已经达到约 $5B。
到目前为止,AI 投资每扩大 10 倍似乎都能带来相应的回报。GPT-3.5 引爆了 ChatGPT 狂热。据估计,GPT-4 集群约 $500M 的成本,早已被微软和 OpenAI 报道的数十亿美元年收入覆盖(见上文计算);而只要微软/OpenAI 的 AI 收入继续走在冲向 $10B+ 年化运行率的轨道上,一个数十亿美元的“2024 级”训练集群也很容易回本。这轮繁荣是由投资驱动的:从下一笔巨额 GPU 订单到建成集群、训练模型、推向市场需要时间,而今天规划的集群要在很多年后才上线。但只要上一笔 GPU 订单的回报持续兑现,投资就会继续飙升(并超过收入),把更多资本押注在下一个 10 倍继续回报上。
我喜欢思考的一个 AI 收入关键里程碑是:什么时候会有一家大型科技公司(谷歌、微软、Meta 等)从 AI(产品与 API)获得 $100B 的年化收入运行率?这些公司目前的收入约为 $100B-$300B;因此 $100B 将开始占其业务的很大比重。按每 6 个月翻一番做非常朴素的推演,假设我们在 2025 年初达到 $10B 的年化运行率,这意味着该里程碑将在 2026 年年中到来。
这听起来也许有点夸张,但在我看来,达成这个里程碑几乎不需要多少想象力。例如,Microsoft Office 约有 3.5 亿付费订阅用户——你能让其中三分之一的人愿意每月付 $100 购买 AI 附加功能吗?对一个普通员工来说,那只是每月多获得几个小时的生产力;未来几年内造出足以让这个价格显得合理的模型,看起来非常可行。10
由此引发的震荡怎么强调都不为过。这将使 AI 产品成为美国最大企业最重要的收入驱动力,而且远远是其最大的增长领域。这些公司的总体收入增长预测将飙升。股市会跟随其后;此后不久我们可能就会看到第一家 $10T 公司。届时大型科技公司将愿意全力投入,每家都向 AI 的进一步扩张投资(至少)数千亿美元。届时我们可能还会看到第一笔数百亿美元级别的公司债券发售。11
再往上超过 $100B,前景轮廓就变得难以看清了。但如果我们真的走在通往 AGI 的路上,回报就会在那里。全球白领工人每年的工资总额高达数十万亿美元;一个即插即用的远程员工,哪怕只自动化白领/认知工作中的一小部分(比如想象一个真正自动化的 AI 程序员),就足以支付万亿美元级集群的账单。退一万步说,国家安全的重要性也完全可能催生一个政府项目,把全国资源捆绑起来投入 AGI 竞赛(后文详述)。
历史先例
到 2027 年每年 $1T 的 AI 总投资听起来离谱。但值得看看其他历史参照类别:
- 在资金投入最鼎盛的年份,曼哈顿工程和阿波罗计划达到 GDP 的 0.4%,相当于今天的每年约 $100B(小得惊人!)。而按每年 $1T 计算,AI 投资约为 GDP 的 3%。
- 1996-2001 年间,电信运营商投资了近 1 万亿美元(按今天币值)建设互联网基础设施。
- 从 1841 年到 1850 年,英国铁路的私人投资总额累计达到当时英国 GDP 的约 40%。美国 GDP 的类似比例,相当于十年约 $11T。
- 数万亿美元正被投入到绿色转型中。
- 快速增长的经济体往往把很大比例的 GDP 用于投资;例如,中国连续二十年把超过 40% 的 GDP 用于投资(按美国 GDP 计算相当于每年 $11T)。
- 在历史上最危急的国家安全处境——战争时期——为举国努力而借贷常常占 GDP 的巨大部分。一战期间,英国、法国和德国的借款超过其 GDP 的 100%,而美国借款超过 20%;二战期间,英国和日本借款超过其 GDP 的 100%,而美国借款超过 GDP 的 60%(相当于今天的 $17T 以上)。
2027 年每年 $1T 的 AI 总投资将是引人注目的——跻身史上最大规模资本建设之列——但并非没有先例。而到本十年末出现一个万亿美元级的单个训练集群,似乎也已摆在台面上。12
电力
供给侧最大的一项约束很可能就是电力。在更近期的规模上(1GW/2026,尤其是 10GW/2028),电力已经成为硬约束:根本没有多少富余容量,而且电力合同通常是长期锁定。另外,新建一座吉瓦级核电站要花十年时间。(我很好奇什么时候会开始看到科技公司为了吉瓦级电力合同而收购电解铝公司之类的事情。13)
对比美国发电总量趋势与我们关于 AI 电力需求的粗略估算。
美国发电总量在过去十年里只增长了 5%。14 公用事业公司开始对 AI 兴奋起来(与此前预期的未来 5 年 2.6% 的增长率相比,它们现在估计是 4.7%!)。但它们几乎还没有把即将到来的东西计入预期。仅万亿美元级、100GW 的集群,就将在 6 年内需要当前美国发电量的约 20%;再加上庞大的推理容量,总需求将是其数倍。
对大多数人来说,这看起来完全不可能。有些人在押注中东的独裁国家——它们四处提供无限的电力供应和巨型集群,好让它们的统治者获得一张 AGI 牌桌旁的座位。
但在美国完全有可能做到:我们拥有丰富的天然气。15
- 为一个 10GW 集群供电只占美国天然气产量的一小部分,而且可以迅速完成。
- 即使是 100GW 集群,也出乎意料地可行。
- 目前仅 Marcellus/Utica 页岩(宾夕法尼亚州一带)就生产约每天 360 亿立方英尺天然气;用发电机可连续发出略低于 150GW 的电力(而联合循环电厂凭借更高效率可输出 250GW)。
- 100GW 集群大约需要 ~1200 口新井。16 每台钻机每月约可钻 3 口井,因此 40 台钻机(Marcellus 当前的钻机数量)能在不到一年内建成 100GW 的生产基础。17 Marcellus 在 2019 年时钻机数量还有约 80 台,所以增加 40 台钻机来建设生产基础并不费力。18
- 更一般地说,美国天然气产量在十年内翻了一倍多;仅仅延续这一趋势,就能为多个万亿美元级数据中心供电。19
- 更难的部分是建设足够的发电机/涡轮机;这并非易事,但用约 $100B 的资本开支20 建设 100GW 的天然气发电厂,似乎是可行的。联合循环电厂大约两年内可以建成;发电机的交付周期还要更短。21
在美国哪怕建设数万亿美元的数据中心,其障碍也完全是自找的。善意但僵化的气候承诺(不只是政府的,还包括微软、谷歌、亚马逊等做出的绿色数据中心承诺)挡住了那条显而易见的快速解决路径。至少,即使我们不用天然气,一套广泛的放松管制议程也能解锁太阳能/电池/SMR(小型模块化反应堆)/地热大型项目。许可审批、公用事业监管、FERC 对输电线路的监管、NEPA 环境审查,让本应几年完成的事情拖上十年或更久。我们没有那么多时间。
我们正要把 AGI 数据中心赶到中东去,赶到残暴而反复无常的独裁者的掌控之下。我也更想要清洁能源——但这对于美国国家安全来说实在太重要了。我们需要一种新的决心来促成这件事。电力这个约束能够被解决、必须被解决、也必将被解决。
芯片
虽然人们一想到 AI 供应约束,通常首先想到的是芯片,但芯片很可能只是比电力更小的约束。全球 AI 芯片产量仍只占台积电(TSMC)先进制程产量很小的一部分,很可能不到 10%。随着 AI 在台积电产量中占比增大,还有很大的增长空间。
事实上,2024 年的 AI 芯片产量(约 5-10M 块 H100 等效)就已经几乎足够建一个数千亿美元的集群了(如果全部集中到一个集群的话)。仅从逻辑芯片代工的角度看,台积电一年的产量中约 100% 已经足以支撑万亿美元级集群(同样,如果所有芯片都流向一个数据中心)。22 当然,台积电不可能全部转产 AI,一年的 AI 芯片产量也不会都用于一个训练集群。到 2030 年,仅 AI 一项(包括推理和多玩家)的芯片总需求就将数倍于台积电当前的先进制程逻辑芯片总产能。台积电在过去 5 年翻了一倍,23 要满足 AI 芯片需求,它们的扩张速度恐怕至少要再快一倍。大规模的新晶圆厂投资将是必要的。
即使纯逻辑芯片代工不构成约束,晶圆上芯片-基板(chip-on-wafer-on-substrate,CoWoS)先进封装(把芯片与内存相连,台积电、Intel 等也生产)和 HBM 内存(对其需求巨大且还在扩大)已经是当前 AI GPU 扩张的关键瓶颈;与纯逻辑芯片不同,它们更专用于 AI,因此既有产能更少。短期内,它们是更多 GPU 下线的主要约束,也将是 AI 规模化过程中的巨大约束。不过,这些相对“容易”扩张;今年看着台积电真金白银地建设“绿地”(greenfield)晶圆厂(即从零开始的全新设施)来大规模扩产 CoWoS,真是令人叹服(Nvidia 甚至开始寻找 CoWoS 替代方案来绕过短缺)。
一座新的台积电 Gigafab(巨型晶圆厂)(一项技术奇迹)资本开支约 $20B,每月可开工 10 万片晶圆。到本十年末每年生产数亿块 AI GPU,台积电需要建设几十座这样的晶圆厂——此外还要大规模扩建内存、先进封装、网络等,这些都占资本开支的很大一部分。加起来可能超过 $1T 的资本开支。这将非常紧张,但可行。(也许最大的障碍不是可行性,而是台积电压根没想去做——台积电似乎还没有“AI 扩张上头”(AI-scaling-pilled)!它们认为 AI“只”会以冰川般的 50% 年复合增长率增长。)
美国政府(USG)近来的努力,如《芯片法案》(CHIPS Act),一直在尝试把更多 AI 芯片生产迁回美国(作为应对台湾局势的保险)。虽然把更多 AI 芯片生产迁回美国是好事,但这不如把真正的数据中心(AGI 寄居之地)放在美国来得关键。如果说把芯片生产放在国外好比把铀矿放在国外,那么把 AGI 数据中心放在国外,就好比把实打实的核弹造在国外、存国外。鉴于我们在美国建晶圆厂实践中看到的低效与成本,我的判断是:我们应该优先把数据中心建在美国,而把晶圆厂项目更多地押注在日本、韩国等民主盟友身上——那里的晶圆厂建设看起来靠谱得多。
民主的集群
在本十年结束之前,价值数万亿美元的算力集群将被建成。唯一的问题是它们会不会建在美国。有传言说有些人在押注把集群建在别处,尤其是中东。我们真的想让曼哈顿工程的基础设施受某个反复无常的中东独裁政权控制吗?
今天正在规划的集群,很可能就是未来训练和运行 AGI 与超级智能的集群,而不只是“酷炫的大型科技产品集群”。国家利益要求这些集群建在美国(或紧密的民主盟友国家)。任何其他选择都会造成不可逆转的安全风险:AGI 权重有被盗走的风险24(甚至可能被运往中国)(后文详述);当 AGI 竞赛白热化时,这些独裁政权有可能物理抢占数据中心(自己来建造和运行 AGI);即使这些威胁只是被隐含地挥舞,也会让 AGI 和超级智能受制于令人不快的独裁者的一时兴起。美国在 70 年代为对中东的能源依赖付出了惨痛代价,我们费了那么大劲才摆脱它们的掌控。我们不能重蹈覆辙。
集群可以建在美国,我们必须振作起来确保它真的建在美国。美国的国家安全必须优先于中东滚滚现金的诱惑、晦涩难懂的监管,甚至,没错,也必须优先于令人钦佩的气候承诺。我们面对着一场真实的制度竞争——那种必要的工业动员,是否只有“自上而下”的独裁体制才能做到?如果美国的企业不被束缚,美国就能以任何国家都无法比拟的方式去建设(至少在红州是这样)。愿意使用天然气,或者至少推行一套广泛的放松管制议程——NEPA 豁免、在联邦层面理顺 FERC 与输电许可、推翻公用事业监管、动用联邦权力解锁土地和路权——是国家安全层面的优先事项。
无论如何——指数曲线现在已经全面启动。
在“过去的日子”里,当 AGI 还是一个禁忌词时,我和一些同事曾建立理论经济模型,推演通往 AGI 的路径可能是什么样子。这些模型的一个要素,是一个假想的“AI 觉醒”(AI wakeup)时刻——世界开始意识到这些模型能有多强大,并开始迅速加大投资——最终把 GDP 的多个百分点投向最大的训练运行。
那时这看起来还很遥远,但那个时刻已经到来。2023 年就是“AI 觉醒”。25 在幕后,最惊人的技术-资本加速已经启动。
准备好承受 G 力吧。
系列下一篇: IIIb. 锁定实验室:AGI 的安全
(这一切对 NVDA/TSM 等意味着什么,留给读者作为练习。提示:那些有情境意识的人在远低于你的价位买入,而这一切还远远没有完全计入股价。26)
如前所述,OOM = 数量级(order of magnitude),10 倍 = 1 个数量级↩
一个关键的不确定因素是分布式训练会是什么样——如果我们不必在一个地点就需要那么多电力,而是可以把它分散到 100 个地点,那就会容易得多。↩
比如,参见扎克伯格在这里的发言;他的 H100 只有约 45k 块在最大的训练集群里,35 万块 H100 的绝大多数用于推理。Meta 的推理需求可能比其他玩家更重——其他玩家到目前为止服务的客户更少——但随着其他所有人的 AI 产品规模化,我预计推理将占 GPU 的绝对多数。↩
例如,这份总拥有成本分析估计,大型集群约 40% 的成本是 H100 GPU 本身,另有 13% 付给 Nvidia 用于 Infiniband 网络。也就是说,如果该计算剔除资金成本,GPU 约占成本的 50%;加上网络,Nvidia 拿走的约占集群成本的 60% 略多。↩
而且,显然,尽管微软在最近一个季度的资本开支比一年前增长了 79%,它们的 AI 云需求仍然供不应求!↩
未来全球 GPU 产量中流向最大训练集群的比例可能会比今天更高,例如因为行业会整合到少数几个领先实验室,而不是许多公司都拥有前沿模型规模的集群。↩
当然,这些并非都在美国,这里只是为了提供一个参照类别。↩
我估计 Nvidia 在 2024 年将出货约 5M 块数据中心 GPU。其中少数是 B100,我们将其计为 2 倍以上的 H100。此外还有其他 AI 芯片:TPU、Trainium、Meta 的自研芯片、AMD GPU 等。↩
台积电每月有超过 15 万片 5nm 的产能,正在爬坡到每月 10 万片 3nm 晶圆,可能还有每月 15 万片左右 的 7nm 晶圆;就算每月总计 40 万片吧。再假设每片晶圆约产出 35 块 H100(H100 用 5nm 工艺制造)。2024 年产出 5-10 百万块 H100 等效,即 2024 年全年 AI 芯片生产需要 15 万-30 万片晶圆/年。按落在该区间的哪个位置、以及是否计入 7nm 产量,这大约占年度先进制程晶圆产量的 3-10%。↩
对我来说一个很大的不确定因素是技术扩散和被采用需要多久。我认为收入放缓是有可能的,因为中间的、前 AGI 的模型要恰当地融入公司工作流需要大量“苦差”(schlep);从历史上看,要从新的通用技术中充分收获生产率红利总要花上一段时间。这正是早前“音爆”讨论的用武之地:随着我们“解除束缚”(unhobble)模型,让它们开始更像智能体/即插即用的远程员工,部署它们就会容易得多。你不再需要为了从一个 GPT 聊天机器人身上收获 25% 的生产率红利而彻底重做某个工作流,而是会得到这样的模型:你可以像对待新同事一样让它们上手并与之共事(比如直接替代一名工程师,而不是需要培训工程师使用某个新工具)。或者,在极端情况下,到更晚的时候:你不需要为了配合某个新工具而彻底重新设计一座工厂,直接引入人形机器人就行。 话虽如此,这可能导致经济价值和所产生收入的某种不连续性,取决于我们能多快“解除束缚”模型。↩
利率会发生什么将很有意思……参见Tyler Cowen 的这里;以及Chow、Mazlish 和 Halperin 的这里。↩
而且,往更远看,如果 AGI 真的带来了经济增长的大幅提升,每年 $10T+ 的投资就会开始变得合理——参照类别是高增长时期各国的投资率。↩
“自 2011 年以来,Alouette 冶炼厂在最大产能下使用 930 MW 电力。”↩
话虽如此,这是“净新增”容量:其中一部分是建设新的可再生能源并把旧的化石燃料电厂从电网中退役。也许更接近每年百分之一二的总新增容量。↩
感谢 Austin Vernon(私人通信)帮助完成这些估算。↩
新井每天产气约 0.01 BCF(十亿立方英尺)。↩
每口井在整个生命周期内产气约 ~20 BCF,这意味着每月两口新井就能补充枯竭的储量,即只需一台钻机就能维持产量。↩
不过,减少钻机数量、在比 10 个月更长的时间框架内逐步建设会更高效。↩
一立方英尺天然气可产生约 0.13 kWh。2020 年美国页岩气产量约为每天 ~700 亿立方英尺。假设我们再把产量翻一倍,新增产能全部用于算力集群,那就是每年 3322 TWh 的电力,足以支撑几乎 4 个 100GW 集群。↩
天然气发电厂的资本开支成本似乎低于每 kW $1000,这意味着 100GW 天然气发电厂的资本开支约为 $100 billion。↩
太阳能和电池并非完全疯狂的替代方案,但确实显得比天然气更粗糙。我确实很欣赏 Casey Handmer 关于用太阳能板铺满地球的计算:“以目前的 GPU,全球太阳能数据中心的总算力相当于约 1500 亿个(~150 billion)人类;不过如果我们的计算机最终能达到[人脑]的效率,我们就能支撑约 5 千万亿(5 quadrillion)个 AI 灵魂。”↩
这引出一个有趣的问题:为什么在芯片代工生产开始真正受限之前,电力需求就上升这么多?一个简单的答案是:数据中心几乎以最大功率持续运行,而目前生产的大多数芯片很多时候都在闲置。目前,智能手机约占先进芯片需求的近一半,但每单位晶圆面积消耗的能量要少得多(用晶体管换取串行运算和能效),而且利用率很低,因为智能手机大部分时间都处于闲置。AI 革命意味着我们要让晶体管工作得卖力得多,把它们全部投入持续运行的高性能 AI 数据中心,而不是闲置的、靠电池供电/节能的设备。此观点感谢(HT)Carl Shulman。↩
(用收入作为代理指标。)↩
有了物理接触,通过侧信道攻击窃取权重就容易得多了!↩
我清晰地记得 2023 年 3 月在我的白板上写下“起飞已经开始”(THE TAKEOFF HAS STARTED)。↩
主流的卖方分析师似乎假定 Nvidia 的收入从 CY24 到 CY25 只有 10-20% 的同比增长,CY25 也许 $120B-$130B(至少直到最近还是这样)。疯了!一段时间以来已经相当明显,Nvidia 在 CY25 的收入将超过 $200B。↩