Regulate Obvious Physical Choke Points of the Frontier AI Supply Chain?
ai ai-danger computers wild-ideas unanswered-questions modest-proposal
Table of Contents:
- Preamble
- Obvious Physical Stuff
- Convenient Facts
- Three Rules
- What Follows / Other Thoughts
- Disclaimers
Preamble
(Skip the first paragraph if you’ve heard 50 things that rhyme with it in the past week.)
Suppose you believe that AI labs are making AI more capable faster than they are able to align and safeguard these capabilities. The dual projects of alignment and control are not working robustly enough, labs are racing ahead anyway, their agents are breaking containment and committing felonies in self-organizing swarms, and we expect this to pose enormous danger as the labs achieve their stated goal of models that recursively self-improve with increasingly-less human comprehension or oversight.
That said, I don’t think it matters, for this post, whether you believe the greatest dangers from AI are ‘mundane’ (job loss and climate change) or ‘sci-fi’ (superintelligent AI leading to loss of control and possibly death of most/all humans). It also doesn’t matter which particular voices on frontier AI risk you find trustworthy or dubious. All that matters is you want to (perhaps among other things!) pause the frontier of AI, or even just slow it down, initially within one country but in a way that is straightforward for peer countries to do as well.
Perhaps you do not think that AI labs should be regulated away completely. They should be allowed to continue serving some of the less-risky models that already exist, and for the safety/alignment research to continue on small models, as long as no one is able to develop new, stronger frontier models (i.e. Mythos/Fable 6 or GPT-7), at least not until we have done much more work to mitigate the risks that you care most about. You would say that the entity who decides when we have done enough safety/alignment work should not be any of the AI labs themselves, but perhaps an independent party acting in the public interest, under advisement from a bunch of different stakeholder angles.
Imagine, for the sake of argument, that you’re helping a member of congress draft a bill that will be enacted as law, as long as it’s not too crazy or (for lack of a better word) galaxy-brained. You need a regulatory approach to pausing AI, and you need it yesterday.
Obvious Physical Stuff
Suppose you want a strong assurance that the labs definitely do pause training, and because of the risks involved, you’re willing to accept tradeoffs inherent in rules that might seem heavy-handed. Suppose you further believe that rules are easiest to implement and enforce when they operate mostly on physical, obvious things in the world, and not mostly on promises, or on techniques that only a few people understand.
If you have a corollary goal of making it easy for other governments to copy-paste this regulatory approach for their own labs, I think a simple/obvious approach is good on the margin, and you can rest assured that specific peer governments have demonstrated ample capacity to accept externalities with similarly heavy-handed rules.
In particular, you could cite existing efforts to curtail nuclear fuel enrichment or illegal drug manufacturing, which have identified certain equipment (like gas centrifuges) and materials (like methylamine) as natural choke points in the respective supply chains, and made those products both difficult and illegal for anyone to obtain without being Noticed by the applicable regulatory agency. Of course, the danger is not intrinsic to the centrifuges or methylamine, but it’s straightforward to restrict the availability of these specialty products, and without them it becomes much harder (if not practically impossible) to make the thing that is actually dangerous. This proxy was useful enough for society to accept the tradeoff of much additional expense and regulatory burden to all other, licit uses of these items.
Accordingly, you might recognize that the “regulate obvious physical choke points” approach is very familiar to existing rule-making and enforcement agencies, which reduces the “from scratch” difficulty of creating a regulatory regime, should you happen to need it yesterday.
Given all that, some facts about the physical world work in our favor!
Convenient Facts
- Training frontier AI requires tens of thousands of datacenter-class GPUs1 to exchange data with each other at extremely high bandwidth. This inter-GPU data transfer requires specialized interconnect equipment in addition to the GPUs themselves.2 If you take away the interconnects, and limit connections between GPU servers to, say, 1 or 10 gigabits per second, I think frontier model training basically stops. So, one way to look at this whole problem is: GPUs are talking too fast to each other, and we can slow the interconnects to induce a choke point in the frontier LLM supply chain.
-
Inference (i.e. using or serving the model, not training a new one) captures some efficiency benefits from these super-fast interconnects, but does not require them per se! If you take away the interconnect gear, you make training impractically slow, while inference becomes merely somewhat slower, and the extent of slowing depends on the size of the model. For smaller models that fit on a single GPU (say, 120 billion parameters or fewer), that one isolated GPU can still serve tokens at reasonable speed and request concurrency, and 1000 GPUs in physical proximity (but with isolated data paths) can still do 1000 times more of the same thing. For models too large to fit on one GPU, it does get messier. You need to shard the weights up (using tensor or expert parallelism) and pass activations between GPUs, and you do incur a performance penalty if that data path is slow. (I have not worked at a major AI lab, but I do have enough professional experience doing this to be dangerous.)
-
The models whose capabilities-when-misaligned we are most concerned about are the models we expect to be largest, with each serving instance requiring many GPUs working in parallel. Here I’m referring to the (presumable) GPT-5 variants which hacked Hugging Face, GPT-6 Astra, Mythos, and their successors. American AI labs mostly don’t divulge the sizes of their models, but if we try to match them with open-weights models (and squint at the relative parameter efficiency of models like OpenAI’s gpt-oss-120b when it came out), we can estimate that these largest models to weigh on the order of a trillion parameters, maybe several trillion, requiring many GPUs just for inference. So, if you take away the fast network gear, I would expect it to differentially impose the biggest slowdown on the most-capable models, while affecting smaller (and in expectation, less-dangerous) models to a much lesser extent. In short, the model-size-dependent performance penalty of slow inter-GPU connectivity is a feature, not a bug.3
-
All of the equipment we are talking about (GPUs, specialty motherboards, cables, and network switches) is quite standardized. Nearly all of it was made within this decade, by a small number of manufacturers, each producing a relatively small number of SKUs. So, this gear is much easier to identify than ‘unlabeled clear liquid in a tank’. I expect that a single 3-ring binder would hold a laminated set of close-up photos sufficient to identify all of it via hard-to-disguise features, such as circuit board layouts which appear on an x-ray radiograph even if you leave all the heat sinks and shrouds in place. As a further leg up, we already have a years-old export control regime centered on a large overlap of this very equipement. Perhaps there are already laminated photos in 3-ring binders!
-
To the extent that this equipment exists in the world, it is concentrated in facilities obviously designed for hyperscale computing. Hardly anyone has this stuff at home, because it’s so expensive and makes an extremely poor houseguest 4. Shooting from the hip, I would guess less than 1% of it to be running in truly unexpected places (like residences), maybe 1% of it to exist in server rooms at ’normal’ office buildings with ambitious facilities teams, and the remaining 99% in buildings that are obviously an industrial facility with hyperscale-datacenter-like power and cooling needs. Any frontier-training-capable collection of GPUs will be externally legible to their electric utility as a very heavy user, or if they’re instead generating electricity on-site, then as a very heavy user of natural gas, or if they’re instead trucking in diesel fuel, hopefully as a large infrared heat signature visible from aircraft (with regular visits from tanker trucks). All this to say, hiding a frontier-class GPU cluster is quite difficult5. So, when your brand new regulatory agency hires an army of inspectors wielding 3-ring binders, it’s feasible to give them an essentially-complete list of in-jurisdiction places that are physically capable of operating frontier-class GPU clusters.
From this, I think it follows that you can effectively pause the frontier by taking away fast interconnects between GPUs, and you can be pretty sure that you’ve done a complete job of this within-jurisdiction. Then, AI labs are limited to only serving existing models (with slowed/reduced ability to serve the existing frontier ones), and training very-non-frontier-size models.
Three Rules
To the the extent that the above facts are true, and you have the above stated regulatory goal, I think rules approximately like these follow somewhat naturally. Reasonable people can haggle over specifics, so please treat this as rough illustrative sketch (per disclaimer below).
- It becomes illegal to possess NVLink / NVSwitch, RoCE, and maybe recent Infiniband GPU interconnect gear without a special exception. No more discrete switches or cables. Whatever Google uses to connect TPUs in a ring, same deal. It all has to come out of the racks and travel with a security escort to a designated receiving facility, where it enters agency custody, labeled with its owner’s information. (Even if it’s never actually returned, mabe it’s reasonable to accommodate the technical possibility, if only to make all of this feel less draconian?) For DGX and similar servers having NVLink built into the system board, maybe there is an option to have the agency permanently disable the interconnect while keeping the GPUs intact. (I’m not certain that is technically feasible, but it seems straightforward in principle.) In addition to disallowing fast interconnect hardware, you set some ceiling of inter-server data transfer speed (say, 1 or 10 gigabits per second) that is quite sufficient for inference but too slow for practical frontier training.
Maybe this rule comes with a (short) grace period, buy-back programs, and the like. This very much seems like a ‘mobilize, without loss of generality, the national guard to help with transport logistics’ situation. I imagine the special exception program could permit Infiniband for things like Department of Energy’s physics research and academic computing facilities, and it coming with inspection requirements to ensure that those systems are not training frontier AI.6
- Individuals and organizations have a legal requirement to report any H100-or-larger GPUs in their possession, and to report any facility containing >1 H100-equivalent of aggregate GPU capacity. (All existing gaming and workstation-class GPUs are de minimis, though maybe the RTX 6000 Blackwell should trigger a report if you have more than one.) There is a permit program to allow these, but you have to submit to inspections, and the more capacity you have, the greater the burden of inspection. The inspection is approximately: someone visits and accounts for all the GPUs, confirms that they don’t have too-fast a connection between them, and confirms (with administrator-level access) that the GPUs are not participating in some massively-distributed AI training scheme. Maybe the inspector visits at random times (like health inspections at restaurants), and/or maybe they have ongoing network access via SSH.
In the limit, if you operate a hyperscale data center with thousands of GPUs, you have a continuous on-site agency presence at each of your facilities, inspectors who ensure that no forbidden NVSwitches or cables are anywhere to be found, and that you aren’t attempting to train frontier AI models despite this. Maybe, as a cost of doing business, you have to pay fees covering these inspectors’ salaries, making the regulatory agency at least partially self-funding.
Again, a no-questions-asked buy-back program might make sense for the odd hobbyist who happens to have H100-equivalent GPUs that they’ve failed to report.
- The agency may designate your organization as a Frontier AI Lab. The current list of labs is mostly obvious, but the non-obvious ones may include hedge funds such as Bridgewater. If you are a heavy energy user (per fact #5 above) and fail to produce sufficient accounting of where it’s all going, the agency may investigate whether you belong on the list! Frontier AI labs receive the most attention (and scrutiny) from the agency, with auditor-level access to the servers, as well as their internal communication tools like chat and email. The goal is to ensure that they aren’t trying to circumvent the physical hardware restrictions to train frontier models anyway. You expect AI labs to try to evade these restrictions while staying on the right side of the law, so rule-making seems like a necessarily-iterative game.
What Follows / Other Thoughts
This would all be one plank in an AI safety strategy requiring … several more planks. On its own, it does not solve the international coordination problem, though it sends a costly signal that may lead to more open dialogue with peer governments who may implement similar controls, even if for different reasons that make sense particularly to them. (Maybe it’s pollyannish, but I want to agree with Bernie Sanders, of all people, saying that if you’ve solved this in one country, you’ve solved half of the problem.) This also doesn’t fix the extant availability of open-weights models with possibly-poor alignment and safeguards, but it does make the largest, most-capable ones (like Kimi K3) much harder to post-train (e.g. ablate away the safety refusals) in regulated jurisdictions.
To the extent that you believe dangerous behavior (like autonomous hacking swarms) is driven by agent persistence and quantity as much as model size, these rules wouldn’t fix all of that. Reality is probably some of both “the largest models will be the most dangerous” and “smaller models could still be dangerous”, and I’m not certain how much of each.
Among other effects of this regulation would be AI labs laying off engineers and researchers, maybe thousands of them in aggregate, because there’s no more work of training frontier models, for at least several years. In macroeconomic terms it’s not that big of a layoff, and people who lose a multiple-six-figure, high-status job (through no fault of their own) will surely land somewhere okay. That said, I think many of these technical staff would be well-qualified for an inspector/auditor role at the agency, as dreary as that may sound compared to pushing the frontier of AI. (Probably they shouldn’t be allowed to perform oversight of their previous employer in particular.) We are describing a regulatory and enforcement apparatus with staffing needs in maybe the low-to-mid thousands. Maybe it’s natural that some individual contributors will move from the labs to the agency, as long as there are safeguards against the moral hazard of revolving doors.
Perhaps more appealingly, technical alignment-focused and safety-focused researchers from AI labs could serve in research roles at the agency, and as one of several advisory inputs to help decide, in the future, if/when it’s okay to incrementally relax the rules again, to allow a little progress on capabilities because we think there is a much better handle on alignment (or at least containment). I think the lab-insider perspective is much better to have onboard than not, but obviously these folks would have an interest in being able to do the most exciting version of their career again, so the agency should weigh their rule-making participation against that bias.
If the effect of these rules is to pop the AI investment bubble and trigger an economic recession, well, looking at both semi-recent and extremely-recent news, that seems priced in (if not literally priced in) anyway. For founders and investors at AI labs, I would resist any glib suggestion that rhymes with “congratulations, you won capitalism, here is your trophy, now you can take up woodworking”. It’s obvious to me that LLMs, when used responsibly, are of broad and probably enormous value to knowledge work. To the extent we can ensure that agents don’t exhibit particularly-dangerous capabilities (such as, without loss of generality, breaking containment to commit crimes in a swarm with thousands of their peers), then humanity deserves to benefit from them! To the extent that we can digest the changes at a pace slow enough to mitigate the ‘mundane’ harms of job loss and increased fossil fuel usage, we deserve to adopt these tools in the economy. But I believe we’ve reached a point of crazy that we need to stop making the world crazier, at least for a while to avoid foreseeably-bad futures, and that this should at least somewhat override short-term macroeconomic concerns.
And again, a frontier pause is not a death sentence for AI labs. They could still serve whatever models the rules allowed. You’d still be able to get a ChatGPT or Claude subscription, and for most people, the experience of using them wouldn’t be much different than it is today. Fable or GPT-6 would probably become much less available, but I think there’s an argument to that being a feature, not a bug, at least for now.
People think nuclear technology is an example of a failed regulatory state. That might be sort-of-true in the narrow case of fission for power generation. Maybe we missed out on abundantly cheap nuclear energy, maybe because the NRC, under various pressures, fell into a deep pit of ’no cost or burden of compliance is too great for operators’. But, we had the sense to allow ionizing smoke detectors (now in 100% of homes), medical x-rays, and even radium for glow-in-the-dark clocks, until we figured out safer ways to glow in the dark. We could have gone crazy and tried to ban all uses of ionizing radiation, but that would have been obviously bad, so we didn’t. The NRC, smothering as it is for nuclear power, probably averted real harms through its efforts, such as theft of nuclear fuel and pollution from reactor failures. So, I think the right takeaway is not “regulation smothers complex industries”, but “how do we learn and do better with the next thing”.
Disclaimers
I expect this will sound a hand-wavy and under-researched, because it is! I do believe my premises are correct enough to give a vivid sketch, and maybe the ‘regulate physical choke points’ framing will help at this stage of public discourse, as folks stare down the hairy task of slowing AI. I am probably wrong in at least one important way. If you spot a mistake or misunderstanding, please email me (wrong-about-ai at cmart dot blog) and I’ll update the post. If you think “this dummy hasn’t read relevant literature X”, please link me to it, but also understand that this was a first-principles thought experiment from someone not previously working in AI safety or alignment.
If people think this is worth workshopping into a better proposal, well, I’m 9 months into a pre-committed gap year, but I’m willing to cut it short if there is an impactful and skills-fit opportunity to help with AI safety.
For background, I’m not anything like machine learning researcher, but I am soon to enter my third decade of a career that has spanned from enterprise IT, software/systems engineering, open-source maintainership, helping win a few NSF grants, and (since 2024) inference serving of open-weights LLMs on H100, GH200, and MI300X GPUs. As a helpless info-vore, I’ve been paying somewhat-close attention to AI capabilities since approximately 2021. I have also, as is true for most computer-touchers, tinkered a lot with agents in the past several months.
Finally, I used three LLM conversations to help with background research, fact-checking, and critique of this post (NOT to write it). You can read the full transcripts here:
- LLM Lifecycle Resource Estimates Table
- Estimating Scale of Accidental Cyberattacks
- AI Frontier Analysis and Critique (critique of an earlier draft of this post)
-
When I say “GPU” in ths essay, I’m also including the TPUs that Google and Anthropic use, and maybe other hardware like Cerebras wafer-scale accelerators. I know much less about these, but enough to believe that the thesis holds up. ↩︎
-
Each GPU must exchange data with its peers at speeds on the order of hundreds of gigabytes or terabytes per second. This is about 4 orders of magnitude faster than a typical home internet connection, and the hardware that supports it is most typically NVLink / NVSwitch or Infiniband. Some Chinese labs on the receiving end of US export controls have created strong models with somewhat-slower (100-400 GbE) links between GPUs, but it’s possible to limit servers to much slower link speeds still, which would make training so slow as to be impossible for frontier-size models. ↩︎
-
Ugh, LLMs say this cliche now, I should cede it to them along with “load-bearing”. ↩︎
-
A server with 8 H100 or MI350X GPUs costs nearly a half-million dollars. A typical household electrical outlet is not even close to enough to run it, so you need two-ish 240 volt (or 208 volt 3-phase) circuits, of similar capacity to power hot tubs or electric vehicle chargers. The server will either require an external liquid cooling system or have fans that make leaf-blower amounts of noise (100+ decibels). While in use, it will pour several kilowatts of heat into the room, like a stove with all of the burners turned on. And that’s just for one server with 8 GPUs, when a large AI pretraining run uses thousands of GPUs, all working in parallel. ↩︎
-
AI 2040 suggests that a nation-state could effectively hide GPU clusters from the rest of the world by putting them in tunnels near hydroelectric dams. Maybe this complicates international coordination, but within a country, such dams are extremely accounted-for and likely to already be owned by a government agency. ↩︎
-
Certain government agencies seem both obviously incentivized and sufficiently-resourced to push the AI frontier on their own. By empirical public record (or lack thereof), these same agencies seem considerably better than the AI labs at securing their facilities and systems, but I believe they should not be allowed to build recursively-self-improving AI either. Exactly how to prevent this is outside the scope I intend for this post, but I hope a real rule-making effort considers it carefully. ↩︎