Back home, people like launching fireworks on New Year’s Eve. Ambitious pyrotechnicians sometimes forgo proper batteries and fire rockets from empty beer bottles half-buried in the December snow. And the most daring and inebriated instead discharge the fireworks from their bare hands. One fateful night in 2019, my dear brother launched a rocket into a balcony, incinerating ski supplies and inflicting thousands of euros of property damage.
Right now, we’re in the launching-fireworks-from-hands era of AI policy.
The labs are moving faster than government can follow. And while Washington is waking up to that, it could easily make things worse. The material problem isn’t even any one capability, any one swarm or hack or risk vector, but a pace of AI progress and of AI news cycles that government isn’t equipped to handle. I’ve always favoured a ‘muddling-through’ approach to AI policy. We keep reacting to iterative deployment of more and more powerful systems, don’t overreach, don’t underreact, figure things out as we go. I’d still like to keep doing that.
But muddling through requires time between surprises. Government is running out of it. In the last few weeks, AI policy has repeatedly been shaken up by dire expert warnings, autonomous swarms of hacking agents and a generational math breakthrough engineered by an unreleased model far beyond the already-impressive current frontier. All in one Congressional recess.
This gap is untenable, and the U.S. government is quickly realising it needs to keep pace. The best way to do that is to empower third-party oversight over frontier AI developers.
Closing the Loop
AI lab employees believe ‘recursive self-improvement’ is imminent—or perhaps already underway. The models they’ve built are exceptionally good at building even more capable AI models, which in turn will be even better at autonomously building even better models. A week after GPT-6 Astra’s capabilities surprised a world of new users, OpenAI revealed it already had an even-better model deployed internally to solve mathematics’ Millennium problems.
We can and should have discussions about the exact meaning of RSI: perhaps it describes the beginnings of a software-only intelligence explosion; perhaps it’s a more prosaic flywheel effect. Maybe it’s a fleeting period of acceleration . These are enormously different, but have the same immediate policy implication: we’re headed for a widening gap, where the pace of industry accelerates past the pace of government, leaving the latter effectively uninformed and disempowered.
Even if things were going well in technical AI safety, this gap would be deeply objectionable. In effect, this pace erodes the state’s sovereignty: a handful of companies are the only ones in control, everyone else lives with the consequences. I’d still want the government to have some insight into what’s happening, some ability to shut it down and some ability to steer it, even if I expected things to go very well by default.
However, things aren’t going well in AI safety. Those closest to the models tell us they aren’t nearly aligned enough. Many models are too unreliable to warrant wide deployment even today. Deploying them as a launchpad, having them align and develop the next generation of agents is even more risky. Read the METR report, look at the degree of strategy and long-term planning already employed by the swarm responsible for the Hugging Face attack; would you really trust systems like this to build better versions of themselves?
We’re approaching the loss of meaningful democratic oversight over AI takeoff. If we continue down this route, we’ll lose the ability to intervene selectively; we’ll only be able to let it rip or shut it down. Talk about an unsteady hand: in reaction to real risks and policy gaps, everyone is quickly firing off their ill-conceived policy solutions. That could do much more harm than good.
You’ll see many new draft bills floated in the next few weeks, I’m sure: AI risk is going mainstream, and policymakers will want to pile on. That is not an indication that Congress will actually move—especially not before the midterms—or that the White House will tolerate substantive legislation once the current wave of attention has quieted down. And in the meantime, much can happen: an election in the fall, a new Congress in the new year, subpoenas, investigations, hearings, a government divided and unlikely to pass a law. It will be easy to mistake activity for progress in that environment.
Executive intervention is just as volatile. Too aggressive an intervention will struggle to survive opposition from the still-influential accelerationist faction. Too ham-fisted an oversight regime could destabilise the industry, crash the stock market and ultimately imperil America’s geopolitical position. Too big of a role for national security agencies would concentrate power far away from democratic oversight, replacing reckless labs with an ill-equipped military-industrial complex running the show.
There’s a narrow path between these that allows reestablishing democratic oversight with some degree of moderation. Sooner or later, the administration, pressured by Congress, the public, and reality, will come looking for it.
Here’s where it should start.
In the Loop
The U.S. government should:
empower third-party evaluators,
secure their access to frontier labs to
conduct reactive incident assessments and
proactively oversee AI labs’ safety practices.
First, government needs an ability to comprehensively investigate incidents. In the last few weeks, we’ve repeatedly seen AI developers struggle through haphazard and premature deployments; when trying to figure out what exactly is going on, the general public rarely receives any comprehensive readout. When lawmakers write letters, they get company responses that are shortly thereafter disavowed by that same company’s employees. When developers invite evaluators to find out what happened, the evaluators get all of six days to investigate and have to go on tour explaining all the things they weren’t allowed to do. And even that minimal level of external access is neither mandated nor even incentivised by government; it’s a purely voluntary action undertaken by the developers themselves. That is no system of governance.
Yet after the next incident, someone in the White House will want to send someone into the lab.
For better or for worse, the principals that think themselves in charge of AI policy—cyber director Sean Cairncross, Treasury Secretary Scott Bessent, and the notoriously quiet OSTP director Michael Kratsios—will want to know what’s going on. If they feel like something’s happening that they don’t like, they could overreact in many counterproductive ways: slap export controls on models again, send in the NSA, compel information through the Defense Production Act, and so on. This all doesn’t sound very helpful to get at the deeper problems, so we need to provide this impulse with a productive release valve.
They should be given some way to conduct an actual incident investigation: we need to hand the U.S. government a way to send someone into a lab who can come back with a report on what happened and what needs to be done differently.
Second, government needs permanent oversight over AI labs as entities. That means that at any given time, a few people from a third-party organisation hang out in the OpenAI Slack, sit in the cafeteria, sniff around logs and request conversations with executives. An entity-based approach is prudent because it’s very difficult to break down questions of frontier AI safety to questions about distinct models or products. Many of the most dire risks emerge from complex institutional dynamics: decisions that are made about how to set up a training run or an RL environment. The critical moment for government intervention might arise long before deployment. If government wants to be able to act in time, it needs embedded, continuous oversight with a direct line to the agencies.
The object of oversight, the thing we should monitor, is the entity that makes these choices. You can only do so much rocket evaluation. At some point you have to figure out whether the guy is drunk enough to actually try launching fireworks from his hand.
Ipsi Custodes
Third-party evaluators are best suited to these tasks. First, because no one else is up to it. The U.S. government currently lacks the talent and expertise to do the job. The NSA has some capacity to run evaluations of models; CAISI is full of talented and committed people, but it is embattled by the complex politics of the Department of Commerce in which it sits and seems unlikely to enjoy the full trust and faith of the administration. Third-party organisations, in the meantime, have proven time and time again that they are at the cutting edge of their field. They can hire and scale quickly, and have met increasingly challenging evaluation tasks to near-universal acclaim. Almost everyone likes them; at least better than the government.
Second, because third party evaluators are not tied to any party in the inter-agency knife fight over who gets the AI portfolio. Many principals have some interest in running AI policy. Some of them might want the political limelight; others simply think they’re the most qualified; others again fear the political volatility of letting anyone else handle it. I don’t know where effective AI authority will ultimately sit between different parts of the White House, Treasury, Pentagon, and Commerce—and neither do you. So if we build capacity in any one of these places, but a different principal wins the fight, the capacity is sidelined too. We might be able to transfer it, but a department that lost its authority might be reluctant to surrender its expertise as well.
Third parties can build relationships with all principals, and be deployed by whoever comes out on top. That’s how we make sure the expertise we build survives the turf war.
Third, evaluators already have an established working relationship with the AI developers. Even if you could send in a team of NSA AI safety experts tomorrow, I’m not so sure they’d be able to effectively carry out the task at hand. They’re not familiar with the culture of frontier labs, the language, what’s meant jokingly and what’s serious. They’ve never been to Constellation. If they sat down at a table with the safety team, the safety team would get up, grab their pizza, and leave. If they were embedded into a Slack channel, the channel would go quiet. Third-party organisations have a real rapport with the relevant teams in the leading labs, and they’d have a huge head start in orienting themselves around the organisation they’re supposed to monitor.
And fourth, I suggest third-party oversight because what I describe is too intrusive to be done by government directly. If embedding is necessary—and I believe it is—we shouldn’t rush to have the government do it. The CCP does this across most of the industry, and America has rightly identified that there’s a problem involving political officials in the day-to-day of its private sector. My desire to have third parties do this instead is fundamentally a libertarian impulse: it creates a layer of separation between the direct governmental chain of command and individuals that gain deep access to some of the most important companies in the world. It’s the least intrusive way I can think of to gain the access we need to keep our hand on the steering wheel.
Quis Custodiet?
When I discuss this approach to policymakers, especially those with an accelerationist bent, one frequent question is ‘but wait, what do these embedded evaluators actually do?’. The narrow remit above—a red button, incident investigation, and safety assessment—is a good start. It’s not in itself a fully fledged system of AI governance, but we were always going to have to develop that later; through industry standards, in the states, in Congress, through executive orders, but separately from this immediate stopgap intervention.
In the meantime, third-party access is more constrained than agency embedding: third parties don’t get much practical authority to do anything except what the government asks of them. If the government wants information, it can talk to the third-party orgs; if the government wants to figure something out, it can consult them; if the government would like something done differently, it can ask the labs to do it and the third parties to evaluate.
The embedding is really only the enabling mechanism for that action; for any further action, the government still needs to make a discrete decision. That leaves a large part of AI governance up to the executive. If you distrust the administration, you’ll be less at ease than with a substantive framework; but I don’t think a substantive framework is in sight, so this strikes me as the best way to restore the balance of power between government—any government, even this government—and rampant AI progress.
In the meantime, there’s a legitimate concern that these organisations will get captured. The same familiarity with the labs that makes for useful rapport also makes them less likely to be exceedingly adversarial toward their hosts. But this is a much bigger problem now, while the labs still choose whether they want to be evaluated. Once government mandates this evaluation, that changes. Third parties can grow more adversarial as they find their externally-mandated footing; they can afford to erode their existing familiarity as their official authority increases.
Still, beyond all incentives, the social capture risk remains. There are some parallels from other industries: we can rotate evaluators, we can grow the ecosystem, and so on, but they’ll still always have their blind spots. That means they’ll require oversight by the government that empowers them, which in turn means we will need some executive capacity after all. This is an open problem, especially in the long run. But overseeing evaluators is still a bit easier than overseeing frontier model training runs.
Government by Phone Call
It’s 2026, and what used to be a law is now an executive order, and what used to be an executive order is now a phone call from Susie Wiles. I’d prefer to codify all this through legislation; but I’ll offer a minimal way to get it done first.
The current paradigm of AI policy revolves around the executive indicating what it wants, and the labs being too risk-averse to figure out what happens if they don’t comply. My understanding is that this is not precisely constitutional, but it nevertheless seems like a surprisingly sturdy house of cards.
We can laugh about this being absurd, we can question if it’s a good idea, but we should move on to practical considerations. Realistically, I don’t think OpenAI and Anthropic are going to risk their IPOs by suing, and Meta and xAI aren’t going to risk their relationships to the administration, either. To be entirely clear: this is an objectionable state of affairs, and it trades one shape of objectionable power concentration for another. But I think it’s a decent trade regardless, and the only one that might be on offer as policymakers look for a response.
The institutional pathway to making the wishlist come true is simple: the administration declares, to labs and perhaps the public, that it would like to see the above happen. That alone would already make a big difference. On the side of the labs, it would incentivise them to start granting access for fear of being compelled to do so. On the side of the evaluators, it would send to them a direly needed signal of being part of the plan—allowing them to scale much more ambitiously. Official recognition would also help evaluators recruit: it would show more elite researchers that work outside a lab can have as much or more influence than being on the safety team inside.
The mechanism of the resulting governance structure is the White House phone call: if something goes wrong, Sean Cairncross calls Sam Altman and tells him he should let ‘the guys who did the Hugging Face thing’ look into it. If something looks off, Susie Wiles can call Tom Brown and tell him to procure a readout of current safety standards from whoever their embedded auditor is. If the White House wants to check labs’ compliance with an updated version of their secret framework, they can task embedded evaluators to verify ongoing compliance. And if it looks like the developers are trying to prematurely close the loop on recursive self-improvement or seem at risk of developing another swarm of hackers (what the FRONTIER Act described as imminent risk of catastrophic harm), the embedded auditors have the number of someone in the government who can intervene.
I’d much rather do all this in more orderly ways. But we’re going straight from recess to midterms to a likely split government, so I’m not sure there’s actually a legislative window any time soon. If I’m wrong, if Congress agrees to act swiftly, I suspect we could take much of the FRONTIER Act’s language as a basis, sharpen it, tack it onto a more politically promising vehicle like the unpublished Thune-Klobuchar draft bill, and pass it.
I’d be in favour of trying that; I just don’t think it’ll work, at least not in time. Taking the pace of progress seriously means that we have to come up with at least a minimum viable executive version under the administration we have.
Making it Last
Assuming we don’t pass a law, there are still executive pathways to solidify this approach. We might start by anointing qualified third-party organisations through formalising their relationship with the U.S. government: we can start through memoranda of understanding with the relevant departments and intensifying cooperation on security practices. We can continue with contracting them more officially, or even drawing them closer into the government orbit through Other Transaction Authority or moulding some organisations into FFRDCs, which would allow them more leeway in national security matters.
We can also run this through a self-regulatory organisation. Depending on when you’ve last talked to the Treasury Department, you might call this organisation SRO, FINRA for AI, FARO, or SAFA, but they all mean the same thing: industry agrees on minimal standards and finds a way to demonstrate compliance to the government through intermediaries. These days, many observers expect the default version of this idea to either not happen at all or to turn out fairly ineffective, with limited enforcement capacity and a distorting presence of safety-skeptical participants to water down the frontier labs’ concerns.
But done right, an SRO could be a good home for the third-party setup I describe. Guided by government, industry could agree on standards for both incident investigation and continuous oversight, and designate suitable third parties to audit against these standards. Truly codifying this would require passing a law, but if all goes well until then, I could see a smooth on-ramp: we start with governance by phone call, continue through semi-codification in a makeshift SRO, and eventually truly codify the approach by law.
From there, a lot of endgames are conceivable. Some of them turn the ad-hoc third party setup into a true private governance model with all the perks and benefits of regulatory markets contained in the IVO approach deployed in a few states—including California, as of yesterday.1 Others eventually absorb the third parties into a governmental structure, nationalising parts of the ecosystem and putting others under more and more exclusive and restrictive contracts.
But none of the short-term actions I suggest put us onto a clear path to one endgame or the other. You don’t need to be an IVO enthusiast to appreciate the practical upsides of using third-party organisations to solve the immediate problem at hand; and you shouldn’t get your hopes up for the private governance endgame just because third parties run the show for a few weeks. We’ll have debates about the endgames, I’m sure. But not yet.
Building the Bench
Getting this done requires getting the evaluators up to speed and popularising the idea itself across the policy environment.
The ecosystem’s main problem is that it’s not quite an ecosystem yet. METR is an excellent organisation and the closest to a full-stack embedded evaluator and investigator we have. Apollo Research does not conduct quite the same breadth of activity, but it’s getting there. A few others with narrower remit exist, some of whom organised in AI Evaluators Forum, though they’re far from being able to do the whole job. But this doesn’t work if it’s just one organisation; it’s unreasonable to ask the government to put all the eggs into one basket, just as much as it’s unreasonable to ask the AI safety ecosystem to trust one organisation to never stumble or misstep.
The nascent ecosystem as a whole needs to be scaled up quickly, both with technical talent and—perhaps more importantly—political talent that can make the ecosystem’s case to the principals that would empower them. But we need to move, now. If this idea takes root, the Commerce Secretary calls, says he’s ready to go and wants to meet the ecosystem, and even one of the organisations is just a few Bay Area researchers in a trenchcoat, this won’t work.
The second problem is political legibility. For very good reason, third-party evaluators are not policy advocacy organisations. They’re not supposed to be polished guys in suits that speak the language of Capitol Hill or even the MAGA movement. But their approach still needs advocacy. I genuinely believe there’s a win-win to be had here, that the third-party approach best suits the administration’s stated AI policy priorities. But in the past, it has been difficult to make that case in Bay Area parlance alone.
People with access to the administration can help make the case for them. They should spend some of their limited time, attention and energy to lobby for this solution in particular; and the rest of the ecosystem should support them however they can.
These two problems need to be resolved in time, before the administration is compelled to move. That might happen soon: AI progress is accelerating, and the government is out of the loop. Senate investigations, heaps of public and political pressure will compel some action. As much will be clear to everyone—including cabinet officials looking to run for President without this blemish on their record.
They’ll soon find that, if this trend is allowed to continue, something important will be lost: perhaps control over the most powerful corporations in the world and the technology they build; or the prospect of a free American AI ecosystem as the government is ultimately forced to crack down hard instead.
The policy window is open, and more incidents will come. Government will need to react. Its best shot is empowering third-party evaluators now.





This was very good. I would add, in addition to "The narrow remit above—a red button, incident investigation, and safety assessment—is a good start"
...that, to your point about the current climate for better or worse being amenable to certain kinds of political/policy entrepreneurship, one straightforward intervention could be just "call Susie Wiles".
Like, I think it's really easy to overcomplicate the mechanism of intervention - trying to get too caught up in the relationship between process and actuated change, with a level of precision that is illusory to begin with - when in reality much of what is going to actually motivate the labs is to "avoid a variance-laden pain in the ass" of whatever variety. Inasmuch as the variance is located in 'some civil litigatory/arbitration process' versus 'the side of the bed Susie/DJT woke up on that day', that matters when one is actually in the process, but the whole point is to incentivize avoiding that kind of variance/friction.