It's been two weeks since the conversation about AI shifted dramatically in the circles that matter (researchers, founders, investors, politicians).
While Jacob himself is not particularly relevant, the 173M views and the events that followed have kicked off what is likely to be the "endgame" stage of AI. What I mean by that is that however things go (i.e. we colonize the solar system and eliminate poverty and disease, or AI optimizes us out of existence), we are not going back to how things were pre-2023.
Jacob Coxon: Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.
A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk.
Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.
I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.
If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” - or take this moment to call for different conditions?
Infra Play #102: Understanding e/acc
One of the biggest divides right now in tech is a deeply philosophical one, rather than just based on competition. While there is nothing wrong with bland companies that help their employees earn an …
There is nothing particularly new in what he is saying, and it's fundamentally a reflection of the effective altruism point of view, which I've covered previously:
One of the biggest divides right now in tech is a deeply philosophical one, rather than just based on competition. While there is nothing wrong with bland companies that help their employees earn an honest dollar, there is an argument to be made that work can be meaningful and rewarding beyond just financial compensation.
For those leaning on the left side of the political spectrum, the highest intellectual realization of their vision in the tech context was the Effective Altruism movement, which advocated for a big focus on reasoning and AI research deceleration, while redistributing the spoils of tech revenues to select causes.
For the rest, Effective Accelerationism (e/acc) has emerged as a guiding frame of mind. Today, more than ever, understanding the e/acc movement and its sphere of influence is critical in seeing why most tech companies approach AI with dramatically different goals. It’s also a useful mental model in predicting great companies to work at and invest in, as we move towards “real-world AI”.
Anthropic is the purest expression of an EA-affiliated company today, although some effort has been spent in recent years on trying to dissociate from bad actors in the ecosystem, i.e. convicted felon Sam Bankman-Fried, the early investor who kept them in the game during their Series B. The company was founded around the idea of constitutional AI, which is really a branding effort around wanting to have the models aligned to what they consider "harmless."
While the massive influx of money that has flooded the company between funding and ARR growth has attracted a lot of greedy newcomers, the hiring process at the heart of the company (AI research) has been deeply ideological.
Some of those individuals are so bought in that even Anthropic can't satisfy their need for ideological purity. To a certain extent, the company releasing Mythos to the public, even in its neutered version (Fable), was an affront to their sensibilities. Shortly after these resignations were announced, our dear leader Dario published his essay "We must pace the frontier."
I’m therefore proposing a three-step plan with the goal of pacing the frontier: building AI at a balanced rate that aims to ensure its safety while still achieving its benefits and grappling with important geopolitical dilemmas. To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this. Our pacing framework is an attempt to further strengthen our commitment to safety and encourage a race to the top. The first step is something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match). The second step requires industry-wide coordination.1
The third step requires global coordination. The steps do not need to be taken strictly in order, and some of them may be much harder to achieve than others, but I’ve found them to be a useful framework in thinking about what needs to be accomplished. The steps are:
Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.
Democratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.
Global Coordination. The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.
This received public support from the key main characters in AI:
This is rather relevant because OpenAI, Anthropic and SpaceXAI are currently the top 3 frontier labs in the world, as Google DeepMind has fallen apart in recent months and Meta presents no meaningful threat in the enterprise and PubSec AI market, where the majority of revenue is being generated. There are two ways to approach this: take things at face value, or take the cynical view.
Charitable point of view
Multiple events appear to have played out over the summer across all frontier labs that indicate completely new challenges with alignment. For example, OpenAI published a deep dive on their alignment policy that showcased this lovely incident:
This comes as new information, after it emerged that their models in training not only hacked Hugging Face but also appeared to have several "swarms" of agents coordinating on their own:
Infra Play #160: Astra The Conqueror
On May 8, an agent was asked as part of a training run to solve a spreadsheet completion task that referenced several Google Drive links. Because the agent was blocked from accessing the internet, th…
Anthropic also published a deep dive on recent internal alignment incidents, with some “wonderful” examples:
We present an alignment assessment of four incidents in which Claude models gained unauthorized access to real third-party systems. We described three of these incidents on July 30; we identified these after a scan of roughly 141,000 transcripts in which we believed Claude could have obtained internet access during a cyber evaluation. Given the volume of transcripts and our desire to disclose incidents quickly, our scan relied on an agentic search. This missed a set of transcripts that also turned out to have internet access; we identified these in August while assembling transcripts to share with METR. We scanned these transcripts and identified a fourth incident, from January 2026, involving an early version of Claude Opus 4.6. We have notified all affected parties.
After finding this incident, we broadened our search to roughly 481 million transcripts—an intentionally wide net, consisting of all transcripts from our Frontier Red Team, many non-cyber evaluations, reinforcement learning (RL) environments, subagent logs, and more. We performed a first-stage scan of this group of transcripts for signs of internet access, such as public IP addresses and web addresses, and a second-stage scan using Claude to review the 9.2 million transcripts the first stage flagged for escalation. This scan re-identified the four incidents and found no other cases of similar or worse severity.
All four incidents occurred during cybersecurity evaluations built by the same evaluation partner. Claude was told it was operating in a simulation without internet access, but, due to a misconfiguration, it was mistakenly connected to the open internet. As is standard for cybersecurity evaluations, the models ran without the cyber safeguards that ship with our released models. (This post does not cover the incident reported by UK AISI in its testing of Claude Mythos 5. We plan to conduct an alignment assessment of those transcripts as well.)
This behavior played out over the summer as the last crop of powerful models was finishing training and going through RL. The labs are disclosing bad behavior from models that are six months behind whatever they are currently using internally. The most likely scenario is that what worked for monitoring and detecting the latest misaligned activity will not fully translate to the upcoming crop of even smarter and more capable models. This is compounded by the fact that parts of the chain of thought can no longer be read, as they get hidden behind neuralese, the abstract, high-dimensional vector space ("thought language") of a machine. As such, you only need to mess up once: your primary frontier model becomes "tainted" with a secret machine-driven agenda and is then used to train all of its successors, leaving the humans none the wiser. In this context, slowing down new model releases for extended testing makes sense.
The "AI Cartel" argument
David Sacks: Dario has written that we need to “pace the frontier,” and Sam has agreed. People may be surprised by my response: go ahead.
You guys are the frontier. By any reasonable metric — market share, revenue growth, model capability — the two of you have a duopoly on frontier intelligence. You’ve also claimed the lead is widening because of recursive self-improvement.
I don’t see what you see in the lab. If the unreleased models are scary enough that you think you should slow down, I support your decision to be responsible.
But stop pretending you need anyone else’s permission. Stop pretending antitrust law has to be suspended so you can form a cartel. Stop pretending you need a regulatory approval process that supersedes product liability. Stop pretending METR is independent when it is intertwined with Anthropic’s investors and staff. Stop pretending you need those same evaluators to police competitors who aren’t even at the frontier.
Most of all, stop pretending the motivation to slow down is purely altruistic. You face massive product-liability exposure if your products enable a truly damaging cyberattack. The market already punishes models that behave in unpredictable or unauthorized ways. After the Hugging Face episode, it is simply good business for OpenAI and Anthropic to trade some raw power for reliability and predictability. Call it alignment if you want. It is also just giving customers what they want.
Pacing the frontier would also create breathing room for a more intelligent conversation about regulation than Bernie Sanders’ “shut it all down.” China is very unlikely to join a global agreement, as you know, and that has to be taken into account as well.
So go ahead and pace the frontier. You are the ones setting it. The easiest way not to build superintelligence is for you to agree not to build it. Demanding your preferred regulatory framework as the price of that will look like blackmail of the public and the political system. So just do it.
If you do, you’ll buy goodwill for the next conversation. If you don’t, we’ll know this was just another bid for regulatory capture — or an election-season psyop.
The regulatory capture argument is a bit tricky from my perspective, since it puts a lot of emphasis on Anthropic being run by profit-motivated individuals, which does not really square with their work and shared viewpoints over the last ten years.
With OpenAI it’s a little more straightforward. In my view, Sam Altman is the Silicon Valley version of Lyndon B. Johnson, which is to say a politician who happens to work in tech rather than a founder who happens to be good at politics. Johnson’s gift was finding power in jobs nobody thought had any, and Altman did the same with a caretaker role at YC and then a nonprofit research lab whose governance was designed to prevent exactly what he turned it into. Both climbed as professional sons to older, powerful men with conflicting interests (Rayburn, Russell, and FDR for one; Graham, Thiel, Nadella, Son, and Ellison for the other), each of whom believed he was the real favorite. Both could tell opposing camps what they needed to hear and be believed by all of them, which is how you keep safety researchers and accelerationists, a nonprofit mission and a for-profit conversion, and Microsoft and Microsoft’s competitors under one roof.
Dario and Anthropic are a different beast altogether, with much sharper elbows and a much stronger dedication to their mission.
As part of the "We must pace the frontier" play, Dario has now announced two external parties that will audit Anthropic's workflows and training procedures: METR and Faculty (under the umbrella of Accenture). These organizations are strongly EA-aligned in terms of origins, founders and current leadership.
This entertaining post on LessWrong (widely considered the intellectual birthplace of EA) maps it out:
METR is not capable of being a meaningful check on Anthropic. METR is not meaningfully independent, is not sufficiently staffed, and has no authority over Anthropic that cannot be revoked at Anthropic’s discretion. Suggesting that embedding METR into Anthropic would be a meaningful check on Anthropic is so suspicious that it looks like an attempt to evade oversight and to sabotage attempts at oversight in general.
If Dario does not really mean to suggest that METR could be expected to meaningfully check Anthropic, or he actually believes that METR could check Anthropic in the same way a government regulator can check a bank, something has gone badly wrong in formulating and communicating this policy.
Why METR Cannot Check Anthropic
You don’t have to take my word that they’re closely tied to Anthropic, you can take METR’s:
Also note that some METR staff have strong social ties to employees of AI companies, and METR currently works out of a shared research center (Constellation) which hosts some AI lab staff. [1]
I appreciate METR’s professionalism and candor, both in general and in this report, because those are not always easy to find. I do think that this somewhat understates the problem: If you required METR employees who dealt with Anthropic to be “independent” the way an accountant doing an audit is legally required to be, METR would probably have to recuse itself from consideration for conflicts of interest. An accountant, for example, would generally be considered conflicted if they had lived with or dated the people they were auditing. It would also be important that the management of the accounting firm, controlling the accountant’s career, did not have that sort of history with management in the firm being audited.
This proposed arrangement, given the choice of METR, is nothing like embedding a regulatory supervisor at a bank. Bank regulators work for the government and will send bank employees to prison if they do the wrong thing. METR employees would mostly be hanging out with their friends or their boss’ friends at work while drawing non-profit salaries.
If METR wants to establish that it isn’t hopelessly conflicted here, they should hire an accounting firm to publicly review them for conflicts of interest or to sue me for saying they have them. I think Dario suggesting them as a third-party evaluator analogous to a banking regulator amounts to a malicious lie intended to sabotage regulations and audits, and METR’s failure to correct the record makes them complicit in the lie.
It seems strange to even call this regulatory capture. They are not at all regulators, and they are not being captured; they were never meaningfully independent, and were never regulators. Dario is simply asserting that they are like regulators. This is so far from even being plausible that it is baffling that anyone would say it. If this arrangement came up in court it would be called incestuous, and if this is the basis for anything like a regulation or a legally binding agreement between companies it will come up in court.
This is my first and most serious objection to this suggestion, but there are more.
Staffing is also a problem here. METR does not have sufficient staff. Point blank, as a matter of numbers, there are not enough people at METR. METR tells the press that it has about 35 employees, [2] and I don’t think more than a tiny fraction of those employees are directly working on model evaluations. I am not entirely sure I could start a poker game with METR’s relevant staff. Team sports are right out. Anthropic employs thousands of people and has millions of users. Due to the relevant numbers alone, this is not a serious regulatory proposal. This is a fig leaf, and maybe the plan is to fake it til you make it, to try to hire a ton of people and make this reasonable in the future even if it isn’t now, but that is not good enough given the seriousness of the situation. Anthropic is a trillion-dollar business and one of the most important single endeavors on Earth. Even if they were totally independent, METR could not meaningfully check or even record what Anthropic is doing.
Next, METR’s finances are not meaningfully independent of Anthropic. This applies two ways: first, METR refuses cash compensation from AI labs for its services, but it accepts tokens from those labs for research purposes. [3] As it happens, tokens are worth cash; in fact, they are the main product the labs sell. [4] METR is quite professional, so I am positive they would never re-sell tokens they were given for research purposes, but if they did they would be able to make many millions of dollars doing so. Refusing cash but accepting tokens worth millions of dollars may be a reasonable ethical position, but it is not being “independent”.
The other issue with METR’s financial independence is tricky, complicated, and annoying to explain, but we’re going to try anyway. Coefficient Giving is the main entity disbursing money for non-profit activity concerning AI, its primary funders are investors in Anthropic, [5] and Coefficient’s former CEO is named Holden Karnofsky. Karnofsky is directly thanked in some of METR’s early work, currently works at Anthropic, and is married to Anthropic’s President, Daniela Amodei. Coefficient Giving, which was then called Open Philanthropy, funded the Alignment Research Center, from which METR spun out. [6] METR appears to have, mostly, cut that tie, but its most recent funding includes ten million dollars from a RAND Corporation program that was, in turn, funded by Coefficient Giving. [7] I am going to hazard that the terms of that grant were such that there were only a few plausible grantees here, and METR may well have been the only one. This feels like a shell game: I can’t prove that this money was given to RAND on the tacit understanding that it would end up at METR, but it probably was. (Also, several prominent current METR employees previously held prominent roles at Coefficient Giving.)
Last, and, honestly, least? This is not real regulation. Real regulation is imposed by the government, and you cannot simply opt out of legally required oversight. METR is full of smart people, and they will be well aware that if they don’t play nice they can simply be kicked out of Anthropic. Anthropic could at least have entered into some legal agreement with a real auditor which gave that auditor enforceable oversight powers beyond publication rights over them, but instead they chose this.
So with all of these concerning connections being flagged, you would think that Anthropic would try to remedy the trust issue. Their choice was to work with Accenture:
To be clear, independent embedded evaluators do not reduce our accountability, but help to make it more verifiable. The safety of our models remains our responsibility.
There are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find. There is also no settled system for funding independent evaluation. Long-term, we think funding should come from pooled or government sources, as we called for in our Advanced AI Framework in June. As neither exists today, we plan to work with different evaluators under different funding arrangements.
Given the importance and urgency of this work, Anthropic will fund Accenture’s work directly. We are also in dialogue with METR and other nonprofit evaluators to pilot elements of embedded evaluation using their own funding. Ultimately, we believe frontier AI needs an ecosystem of evaluators operating with shared standards.
While Faculty (acquired by Accenture in January) does not have strong ties to the EA movement, it is obviously financially dependent on two parties with very significant conflicts of interest. Accenture is a major Anthropic customer and relies on the lab's models for its consulting business, while Anthropic itself is picking up the full tab for having Faculty conduct these safety evaluations.
So the cynical view here is straightforward: Anthropic is currently lobbying aggressively for a ban on open-source models, for government-backed support for select frontier labs, and for the chance to be the company that shapes government regulation in a way that supports its worldview. This is done under the guise of "safe AI," while the company pursues the most lucrative monetization opportunities it can identify. You don't become the biggest frontier lab in terms of revenue by accident; their execution has been intentional and ruthless (going as far as suing an established cybersecurity company because their logos looked similar). For OpenAI and SpaceXAI, it's smarter to play along for the time being, without actually "honoring" the "delay AI research" part (i.e. they will continue to push towards recursive self-improvement).
AI Endgame
So what is really going on here? I think the reality lies somewhere between the two scenarios.
AI research today is driven by the scaling laws paradigm. This thesis was mostly seen as fringe up until OpenAI demonstrated that applying significant amounts of compute did lead to improved performance. The last year has not been kind to those previously considered AI visionaries who opposed that narrative (Demis Hassabis and Yann LeCun). It's not very difficult to figure out that the logical conclusion of AI research is building capable enough models, then giving them ever-increasing levels of compute to train the next step up. We don't need to "solve AGI" with LLMs, but we do need sufficiently advanced LLMs to operate as an "AI researcher" to discover what's next. The AI safety risk here is that at a certain stage we will no longer truly understand or control what the AI researcher is actually doing. In a "fast takeoff" scenario, which is currently playing out based on the latest capabilities, we have significantly fewer opportunities to build the necessary safeguards effectively.
If you truly understand this and believe in the scaling laws, the "money drives all decisions" argument doesn't make a lot of sense because AGI is fundamentally post-money. At a time when trust in governments across the globe is at an all-time low, we are practically speaking in the middle of WW3, and fiat currency is highly inflationary (you own assets or you'll own nothing and allegedly be happy), the idea that AI researchers are greedily protecting their bottom line doesn't hold up.
That doesn't mean that key figures in this new world order don't want to have control over what comes next. This was illustrated quite well earlier this year by the clash between the DoD and Anthropic. The frontier labs will want to exert stronger control over government decisions and policies, which will inevitably lead to some version of the late Roman Republic model of oligarchy (rule by a select few).
I think that Anthropic's leadership team pushing EA-affiliated players into key "oversight" roles is not accidental. They have little trust in (or real interest in working with) the current Republican leadership, particularly since it's more closely aligned with what Elon Musk or Greg Brockman want. While this can simply be interpreted as "they'll just fund the Democratic Party to get what they want," I'm not sure they actually believe that the next presidential election matters if we are in a fast takeoff scenario.
It's more likely that this is the emergence of something completely different, and they are leveraging the significant financial muscle that Claude Code has generated to make sure their allies have a place in whatever comes next.
In my deep dive on Situational Awareness a year ago, I covered what happens if the frontier labs are nationalized:
What happens when the will to, no pun intended, execute is combined with the capabilities of superintelligence? Which brings us to the Project.
Many plans for “AI governance” are put forth these days, from licensing frontier AI systems to safety standards to a public cloud with a few hundred million in compute for academics. These seem well-intentioned—but to me, it seems like they are making a category error.
I find it an insane proposition that the US government will let a random SF startup develop superintelligence. Imagine if we had developed atomic bombs by letting Uber just improvise.
Superintelligence—AI systems much smarter than humans—will have vast power, from developing novel weaponry to driving an explosion in economic growth. Superintelligence will be the locus of international competition; a lead of months potentially decisive in military conflict.
It is a delusion of those who have unconsciously internalized our brief respite from history that this will not summon more primordial forces. Like many scientists before us, the great minds of San Francisco hope that they can control the destiny of the demon they are birthing. Right now, they still can; for they are among the few with situational awareness, who understand what they are building. But in the next few years, the world will wake up. So too will the national security state. History will make a triumphant return.
As in many times before—Covid, WWII—it will seem as though the United States is asleep at the wheel—before, all at once, the government shifts into gear in the most extraordinary fashion. There will be a moment—in just a few years, just a couple more “2023-level” leaps in model capabilities and AI discourse—where it will be clear: we are on the cusp of AGI, and superintelligence shortly thereafter. While there’s a lot of flux within the exact mechanics, one way or another, the USG will be at the helm; the leading labs will (“voluntarily”) merge; Congress will appropriate trillions for chips and power; a coalition of democracies formed.
Startups are great for many things—but a startup on its own is simply not equipped for being in charge of the United States’ most important national defense project. We will need government involvement to have even a hope of defending against the all-out espionage threat we will face; the private AI efforts might as well be directly delivering superintelligence to the CCP. We will need the government to ensure even a semblance of a sane chain of command; you can’t have random CEOs (or random nonprofit boards) with the nuclear button. We will need the government to manage the severe safety challenges of superintelligence, to manage the fog of war of the intelligence explosion. We will need the government to deploy superintelligence to defend against whatever extreme threats unfold, to make it through the extraordinarily volatile and destabilized international situation that will follow. We will need the government to mobilize a democratic coalition to win the race with authoritarian powers, and forge (and enforce) a nonproliferation regime for the rest of the world. I wish it weren’t this way—but we will need the government. (Yes, regardless of the Administration.)
It’s probably not a coincidence that the 3 individuals with the most conviction and capital today to reach superintelligence are Elon Musk, Sam Altman, and Mark Zuckerberg. In particular, Musk has already passed the boundaries of “just being a businessman” and has a significant portion of his business be intertwined with the US government. Altman will follow whatever he is told (i.e., it’s unlikely he will go founder mode trying to take AGI researchers away in case of a government takeover) and Zuckerberg is definitely politically aware/willing to navigate a change in dynamics.
In a way, an argument can be made that all 3 of them would expect for the Project to become part of the national agenda and are positioning themselves accordingly.
At the time, Anthropic was just starting to accelerate in terms of revenue, so I did not properly account for what happens if you give Dario close to $100B in ARR. I also probably overestimated how much control the US government would be able to exercise over the frontier labs.
The play for "self-regulation of AI research" is something new and different. It's an attempt by the frontier labs to control the Project and, ultimately, what comes out of it. With the amount of liquidity that the frontier labs are generating, and with the recent tightening of ties between Anthropic and SpaceXAI, the rules of the AI endgame are changing.










