There's not much more to be said on the importance of safety in an AI world that hasn't already been vociferously shouted from rooftops. On September 12th 2026, Dario, Elon, and Sam broadly agreed it needed to be done. However, we are missing crucial detail in the how. How do we really make AI progress safer?
This isn't a question reserved for the so-called 'Doomers'; even those of us who believe in the incredible potential of AI diffusion for individual empowerment and scientific progress would like to ensure that the worst case scenarios are avoided — ranging from simple transactions gone wrong to the sci-fi-like AI attacks on humanity.
A series of strong opinions, loosely held, ranging from strategic to technical.
Safety cannot be an outcome of a single channel; an ecosystem of players will be needed. No single organisation, reporting structure, law, or technical standard will work. It needs to be the constantly-recalibrated combination of these, but in an orchestrated ecosystem of complementary and collaborating institutions — both at national and international levels, with separate arms for technical and regulatory. Organisations like METR, Apollo Research, ARC, or Resolution can continue focusing on technical evaluations. Departments with national legal authority need to both enforce reporting to evaluator organisations, and use evaluations as inputs to decide a combination of appropriate punitive measures based on the findings. Global "shared transaction rail" provider bodies need to also be included (more in #13). Country-specific safety institutes (like the UK AI Security Institute) can track systemic risk in their national context (see #12) and manage shared national digital infrastructure if needed. Demis has proposed an industry wide Frontier AI standards body. However, the remit of this organisation needs to extend beyond pure evaluation, to creation of digital public infrastructure for the AI arena via enabling actions as laid out below.
The dominant discussions around safety focus on mechanisms of external reporting: essentially telemetry on models (as model cards or in more detailed formats) in pre-training, training, and pre-deployment stages.
- a.The scope and nature of per-model reporting & evaluation needs to be more comprehensive. More on this in #10.
- b.Today external reporting via most third party evaluators is voluntary. Therefore evaluators have to calibrate their critique of models to ensure continued participation by the frontier labs, and labs can usually veto publication of results. Evaluation bodies require the legal teeth of enforcement (primarily in the US and reinforced in other countries) to ensure reporting is mandatory and that there are real implications to bad findings.
- c.External reporting/testing are necessary but far from sufficient! External reporting as a compliance tickbox has limited efficacy; outsiders will be, by design, the last to know, and given the least to know.
- d.Finally, individual model cards can't help with systemic risk reduction — more on this in #11.
Triggering within-company governance. One way to instill safety is to slow things down within the frontier labs by creating stronger internal checks and balances rather than relying on external efforts. Those within the know in the company (Chief Risk Officer or Safety Lead equivalents) will have to sign off on safety before it moves ahead. This should happen at personal risk of significant financial liability or incarceration, given the catastrophic potential harms. We need to ensure those closest to the models are incentivised to trigger alarms. This will ideally need to be placed into California legislation; other countries can frame that models applied in their geographies must be signed off by a local national operating as chief risk officer but this will be less effective. Dario Amodei announced 'Embedded Evaluators' on Sep 12. This is a step in the right direction, but not sufficient.
Aligning safety with financial incentives. Within-company AI safety in the frontier labs must be tied to clear financial outcomes. Fortunately for us, the frontier labs will all be publicly listed companies. They will naturally see dips in their stock price based on an incident, but their lack of cooperation with external reporting or safeguard frameworks must hold immediate financial penalties that need to be devised to ensure their financial incentives are aligned. Profit pools have to be linked to preservation of safety as a guiding principle.
Balancing innovation and safety. Typically, there is a concern that additional regulation or safeguards will a) slow down the pace of innovation, or b) limit innovation to deep-pocketed actors that can afford to invest in more costly compliance processes. On a) Innovation is happening at a breakneck pace. Slowing down developments from improvements every month to improvements every 2-3 months will not do significant harm to the pace of innovation, and is a worthy trade off to make considering the potentially catastrophic and globally scalable risks of unsafe AI. Also, since the pace of diffusion of benefits is still quite slow (usually a 6 month adoption cycle for enterprises), reducing speed of model release won't have as significant a reduction in beneficial applications. On b) There is already a much more significant inequity in the capital required to compete on training foundational AI models in the first place — one not based on compliance costs. Compliance costs will be a few added drops in this bucket, and so should not be a deterrent to imposing a few minimalist, well thought through guardrails.
Introduction of a safety leaderboard for AI companies that ensure they don't just win on being the fastest, but also get public recognition for looking out for humanity. A heated oligopolistic competition between Anthropic, OpenAI, Meta, and Google fuels the speed of this arms race, and these companies only score points in this race when they release a new and powerful model. These companies need to win equal public attention and plaudits (scoring points in their competition with each other) when they make real strides towards safety, not just when they release the newest model. A global public safety leaderboard showing which companies have shown the strongest alignment with the values of safety (measuring how the companies comply with reporting; how long they have held back powerful models to improve safety (published as extra-audit-weeks)) ensures there is public recognition or knowledge of company decisions towards safety that otherwise remain hidden and unacknowledged. This information can in turn reflect in stock valuations.
Ensure and increase liability. We already have a framework in place to hold companies accountable when their creations cause harm: product liability! For agreements signed between frontier labs, hyperscalers, and/or enterprises, terms and conditions forms must have significant financial repercussions as part of product liability clauses. The presence of these clauses should be enforced in national legislation. Frontier labs cannot be allowed contract clauses that shrink their liability — given they have market power, their lawyers are likely to otherwise bulldoze client lawyers. This is an area national AI law could enforce.
Global legal frameworks. Frontier models are covered by legal jurisdictions in California and the US, yet operate prolifically in emerging markets like India or Brazil or Indonesia. To tackle this problem, we may need to deploy a home-host framework with mirror filing to ensure alignment of regulations. Country-specific Safety Institutes ("host") need to lobby California and US federal regulation ("home") to ensure they match up with the degree of regulation (including mandatory evaluations) in country-specific AI laws. Ideally, home and host countries should align on evaluation providers. Frontier lab market access in the host country (for both users of the models and for data centers or other infrastructure) can be conditioned on mirror reporting.
Learn from other sectors. While the scale of threat posed by AI is new, humanity has encountered parts of the problem in other spaces: nuclear non-proliferation (a humanity-wide immense loss of life nudged unprecedented national treaty-level organisation & IAEA); financial regulation architectures (we learned through many failures to manage fast moving, catastrophically-coupled private systems); and civil aviation reporting (the immediate threat to human life forced reporting and management of technical failures via ICAO).
Telemetry & model attestation. Per-model filing & assessments need to be improved, with standards set globally. Top model scientists need to come in to evaluate the various stages and characteristics at which models should be evaluated. Broadly, they should cover chain of thought monitorability; capability tripwires including autonomous cyber offense, situational awareness, and 'sandbagging' or deliberate underperformance on evaluations; propensity signals during training including reward hacking, unauthorised channel-seeking, and evidence tampering attempts; network isolation (do isolated agents still share any writable surface, etc). Pre-release should also include commitments including standard response mapping at the frontier lab to incidents, including allowlisted egress (blocking all outbound network traffic except for a pre-approved list). Surprise evals should also be introduced, rather than lab-volunteered evals. Ideally, we have one global model evaluation body that sets tech standards for evaluation in a consultative process, and is held accountable. Too many evaluators would create competition that may favour the labs' ability to partner with a 'softer' evaluator, or create a fragmentation of X country requires Y evaluator that would be too onerous for frontier labs (and therefore become an excuse to delay safety). This global evaluation body should be able to 'acquire' smaller, innovative evaluators on specific benchmarks to crowdsource quality.
Solve for systemic risk beyond just individual risk. We need to regulate the risky activity wherever it occurs, not just at frontier labs; and watch the links between different systems to prevent compounding. OpenAI/Hugging Face already indicates the ability to cut across multiple platforms to multiply risk. To do this, we need #12 and #13.
Data access. Country-specific AI Safety Institutes need to maintain a registry (that can be seamlessly reported into by agents via open APIs) of all AI industry players or operators in the country, as well as a reporting account of major deals and partnerships (invoice filings submitted to entities like India's GSTN auto reported to the country safety institute). This is required to monitor threat levels and predict potential snowball effects of smaller incidents, and track networks that agents may have access to.
Shared rails for all agents. Existing data won't be sufficient to track incidents in real time and prevent them from scaling. We will need some sort of shared Digital Public Infrastructure rails for all AI agents — where they carry credentials as per standard specs indicating their model, maker, task, and consent, to ensure that incidents can be managed as they unfold. Threats to safety need to be managed based on amendments to the normal transaction rails of AI agents; otherwise, if it stays in an outside reporting rails or separate infrastructure its efficacy will always be delayed.
Use AI to keep AI safe. Safety efforts cannot bring sticks to the gunfight. Nor can it win a race where the opponents are moving at AI-bullet train speeds with only-human drawn carriages. A commitment to safety should not mean a rejection of the technology that can bring much productivity to the safety fight. However, critical functions must always have human oversight and override.
Reserve the heaviest safety-related compliance to the frontier, whilst still encouraging diffusion of plainly useful past generation models. The safety effort should be targeted at the edges of AI capability, as well as monitor areas with disproportionate compounding risks — rather than impose one-size-fits-all rules for all players and activities.
Ensure "Living Wills & Fire Drills" at frontier labs. After 2008, banking introduced protocols for orderly wind downs in case of the worst, and for operational drills to practice some of these procedures. The same set of well understood practices for a containment plan for a rogue model should be required for frontier labs, which includes quarantine procedures, weight rollbacks, revocation of credentials, systemic dependency mapping of the model, etc.
The bottom line: nobody can presume to have all the answers we need, but the bones of a long lasting AI safety solution would need to contain all of the above.