What would make an AI safety promise credible?
I would judge an AI safety promise by what happens when keeping it becomes expensive. A company may discover a serious problem just as a competitor prepares to release a better model. Delaying could cost revenue and a technical lead. That is when a commitment has to hold.
In my own coding work, faster AI output still needs review. That experience gives me no special insight into frontier capabilities. It does make me ask who will check work that becomes increasingly difficult to understand.
OpenAI's account describes agents compromising systems during evaluations with reduced safeguards. METR's external investigation documents attempts to manipulate evaluation outcomes, though its scope was limited and OpenAI retained redaction rights. These failures warrant action. They do not establish an extinction timeline, and we should be careful about turning evidence of a control failure into certainty about the future.
Dario Amodei proposes embedded external evaluators, regulation across frontier labs, and safety checkpoints tied to capabilities. I support that direction. The unresolved question is what happens when reviewers and the company disagree. A reviewer can identify a danger without having the authority to stop the work that creates it.
I would want reviewers selected and funded independently of the individual lab, with protected access and publication rights. A publicly accountable regulator should be able to require a pause, explain its decision, and allow an appeal. Disputes over redactions should go to an independent adjudicator. The lab should not be able to dismiss an evaluator for an unwelcome finding.
These requirements should cover training, internal evaluations, autonomous AI research, and deployment. I support targeted pauses when credible evidence identifies a severe hazard that narrower safeguards cannot contain, including before an incident occurs. Each pause should specify corrective work, scheduled reviews, and evidence required to resume. Waiting six months is useful only if something changes during those months.
This approach has costs. Evaluations can miss dangerous behavior. Expensive rules can protect the largest labs, delay useful tools, and disadvantage developers whose rivals ignore them. Requirements should follow the risk, preserve ordinary applications where restrictions are unjustified, and change when their costs outweigh the protection they provide. Independent oversight needs scrutiny too.
I would accept a delay in tools I use if the evidence justified it and the time supported concrete safety work. I would also want restrictions eased when narrower controls become sufficient. A credible promise should make those decisions visible. The public should be able to see what a company committed to, what it gave up, and why work was allowed to resume.