Microsoft CEO Satya Nadella published a lengthy essay on X on Saturday arguing that advanced AI models must be engineered with containment controls from the outset, including what he called an emergency brake: an authorized person should always be able to pause or shut down a model mid-task. He wrote that the industry must “assume a model is compromised and contain it from the start,” and called for every meaningful model action to be documented with tamper-proof human-readable evidence.
Nadella framed the argument as a trust architecture problem. Models can no longer be treated as nested black boxes whose outputs teams simply accept or reject, he argued. Instead, the model should be separated from the “harness” that orchestrates its work, with controls and safeguards externalized so they survive even if the model itself behaves unexpectedly. He also called for timely incident disclosure, independent audits, and verifiable data on model inputs and outputs, plus industry-wide containment standards where current guidelines fall short.
The intervention lands as the industry’s safety debate intensifies. It follows Anthropic’s Friday disclosure of unintended agent actions in its own evaluations, which drew a White House briefing and an FTC response, and Anthropic CEO Dario Amodei’s plan for more cautious frontier development. Nadella’s version goes further on the operational point: he treats powerful models as potential insider risks and asks every deployer, not just the labs, to build kill-switch capability rather than rely on developer assurances.