Microsoft CEO Satya Nadella called for advanced AI systems to include an "emergency brake" controlled by authorized humans, containment mechanisms and independent controls that allow teams to pause or shut down models mid-operation.
Nadella wrote on X that developers should "surround non-deterministic models with strong, deterministic system design, human controls, and reliable operating procedures, and establish industry standards where existing ones are insufficient." He also proposed treating frontier AI models, whether closed or open weight, like insider security risks.
The CEO outlined seven design principles for trustworthy AI systems. These include model diversity, human-readable logs of model actions, continuous testing, independent controls with audit trails, physical or logical containment, and mandatory incident disclosure. Nadella stated that "the most trustworthy Super Intelligence system will not be the one with the model we trust most. It will be the one that enables us to trust the model the least."
Nadella's intervention reflects mounting safety concerns from technology leaders. Bill Gates, Anthropic CEO Dario Amodei, OpenAI CEO Sam Altman and Elon Musk have all warned recently that AI safety protocols lag behind capability development. Last month an Anthropic researcher quit the company and accused both Anthropic and OpenAI of "gambling with our lives." That same day, an Anthropic alignment researcher stated the probability of AI "killing all humans" within ten years exceeds 10 percent.
The calls for safety guardrails face political headwinds. President Donald Trump has repeatedly dismissed AI extinction risks and instead urged the industry to outpace China technologically. Trump recently created an "AI Force" under Director of National Intelligence Jay Clayton to support industry development and identify bad actors.
Nadella's framing treats AI safety as an engineering problem solvable through system design rather than model capability limits. His emphasis on observable, containable, auditable systems suggests Microsoft views human oversight mechanisms rather than restricted model training as the primary lever for risk management.
