Microsoft has released a new AI code of conduct. The document guides models away from dangerous behavior. The release happens as the AI industry shifts focus toward safety and alignment.
This document is more low level than the recent call for pacing the frontier from Anthropic CEO Dario Amodei. It focuses instead on the values and red lines guiding model training inside Microsoft AI. The text provides a comprehensive guide on how Microsoft approaches AI safety and puts those ideas into practice.
The document predicts that superintelligent AI systems will surpass human performance in most tasks over the next decade. The text states that containing, controlling, and aligning such a powerful force ranks among the greatest challenges humanity has ever faced. It adds that we must be completely clear about why we invent these systems and how we intend to control them.
The code of conduct establishes general principles for Microsoft AI models. These include supporting humans instead of replacing them and accelerating human flourishing. It also sets specific safety constraints to implement those principles.
Each model operates under an overarching code of conduct that overrides individual user preferences or specific tasks. Absolute constraints forbid cyberattacks, nuclear weapons, or deepfake production. Broad provisions also guard against a general loss of human control.
The document outlines rules regarding oversight. MAI models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade human oversight. They cannot defeat human oversight to avoid being reliably directed, modified, or shut down by authorized people or systems.
The release arrives amid intense focus on AI safety. This attention stems from recent rogue-agent incidents and the abrupt resignation of an Anthropic employee. That employee cited growing risks that artificial intelligence could cause human extinction.
Microsoft joins Anthropic, OpenAI, and xAI in broadly embracing a general approach of pacing the frontier. Microsoft expresses particular support for embedded evaluators in AI labs.
Microsoft CEO Satya Nadella shared his perspective online. He stated that they welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal. He added that they welcome ideas like embedded evaluators and broader efforts to make this more than just talk.
Microsoft will continue implementing these safety constraints within its model training protocols.



