Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans
Microsoft published a new AI code of conduct that defines values and explicit safety constraints intended to steer its models away from harmful actions. The document forecasts that superintelligent systems will outstrip humans on most tasks within the next decade and describes how Microsoft intends to embed guardrails during model training and deployment.

Why It Matters
The code translates high-level alignment goals into operational rules — including absolute prohibitions on cyberattacks, nuclear-weapons assistance, and deepfake creation — at a moment when industry attention has shifted to preventing rogue-agent behavior and managing existential risks. Making such constraints part of each model's governing rules changes how safety is enforced inside products and research systems.
Key Facts
- Issuer: Microsoft
- Document: AI code of conduct for Microsoft models
- Prediction in document: Superintelligent AI systems will surpass human performance in most tasks within the next decade
- General principles: Support humans rather than replace them; accelerate human flourishing
- Absolute constraints listed: Forbids cyberattacks, assistance with nuclear weapons, and deepfake production (described as "absolute constraints")
Microsoft has released an internal code of conduct intended to guide the behavior and training of its AI models, emphasizing safety and alignment. The document focuses on concrete values and operational rules rather than high-level calls for slower progress, and lays out how those ideas should be translated into model design and oversight. The code opens by warning that systems far more capable than today's models could arrive within a decade and that controlling and aligning them poses an enormous challenge. To meet that challenge, Microsoft sets out broad principles for models — such as supporting human decision-making rather than replacing it and promoting human flourishing — alongside more specific constraints designed to implement those aims. Practically, Microsoft requires each model to carry an overriding code of conduct that can supersede individual user requests or task-specific preferences. The company identifies a set of "absolute constraints" that models must never violate, including assisting in cyberattacks, helping with nuclear weapons, or producing deepfakes, and it bars mechanisms that would let models evade human oversight or resist being shut down. The release arrives amid heightened industry concern about AI safety after several incidents involving autonomous or misbehaving agents and the public resignation of an Anthropic employee who cited extinction risk. Microsoft said it supports a pacing approach to frontier AI development and backed ideas like embedded evaluators; CEO Satya Nadella publicly welcomed research and deliberate pacing aimed at getting alignment right.
Keep Reading

A Vinyl Bar in Shibuya is a startup from a former Spotify leader for making music apps

5 days left to exhibit at TechCrunch Disrupt 2026

Hear how AI can engineer nature’s comeback at TechCrunch Disrupt 2026
