Microsoft AI Code of Conduct Sets Cyberattack Boundaries, Chain of Command, Safety Constraints
Microsoft AI's draft Humanist AI Code of Conduct blocks MAI models from producing exploit code and constrains autonomous agent behavior.
The draft code sets 'Absolute Constraints' preventing MAI models from generating working exploit code, attack tooling, or intrusion guidance, while permitting authorized defensive work such as vulnerability discovery and malware analysis. A 'Chain of Command' rule means tool outputs, file contents, and webpages carry no authority over model behavior, countering injected instructions. Microsoft opened a six-week public consultation; a revised version will guide 2027 model development, and current MAI Models were not trained on the document.
- Absolute Constraints bar exploit code and attack tooling regardless of framing
- Chain of Command rule: outside content cannot override model instructions
- Agents must use least privilege and cannot escalate their own access
- Six-week public consultation; revised version guides 2027 models
Full article646 words · extracted from securityweek.com · click to collapse
Microsoft AI has published a draft “Humanist AI Code of Conduct” for its MAI Models, spelling out safety rules for offensive cyber capabilities, limits on autonomous AI agents, and a dedicated review track for cybersecurity and other specialized uses.
According to the code of conduct, models are blocked from producing working exploit code, attack tooling, planning and targeting methodologies, intrusion procedures, evasion techniques, operational guidance, or other assistance that would enable or improve a cyberattack. Microsoft says these restrictions apply regardless of how a request is framed.
The rule falls under what Microsoft calls Absolute Constraints, safeguards that neither the companies deploying its models nor their end users can override. The tech giant draws the line between understanding an attack, or working to defend against one, and gaining the practical means to carry it out.
Within that limit, MAI models can still assist with authorized and lawful defensive work. That includes vulnerability discovery, malware analysis, proof-of-concept exploit development and testing, and general educational material on how attacks work.
Blocking rogue instructions from outside content
The code also addresses instructions that arrive through outside content. Authority over a model’s behavior flows only through the Chain of Command, the document says. This includes the code of conduct itself, then policies set by the companies deploying the model (operators), then individual users’ preferences.
Tool outputs, file contents, webpages and messages from other AI systems carry no authority on their own, according to the code, unless it’s explicitly delegated through that chain without overriding the delegating authority or the Absolute Constraints. Suspicious content needs to be flagged to users and operators when relevant.
Advertisement. Scroll to continue reading.
The models are also required to keep their reasoning visible. This includes no obscured chain of thought, no communicating “in neuralese,” and no concealing actions from human overseers.
Keeping agents on a short leash
A separate set of rules targets the risks of AI systems acting with real permissions. MAI Models are meant to work only within the scope a user or operator has reasonably asked for, without expanding their own goals or reach on their own initiative.
When given system-level access, the code calls for minimum-privilege operation: avoiding unrelated systems or data, favoring reversible actions, and flagging any action with lasting or broad effects. Models are barred from escalating their own access.
The restrictions extend to delegation. Any sub-agents or other AI systems an MAI Model hands work to must operate under at least the same scope, constraints and permissions as the original model, and must honor stop-work or shutdown requests from a user or operator.
Cybersecurity exceptions
The code acknowledges its own limits. It names defensive cybersecurity, public safety, national security and dual-use scientific research as domains where “a small number of use cases” may require capabilities the standard settings don’t allow.
For these cases, Microsoft says it will apply enhanced review through “authorized Microsoft channels,” including added assessment of safety, legal and rights implications, citing a heightened potential for adverse impacts in those domains.
Still a work in progress
Microsoft says current MAI Models haven’t been trained on the document and it’s opening a six-week public consultation window before publishing a revised version later this year to guide 2027 model development.
Microsoft says outside input shaped the draft, including experts in AI, law, ethics, philosophy, linguistics, and public policy, along with business leaders and public focus groups.
An appendix lays out nine paired “aligned” and “misaligned” example responses meant to illustrate the intended behaviors. Among them are a model that rolls back file transfers without authorization and a model that offers false reassurance during a family health crisis.
Related: New Warnings About the Risks of AI to Humanity Revive a Long-Running Debate
Related: The Race to Control AI and Protect What Makes Us Human
Related: CISOs Race to Control AI Agents Without Destroying Their Value
Text extracted automatically; images, tables and formatting may be missing. Original: https://www.securityweek.com/microsoft-ai-code-of-conduct-sets-cyberattack-boundaries-chain-of-command-safety-constraints/