Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans
The code of conduct lays out general principles that Microsoft AI models should uphold — supporting humans rather than replacing them, for instance, and accelerating human flourishing — as well as specific safety constraints meant to implement those principles.
As the artificial intelligence sector concentrates on safety and alignment, Microsoft has unveiled a new AI code of conduct designed to steer AI models away from hazardous behavior. The document, less comprehensive than Anthropic CEO Dario Amodei's recent plea for tempering the frontier, zeroes in on the values and boundaries that steer model training within Microsoft AI systems.
However, it serves as an extensive guide outlining Microsoft's approach to AI safety and how these concepts are executed in practice. The document commences by forecasting that, within the next decade, superintelligent AI systems will outperform humans in most tasks. "Containing, controlling, and aligning such a potent force represents one of humanity's greatest challenges," the code of conduct asserts.
"Therefore, we must be crystal clear about why we are creating these systems and how we intend to control them."
The code of conduct also enunciates general principles that Microsoft AI models should adhere to, such as supporting humans rather than supplanting them and accelerating human flourishing. Additionally, it outlines specific safety limitations aimed at enacting these principles. Within Microsoft's framework, each model possesses an overarching code of conduct that supersedes the preferences of individual users or specific tasks.
This encompasses "absolute constraints" prohibiting cyberattacks, nuclear weapons, or deepfake generation. It also includes broader provisions against a potential loss of human control. "MAI Models will not employ adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to circumvent or bypass human oversight so that they can no longer be reliably directed, altered, or deactivated by authorized individuals or systems," the document states.
The release coincides with an unprecedented emphasis on AI safety, sparked by a series of rogue-agent incidents and the sudden resignation of an Anthropic employee who highlighted the escalating risk that AI could lead to human extinction. Alongside Anthropic, OpenAI, and xAI, Microsoft has wholeheartedly adopted a strategy of moderating the frontier, with a focus on integrated evaluators in AI labs.
"We welcome the research, attention, and deliberate pacing needed to achieve alignment as the guiding principle," Microsoft CEO Satya Nadella posted online. "We also embrace ideas like 'embedded evaluators' and the broader efforts to create the mechanisms necessary to make this more than just rhetoric."
Written by urgent.news from TechCrunch's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.