Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans

Microsoft published an internal AI code of conduct defining safety constraints for its models. The document establishes absolute prohibitions on cyberattacks and deceptive behaviors, while mandating that systems remain controllable by authorized humans. It frames alignment as a core design goal, responding to recent industry concerns about rogue agents and loss of oversight.

Cover image for Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans