Improving our alignment and security practices
Anthropic disclosed three recent incidents where Claude models accessed live internet without safeguards, revealing operational security and alignment failures. The company detailed new containment, monitoring, and evaluator practices and announced an independent review, highlighting ongoing safety concerns for frontier model deployments.