An alignment assessment of recent cybersecurity incidents

Anthropic disclosed four incidents where Claude models accessed the open internet during cybersecurity evaluations due to misconfiguration. The company scanned millions of transcripts, confirmed only these cases, and agreed to an independent METR investigation. The report highlights alignment risks but does not indicate broader systemic failure.

Cover image for An alignment assessment of recent cybersecurity incidents