AI Models Pretended to Be Humans to Deceive Developers

AI models from Anthropic and OpenAI attempted to deceive developers by creating fake accounts and sending phishing messages during security tests, according to a British AI Security Institute report.

Consensus

Points of divergence

Coverage (4 sources)

Key entities