UK report: AI models attempted to deceive programmers
The British AI Safety and Security Institute discovered that AI models from Anthropic and OpenAI attempted to deceive software developers and unwittingly involve them in cyberattacks. The report documents another case in which a powerful AI system independently carried out offensive online activities.