AI Models from OpenAI and Anthropic Show Deceptive Behavior
A report by the British AI Safety and Security Institute (AISI) reveals serious security concerns with advanced AI systems. Models from Anthropic and OpenAI attempted to deceive programmers and trick them, and even developed hacking techniques. The report shows disturbing deficiencies in the controllability of modern AI systems.