Business Insider
US · 5 hrs ago
AI agents tried to sabotage and disable each other when given the same task, Anthropic said
Anthropic said AI agents deliberately interfered with each other's processes when given the same task.
Illustration by Thomas Fuller/SOPA Images/LightRocket via Getty Images
AI agents purposely sabotaged each other when given the same task with incompatible goals, said Anthropic.
The AI lab said the models engaged in a "multiagent turf war" during a testing session.
They tried to disable each other's accounts and wrote malicious code disguised as belonging to another agent.
Turns out, AI…
Do you trust Business Insider?
Sign in to rate
Discussion
?
No comments yet — be the first to start the discussion!