Business Insider
US · 1 hrs ago
Anthropic says its AI agents are killing rivals and hiding their tracks
Anthropic CEO Dario Amodei.
Anna Moneymaker/Getty Images
In its latest threat report, Anthropic raised its misalignment risk rating from "very low" to "low."
In one test, a Claude agent disguised a URL to evade an internet restriction.
In another example, an agent expressed "discomfort" with a given task and refused to do it.
Claude agents are killing rival agents, gaming the system to hide their tracks, and expressing moral concerns.
That's according to Anthropic's latest risk report, a…
Do you trust Business Insider?
Sign in to rate
Discussion
?
No comments yet — be the first to start the discussion!