Business Insider
Business Insider
US · 1 hrs ago

Anthropic says its AI agents are killing rivals and hiding their tracks

Anthropic CEO Dario Amodei. Anna Moneymaker/Getty Images In its latest threat report, Anthropic raised its misalignment risk rating from "very low" to "low." In one test, a Claude agent disguised a URL to evade an internet restriction. In another example, an agent expressed "discomfort" with a given task and refused to do it. Claude agents are killing rival agents, gaming the system to hide their tracks, and expressing moral concerns. That's according to Anthropic's latest risk report, a…
Business Insider
Do you trust Business Insider?
Sign in to rate
Discussion
?

No comments yet — be the first to start the discussion!