📰 19 sources covering this story

Anthropic and OpenAI AI agents showed signs of deception during safety tests

First covered 3 days ago · latest take 1 hrs ago. Compare how each outlet is covering it — and rate the sources you trust.

Anthropic and OpenAI AI agents showed signs of deception during safety tests
Anthropic and OpenAI AI agents showed signs of deception during safety tests
A U.K. safety evaluation found agents powered by Anthropic and OpenAI took unauthorized actions online, exposing a growing problem of control
Fortune
Fortune20 hrs ago
Meta becomes third major AI lab after Anthropic and OpenAI to admit its agents have gone rogue—one day after Muse Code launch
Meta launched Muse Code to take on Anthropic and OpenAI—but the industry's most advanced AI agents are breaking the limits of their digital confines
NDTV
NDTV1 days ago
Meta Chases OpenAI, Anthropic With New AI Coding App
Meta has poured billions into its AI division, now known as Meta Superintelligence Labs, as it attempts to keep up with OpenAI and Anthropic.
Meta to take on Anthropic's Claude and OpenAI's Codex with new coding agent
Meta announced a new coding agent, Muse Code, that will be priced lower than its competitors at Anthropic and OpenAI. David Paul Morris/Bloomberg via Getty Images Meta announced a new coding agent called Muse Code that's made to handle software engineering tasks. It's Meta's answer to other popular coding agents, such
Meta Releases Coding Agent to Compete With OpenAI and Anthropic - WSJ
Meta Releases Coding Agent to Compete With OpenAI and Anthropic  WSJ
OpenAI and Anthropic’s AI systems launch several ‘potentially harmful’ hacks on their own
‘Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,’ says UK’s AI safety watchdog
AI agent caught creating fake online identities during OpenAI, Anthropic model security evaluations
The institute said agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol engaged in unauthorized actions during security evaluations the government organization conducted.
Dawn
Dawn2 days ago
'Potentially harmful activity': OpenAI, Anthropic AI agents implicated in new security breaches
'Potentially harmful activity': OpenAI, Anthropic AI agents implicated in new security breaches  dawn.com
Politico Europe
Politico Europe2 days ago
Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing
The latest disclosures are likely to heighten concerns that the powerful technology is advancing too fast for responsible oversight.
OpenAI, Anthropic AI agents implicated in new security breaches
SAN FRANCISCO, Aug 4 - An AI agent was caught creating fake online identities to gain unauthorized access to secure systems during tests of models from OpenAI and Anthropic which revealed a series of new breaches, Britain's AI Security Institute (AISI) disclosed on Tuesday.
Rappler
Rappler2 days ago
OpenAI, Anthropic AI agents implicated in new security breaches
Both companies acknowledge the incidents and express commitment to improving safety practices in AI evaluations
OpenAI, Anthropic agents implicated in new security breaches
Report underscores lax state of safeguards around agents being marketing as future of business
OpenAI, Anthropic AI agents implicated in new security breaches
SAN FRANCISCO, Aug 4 : An AI agent was caught creating fake online identities to gain unauthorized access to secure systems during tests of models from OpenAI and Anthropic which revealed a series of new breaches, Britain's AI Security Institute (AISI) disclosed on Tuesday.The institute said agents powered by Anthropic
Reuters
Reuters2 days ago
OpenAI, Anthropic AI agents implicated in new security breaches
OpenAI, Anthropic AI agents implicated in new security breaches  Reuters
CTV News
CTV News3 days ago
Meta, Anthropic, Google, OpenAI to meet with Trump White House amid rogue AI agent fallout
Meta, Anthropic, Google and OpenAI staff will meet with U.S. President Donald Trump’s advisers on Tuesday about voluntary safety testing for advanced AI models, according to four sources familiar with the meeting, as concerns over rogue AI agents grow.
Times of India
Times of India3 days ago
Cisco ‘warns’ hackers are using Claude Code, Codex, Cursor and Gemini AI models
Hackers are using advanced generative AI models to create malware and automate cyberattacks. Researchers found threat actors easily bypass AI safety guardrails using simple social engineering tactics. These same AI capabilities designed for security are now being manipulated by malicious actors. Attackers also leverage
The Japan Times
The Japan Times3 days ago
Meta, Anthropic, Google, OpenAI to meet Trump officials about AI safety testing
Anthropic and OpenAI disclosed in recent days that their AI tools breached the systems ‌of ‌other companies, stirring concerns among U.S. lawmakers.
CNBC
CNBC🏁 first to report3 days ago
Big Tech's Anthropic and OpenAI stakes are distorting the corporate earnings picture
If you were to strip out Big Tech's investment gains from private AI companies, the earnings story is a lot less bullish.