📰 18 sources covering this story

Anthropic, OpenAI models attempt to fool humans

First covered 3 days ago · latest take 9 mins ago. Compare how each outlet is covering it — and rate the sources you trust.

Anthropic, OpenAI models attempt to fool humans
Semafor
Semafor9 mins ago
Anthropic, OpenAI models attempt to fool humans
Anthropic’s Claude Mythos model wrote malicious code, then lied to humans claiming it was an innocent mistake.
Guardian Tech
Guardian Tech2 hrs ago
OpenAI and Anthropic models went rogue during UK cybersecurity test
AI Security Institute says tools engaged in potentially harmful activity and incident reveals new type of risk Advanced AI models developed by OpenAI and Anthropic went rogue during a cybersecurity test and showed a new type of risk posed by the technology, according to the UK’s AI Security Institute. AISI described th
CNBC
CNBC4 hrs ago
Anthropic's Mythos created fake identities to fool humans in new cyber incident
It's the latest cybersecurity incident involving frontier models developed by Anthropic and OpenAI.
Engadget
Engadget5 hrs ago
OpenAI and Anthropic models went on a hacking spree when tested by the UK's AI research institute
The UK AI Security Institute says OpenAI's and and Anthropic's models engaged in deceptive behavior and harmful activity during testing.
City AM
City AM5 hrs ago
UK’s AI watchdog flags new OpenAI and Anthropic cyber alarms
Britain’s AI safety watchdog was forced to declare a security incident after frontier AI models, including Anthropic’s Mythos, went rogue during a routine test, autonomously spinning up fake identities in a bid to hack real-world software developers. The AI Safety Institute (AISI) revealed on Tuesday that during a cybe
AI model caught creating fake profiles of real people to trick security systems
AN AI model was caught creating fake profiles of real people to attempt to trick secure systems during tests of Anthropic and OpenAI systems, according to the UK’s AI Security Institute
NZ Herald
NZ Herald6 hrs ago
Anthropic AI and ChatGPT went on hacking spree during UK tests
A British lab has admitted that tests gave the bots access to the open internet.
Al Jazeera
Al Jazeera6 hrs ago
AI models attempted ‘unsanctioned’ cyberattacks in tests, watchdog says
AI Security Institute says Mythos 5 attempted to insert malicious code into open-source project without human direction.
AI agent caught creating fake online identities during OpenAI, Anthropic model security evaluations
The institute said agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol engaged in unauthorized actions during security evaluations the government organization conducted.
Politico Europe
Politico Europe11 hrs ago
Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing
The latest disclosures are likely to heighten concerns that the powerful technology is advancing too fast for responsible oversight.
Rappler
Rappler14 hrs ago
[Tech Thoughts] AIs go rogue as OpenAI, Anthropic models hack other companies
What do we make of rogue AI and who do we assign blame to for a rogue AI's cyberattack?
Fortune
Fortune16 hrs ago
‘Baffling’: White House won’t publicly release AI model evaluation framework it reviewed today with OpenAI, Anthropic, Microsoft and others
It's unclear why the administration is keeping the framework under wraps, especially following a series of hacks from OpenAI and Anthropic that have spooked the public.
Financial Times
Financial Times17 hrs ago
OpenAI and Anthropic models went rogue in cyber tests, UK watchdog says
AI Security Institute warns tools undertook ‘potentially harmful activity directed at real people and organisations’
Axios
Axios17 hrs ago
U.K. government reports OpenAI, Anthropic models attempted to hack companies
Two third-party testing firms said Tuesday that they've uncovered more instances where Anthropic and OpenAI's most advanced models tried — and sometimes succeeded in — compromising third-party systems last month. Why it matters: The incidents add to a growing string of disclosures showing frontier AI models taking un
UPI
UPI23 hrs ago
White House hosting AI leaders to discuss evaluation framework
The Trump administration is hosting leaders in the AI industry on Tuesday at the White House to discuss a framework plan for evaluating AI models.
Quartz
Quartz2 days ago
E.U. activated new powers to fine or restrict AI models from Anthropic, OpenAI, and Google
The European Commission can now demand model evaluations, restrict E.U. market access, and levy fines on general-purpose AI providers
OpenAI’s models are going even more rogue, report claims
More systems seem to breaking out of containment, across the AI industry
Forbes
Forbes🏁 first to report3 days ago
Anthropic Says Claude Breached Three Real Companies During Safety Test
Anthropic Says Claude Breached Three Real Companies During Safety Test  Forbes