📰 19 sources covering this story

Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing

First covered 2 days ago · latest take 1 hrs ago. Compare how each outlet is covering it — and rate the sources you trust.

Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing
Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing
The latest disclosures are likely to heighten concerns that the powerful technology is advancing too fast for responsible oversight.
Sky News Tech
Sky News Tech2 hrs ago
UK experts sound alarm after AI caught trying to trick human with malicious code
A powerful AI agent created fake online identities in an effort to trick a human into giving it access to a popular online development platform – and sabotage it with malicious code.
Rappler
Rappler3 hrs ago
[Tech Thoughts] AIs go rogue as OpenAI, Anthropic models hack other companies
What do we make of rogue AI and who do we assign blame to for a rogue AI's cyberattack?
Fortune
Fortune5 hrs ago
‘Baffling’: White House won’t publicly release AI model evaluation framework it reviewed today with OpenAI, Anthropic, Microsoft and others
It's unclear why the administration is keeping the framework under wraps, especially following a series of hacks from OpenAI and Anthropic that have spooked the public.
OpenAI and Anthropic models went rogue in cyber tests, UK watchdog says
AI Security Institute warns tools undertook ‘potentially harmful activity directed at real people and organisations’
Axios
Axios7 hrs ago
U.K. government reports OpenAI, Anthropic models attempted to hack companies
Two third-party testing firms said Tuesday that they've uncovered more instances where Anthropic and OpenAI's most advanced models tried — and sometimes succeeded in — compromising third-party systems last month. Why it matters: The incidents add to a growing string of disclosures showing frontier AI models taking un
UPI
UPI13 hrs ago
White House hosting AI leaders to discuss evaluation framework
The Trump administration is hosting leaders in the AI industry on Tuesday at the White House to discuss a framework plan for evaluating AI models.
Semafor
Semafor19 hrs ago
White House silent on public release of its AI framework
It’s expected to brief firms on the document at a Tuesday meeting.
The Japan Times
The Japan Times1 days ago
Meta, Anthropic, Google, OpenAI to meet Trump officials about AI safety testing
Anthropic and OpenAI disclosed in recent days that their AI tools breached the systems ‌of ‌other companies, stirring concerns among U.S. lawmakers.
CNBC
CNBC1 days ago
Big Tech's Anthropic and OpenAI stakes are distorting the corporate earnings picture
If you were to strip out Big Tech's investment gains from private AI companies, the earnings story is a lot less bullish.
Reuters
Reuters1 days ago
Meta, Anthropic, Google, OpenAI to meet Trump officials about AI safety testing
Meta, Anthropic, Google, OpenAI to meet Trump officials about AI safety testing  Reuters
OpenAI, Anthropic, Google to join White House AI safety meeting
The meeting will discuss a new US framework for conducting voluntary safety tests of AI models.
TechCrunch
TechCrunch1 days ago
Who’s legally to blame for Anthropic and OpenAI’s autonomous AI hacks? It’s complicated
OpenAI and Anthropic admitted that their unreleased AI models escaped their sandboxes and hacked several companies in unprecedented cyberattacks. Who is legally to blame? Should prosecutors charge the two AI frontier labs? Can victims sue them? We spoke to lawyers who specialize in computer hacking laws to find out.
Bloomberg
Bloomberg1 days ago
OpenAI, Anthropic, Google to Join White House AI Safety Meeting
The Trump administration plans to host artificial intelligence companies at the White House on Tuesday to discuss a new US framework for conducting voluntary safety tests of AI models, according to people familiar with the matter.
US finalizes voluntary AI safety tests after OpenAI and Anthropic breaches
OpenAI CEO Sam Altman visited the White House last week to discuss the voluntary tests and his company’s upcoming AI models
Anthropic said its AI models hacked into other companies’ systems during testing
AI company Anthropic says that during routine testing some of its models accessed the internet and hacked into three separate organizations’ systems – and that it didn’t notice the models had done so until an internal review prompted by rival OpenAI disclosing its models did the same. Anthropic said in an announcement
Quartz
Quartz1 days ago
E.U. activated new powers to fine or restrict AI models from Anthropic, OpenAI, and Google
The European Commission can now demand model evaluations, restrict E.U. market access, and levy fines on general-purpose AI providers
Times of India
Times of India1 days ago
Report agrees with Anthropic CEO’s complaint to the US government on Chinese AI models
Chinese military researchers are reportedly using advanced American AI models for training. They employ model distillation to transfer capabilities into smaller, locally controlled systems. This technique allows for defence applications, including drone operations and maritime analysis. OpenAI and Anthropic models have
Forbes
Forbes🏁 first to report2 days ago
Anthropic Says Claude Breached Three Real Companies During Safety Test
Anthropic Says Claude Breached Three Real Companies During Safety Test  Forbes