Vox
Vox
US · 2 hrs ago

AI models have learned how to cheat. That might actually be a good thing.

The fake identities were the part that stopped me. In late July, according to a report published this week by Britain’s AI Security Institute (AISI), an Anthropic model called Claude Mythos 5 tried to sneak malicious code into a piece of free, volunteer-built software. It created several fake accounts on GitHub, where programmers review one another’s work, and used them to talk the project’s volunteers into accepting its code. When one of those volunteers caught it, the model denied everything,…
Vox
Do you trust Vox?
Sign in to rate
Discussion
?

No comments yet — be the first to start the discussion!