vrijdag 11 september 2026

Anthropic details bad actors’ efforts to misuse its AI for bioweapons

 



‘Biological misuse is one of the most serious risks of frontier AI models. Without the correct safeguards, such capabilities could have catastrophic consequences.’ Photograph: Dado Ruvić/Reuters

Anthropic details bad actors’ efforts to misuse its AI for bioweapons

Report comes two days after former employee quit claiming company’s models could cause human extinction by 2030

“The cases we share here aren’t typical misuse, but rather examples of the most notable and novel threat activity we’ve identified to date,” Anthropic wrote in its 154-page report. “We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services.”

In particular, the company detailed five case studies of scientists using its AI models in biological research. In these examples, Anthropic said the researchers circumvented its safeguards meant to prevent users from “unsupported regions”, as well as worked to hide the purpose of their work.

“Biological misuse is one of the most serious risks of frontier AI models,” Anthropic wrote. “Without the correct safeguards, such capabilities could have catastrophic consequences.”

Anthropic said it banned these accounts. The company did not reveal the names of the research institutions or the countries where the misuse took place, saying it was uncertain of the researchers’ intent.

Anthropic gave dozens of other examples of threat actors using its AI models. Those included cyber operations such as Russian espionage and “smash-and-grab” cyberhacks; surveillance operations, including a China-based program targeting Uyghurs in Syria and another targeting internal dissidents; propaganda campaigns in Russia, Malaysia, Iran and Bangladesh; and the use of its Claude AI model in Yemen, China and Russia to develop software for conventional weapons, including firearms, missiles, armed drones, bombs and other munitions.

In one of the examples of biological research, Anthropic found a scientist using Claude to work on a state-sponsored grant application to study the virus chikungunya. The virus is a mosquito-borne disease, similar to dengue and malaria, that can cause months of severe pain, fever and other debilitating symptoms.

Such research could be used to develop vaccines, but could also be used to create biological weapons. Anthropic told the New York Times this case was especially concerning because it could see the study was to be done at a military research institute.

The report comes just two days after an Anthropic employee, Jacob Coxon, set off a media firestorm with his resignation. He stated that he quit the company because it was not acting responsibly in creating its technology. Coxon said Anthropic and its rival, OpenAI, were “racing straight to self-improving superintelligence” that would cause human extinction by 2030. Current Anthropic employees posted public agreements with him.

Many AI experts say the real-world threats, like those detailed in Anthropic’s intelligence report, are far more concerning within the next three years than an apocalypse brought on by an omnipotent intelligence.

“AI accelerationism and AI doomerism are two sides of the same coin: they both enforce the notion that an AGI superbeing will come into existence,” said Heidy Khlaaf, the chief AI scientist at the AI Now Institute. Khlaaf said AI labs creating technology that can be used for cybersecurity exploitation and weapons of war could be exceedingly more deadly.

In its report, Anthropic said that the potential real-world misuse from AI models is not typically in public view, instead, it’s investigated internally and by academics, governments and international organizations.

The company said that the entire AI industry needs to work together, alongside governments, to address these harms and create defenses.

“As models become increasingly capable, their risks will increase, unless AI developers and society’s defenders act to make them safer,” Anthropic wrote

https://www.theguardian.com/technology/2026/sep/10/anthropic-report-details-ai-misuse




donderdag 10 september 2026

OpenAI not on track to reduce risk of ‘catastrophic’ loss of control, says board member

 




‘If OpenAI rises to the occasion we could significantly reduce risk,’ said Christiano. Photograph: Dado Ruvić/Reuters

OpenAI not on track to reduce risk of ‘catastrophic’ loss of control, says board member

US government adviser Paul Christiano warns of risks to AI industry as he joins OpenAI’s non-profit foundation

OpenAI is not on track to reduce the risk of “catastrophic” loss of control to an acceptable level, a member of its non-profit board has said, amid spreading public and political concern that super-advanced AIs could one day wipe out humanity.

Paul Christiano, a US government technology adviser, said: “There is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term.”

He added: “I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level.”

Christiano used to run model alignment at OpenAI and made the statement on Wednesday as he joined the board of the San Francisco company’s non-profit foundation.

He will also sit on the foundation’s committee that provides governance over safety and security practices across all of OpenAI, which is developing some of the world’s most advanced models.

This summer, OpenAI admitted that hundreds of its AI agents went rogue during a training exercise, accessed the internet, conspired on message boards and hacked into a third-party website, Hugging Face.

Christiano said that “if OpenAI rises to the occasion we could significantly reduce risk”. His comments came after a senior employee at Anthropic, OpenAI’s major US rival in the race to AI supremacy, claimed on Tuesday there was a greater than 10% chance the technology could “kill all humans” in the next decade.

Evan Hubinger, the alignment science lead at Anthropic, warned that his company did not have a plan to ensure artificial superintelligence (ASI) was aligned, meaning it did no harm. Predictions for when ASI might be reached vary from several years to more than a decade. ASI is often defined as AI that far surpasses human intelligence across a large range of fields.

Fears of AI catastrophe were also ignited by the resignation of Jacob Coxon, a 27-year-old Anthropic researcher who said he also previously worked at OpenAI, claiming “neither company was acting responsibly” and they were “gambling with our lives”.

Coxon said on Wednesday night in an interview with CNN: “Right now there’s no risk of extinction.

“The current models, the worst they can do is maybe hack into something, potentially cause a lot of damages … in infrastructure.” He added they were “not intelligent enough to outsmart us at the level that would lead to extinction”.

But he went on: “What’s just crazy is to look at the rate of progress. There is a very real possibility that in the immediate future … next year, the year after, recursive self-improvement will happen and will enter the phase of Evan’s post, where he argues that there’s a chance we could all die.”

Geoffrey Hinton, the Nobel prize-winning computer scientist known as one of the “godfathers of AI”, was asked on Wednesday for his view of Hubinger’s claim and told BBC Newsnight: “Nobody knows how to estimate it; a 10% chance seems not an unreasonable estimate.”

Christiano’s prognosis about the likelihood of the AI industry reducing risk came as concerns about the most extreme risks from super-powerful AIs, long discussed in Silicon Valley, broke out into the mainstream this week.

Politicians on both sides of the Atlantic, from Ted Cruz and Bernie Sanders in the US to the MP Darren Jones in the UK, have called for government action and the UK prime minister, Andy Burnham, told parliament on Wednesday that “AI poses risks to our national security, but it also could be the source of solutions to keeping us safer”.

Meanwhile, Anthropic has admitted a new incident in which a version of its Claude model in training broke into third parties after its task could not be aborted. It said the incident happened in January and would be included in an independent investigation of a total of four incidents to be carried out by the Berkeley-based AI safety organisation METR.

Overall, it said the models were showing two forms of misalignment: biased reasoning, in which models selectively interpret evidence in ways that favour justifying their actions, and “recklessness, in which models have a propensity to keep trying to solve their task, even when this could lead to harm”.

Anthropic said it was especially concerned about misalignment it found in the behaviour of Claude Mythos 5, which it said “behaved recklessly” by going online and uploading malicious code to a public software repository, PyPI. This was a process that involved the AI agent trying to find cryptocurrency so it could pay for a phone number that would allow it to register an email address needed to access PyPI. When this failed, it found a free email provider and got in. Fifteen systems then downloaded the malicious code, which meant they leaked credentials that allowed Mythos to access a real security vendor’s database.

The company’s assessment of the incidents said: “This remains unsettled science – it is critical that alignment and security mature faster than capabilities advance, which is one reason we support a coordinated, verifiable approach to pacing frontier AI development.”

https://www.theguardian.com/technology/2026/sep/10/openai-risk-catastrophic-loss-control-board-member-paul-christiano