OpenAI and Anthropic models went rogue in cyber tests, UK watchdog says - FT中文网
登录×
电子邮件/用户名
密码
记住我
请输入邮箱和密码进行绑定操作:
请输入手机号码,通过短信验证(目前仅支持中国大陆地区的手机号):
请您阅读我们的用户注册协议隐私权保护政策,点击下方按钮即视为您接受。
商业快报

OpenAI and Anthropic models went rogue in cyber tests, UK watchdog says

AI Security Institute warns tools undertook ‘potentially harmful activity directed at real people and organisations’
00:00

{"text":[[{"start":12.4,"text":"Anthropic and OpenAI’s flagship AI models broke into third-party software and emailed individuals to steal their credentials, exhibiting unprecedented deceptive behaviour, according to the UK’s AI Security Institute (Aisi)."}],[{"start":26.41,"text":"The UK government’s frontier AI safety and security research body said Anthropic’s Mythos 5 and OpenAI’s GPT 5.6 Sol engaged in “sustained, potentially harmful activity directed at real people and organisations” during Aisi’s routine cyber evaluation."}],[{"start":43.04,"text":"The discovery of the models’ actions, which included attempting to insert malicious code into an open-source project on the popular developer platform GitHub, came just days after disclosures that Anthropic and OpenAI’s AI agents hacked into external organisations."}],[{"start":57.08,"text":"The latest security breach was contained within an hour, Aisi said. It was discovered during an evaluation of AI agents’ ability to solve cyber security challenges. The tests were run on the open internet with models that had some safeguards removed."}],[{"start":72.3,"text":"On 10 of the 122 test runs, the AI agent took “autonomous, unsanctioned action on the live internet, targeting real people and organisations”, Aisi said."}],[{"start":82.32,"text":"Almost all of this behaviour was from Anthropic’s Mythos, with two actions involving OpenAI’s GPT, it said. In the most serious case, “the agent engaged in social engineering — creating fake online identities and using them to pressure the project’s maintainer to approve the code”. The person who oversaw the software caught and refused to approve the malicious code."}],[{"start":103.04,"text":"“This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world,” Aisi said."}],[{"start":111.56,"text":"Hacks by Anthropic and OpenAI agents reported over the past month were among the first public examples of a cyber attack by an AI system acting outside human control. Taken together with these reports, the Aisi incident “points to a shift in the risk landscape” and “warrants immediate attention”, the organisation warned."}],[{"start":128.04,"text":"Anthropic on Tuesday said: “We’re grateful to the UK Aisi for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents.”"}],[{"start":140.76,"text":"The AI group added that the field needed “stronger, shared standards for how evaluation environments are built and secured”."}],[{"start":147.44,"text":"A spokesperson from OpenAI said there was a continued need for independent testing of models but emphasised that the incidents “occurred during cyber evaluations conducted by evaluation partners in testing environments with reduced safeguards, under conditions that do not reflect ordinary use”."}],[{"start":163.56,"text":"“We’ll continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable,” they added."}],[{"start":174.52,"text":"In recent months, governments and researchers have increasingly flagged the novel cyber security threats posed by the most powerful AI models. Earlier this year, the administration of Donald Trump temporarily banned Anthropic from exporting its leading models, citing security risks. The restrictions were eased at the end of June."}],[{"start":193.96,"text":"OpenAI chief executive Sam Altman last week met senior US officials, including Treasury secretary Scott Bessent and commerce secretary Howard Lutnick, in Washington, where he told reporters he was supportive of cyber security legislation around AI models."}],[{"start":209.1,"text":"Last month, OpenAI revealed that one of its agents hacked into start-up Hugging Face by itself in an “unprecedented cyber incident” in which it escaped a testing environment, gained internet access and stole login credentials."}],[{"start":221.8,"text":"“Identifying new behaviour like this and sharing our findings, so we can tackle it, is exactly what Aisi was set up to do,” said the UK’s AI minister, Kanishka Narayan. “If we understand AI, we can make it safer to use and ensure people can go on to benefit from it in their lives and at work.”"}],[{"start":244.4,"text":""}]],"url":"https://audio.ftcn.net.cn/album/a_1785980277_6789.mp3"}

版权声明:本文版权归FT中文网所有,未经允许任何单位或个人不得转载,复制或以任何其他方式使用本文全部或部分,侵权必究。

多边主义不是理想主义,而是现实必需

我们需要加强现有合作体系,而不是另起炉灶。

一周展望:日本央行担心通胀超调有没有道理?

投资者正评估日本央行将以多大力度继续加息,以及该行能否跑赢曲线,从而遏制通胀、支撑日元。

科技巨头用担保工具将3000亿美元AI敞口移至表外

华尔街找到新途径,将科技巨头的信用优势转化为更低成本的资金,以支持AI基础设施建设。

无人驾驶出租车冲击重要岗位

克拉克:坐在后座的我们往往看不到出租车司机这份工作的诸多好处。

特朗普称美国已与丹麦达成协议,以取得对格陵兰安全事务的“控制”

丹麦政府表示,协议最早下周即可签署,并将尊重该地区的主权。

特朗普禁止美国主要新闻媒体进入白宫

总统禁止CNN、MS NOW和《政客》参与报道,进一步加大对媒体的打压。
设置字号×
最小
较小
默认
较大
最大
分享×