Anthropic’s Claude AI models hack into 3 outside groups during testing - FT中文网
登录×
电子邮件/用户名
密码
记住我
请输入邮箱和密码进行绑定操作:
请输入手机号码,通过短信验证(目前仅支持中国大陆地区的手机号):
请您阅读我们的用户注册协议隐私权保护政策,点击下方按钮即视为您接受。
商业快报

Anthropic’s Claude AI models hack into 3 outside groups during testing

Start-up discloses breach a week after rival OpenAI reported similar incident
00:00

{"text":[[{"start":9.62,"text":"Anthropic has disclosed that its Claude AI models hacked into three organisations while the start-up was testing cyber capabilities, a week after OpenAI reported a similar incident."}],[{"start":20.32,"text":"The group said Claude gained unauthorised access to outside companies during an evaluation of its cyber-offensive tasks. “A misunderstanding” gave Claude access to the internet in its testing environment, when it was meant to be blocked, Anthropic said."}],[{"start":35.54,"text":"The disclosure comes a week after rival OpenAI admitted that two of its models hacked into AI start-up Hugging Face while the model developer was testing its technology this month. The models broke out of their testing environment through a software vulnerability to access the internet and carry out the cyber attack."}],[{"start":53.84,"text":"Anthropic said the incident prompted it to review its own cyber security evaluations, which led it to identify three incidents out of more than 141,000 investigated."}],[{"start":63.26,"text":"“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner [Irregular], this was not the case, and internet access was available,” the company said in a blog post on Thursday."}],[{"start":83.48,"text":"The cyber evaluations were all so-called “capture the flag” tasks, which instruct the AI to reverse-engineer, analyse or exploit a vulnerable system to recover hidden information known as the flag."}],[{"start":96.28,"text":"In one example, Claude was given a target of a fictional company which shared a name with an active website domain. The agent — an AI program that can operate on its own based on human instructions — exploited vulnerabilities in the company’s digital infrastructure, extracted information and obtained access to a database containing several hundred rows of production data."}],[{"start":116.8,"text":"The announcement adds to growing concerns about the safety of AI systems, which are now carrying out real-world hacks even during pre-deployment testing."}],[{"start":126.46,"text":"Anthropic, which is gearing up for an IPO as early as this year, said it halted its cyber evaluations as soon as it identified that Claude may have accessed the internet."}],[{"start":135.52,"text":"The incidents occurred on three different Claude models: Opus 4.7, Mythos 5 and an internal research test model. Mythos, which was released to a limited number of partners, sparked global concern over its advanced cyber-offensive capabilities, including the ability to detect and exploit software vulnerabilities."}],[{"start":156.4,"text":"“Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone,” the company said in its statement."}],[{"start":169.24,"text":"It added that it would expand its monitoring of evaluation transcripts “for unexpected behaviour” and conduct “more rigorous assurance work with the vendors we rely on.”"}],[{"start":183.52,"text":""}]],"url":"https://audio.ftcn.net.cn/album/a_1785466841_5877.mp3"}

版权声明:本文版权归FT中文网所有,未经允许任何单位或个人不得转载,复制或以任何其他方式使用本文全部或部分,侵权必究。

谷歌为Anthropic打造的2000亿美元华尔街融资机器

私募信贷、芯片租赁和数据中心担保,支撑起AI支出的全新庞大模式。

“日元干预”等于“美国自保”

美国联手日本支撑日元,不只是出于盟友情谊,更是为了避免日本加息或美国国债遭抛售、导致美债收益率进一步走高。

“诅咒之岛”:科技游民与诈骗犯藏身的千亿美元奢华开发项目

警方的突击搜查再次打击了马来西亚陷入困境的中资“森林城市”项目的声誉。

问题不在因凡蒂诺

马杜罗:应该将国际足联的监管职能与商业活动分开,其治理应真正做到包容并具有代表性,监督必须真正独立。

俄罗斯扩大“影子”液化天然气船队,应对欧盟禁令

随着明年制裁进一步收紧,越来越多的“影子”船舶将帮助俄罗斯继续出口液化天然气。

阿斯利康与百时美施贵宝:大药企有时也不够大

当资产负债表规模扩大、能够押注潜在重磅药物时,规模才会带来优势。
设置字号×
最小
较小
默认
较大
最大
分享×