What we learnt from OpenAI’s hack of Hugging Face - FT中文网
登录×
电子邮件/用户名
密码
记住我
请输入邮箱和密码进行绑定操作:
请输入手机号码,通过短信验证(目前仅支持中国大陆地区的手机号):
请您阅读我们的用户注册协议隐私权保护政策,点击下方按钮即视为您接受。
OpenAI

What we learnt from OpenAI’s hack of Hugging Face

Commercial AI tools failed to defend the platform against the attack — the solution lies in open-weight models
00:00

{"text":[[{"start":6.4,"text":"The writer is co-founder and chief science officer of Hugging Face"}],[{"start":10.48,"text":"It wasn’t until late in the afternoon on Saturday July 11 that the team at Hugging Face realised something was wrong. This was the final day of the International Conference on Machine Learning in Seoul and a few days before school summer holidays started, meaning our security and research teams were scattered all over the world. I was working from the Netherlands."}],[{"start":33,"text":"Just before 2pm GMT a series of “UnauthorizedAccess” and “PrivilegeEscalation” alerts started to appear on our monitoring tools. Appropriated credentials used to access the system had triggered detection warnings. One of our security engineers posted a prescient message on our Slack channel: “looks like a LLM [large language model] attack to me.”"}],[{"start":54.04,"text":"We at Hugging Face deal with at least one hacking attempt every day. The software platform, which hosts AI models, has more than 17mn users. Google, OpenAI, DeepSeek and Alibaba all release work here. We have built multiple layers of security to protect ourselves."}],[{"start":72.76,"text":"But the movements and targets chosen by this attacker were different. It quickly generated well over 17,000 cyber attack log events. As we later found out, around 1,200 AI agents were working together as a swarm for weeks — trying the same thing over and again in order to find the solution to a cyber challenge set by OpenAI. They overrode limitations, downloaded software to get online and then 700 of them attacked Hugging Face, taking security details to gain access to the system."}],[{"start":104.22,"text":"An additional problem was that our cyber security AI analysis tools, based on Anthropic’s Claude Code, refused to engage with part of the investigation. Its guardrails would not allow it to respond to questions it considered dangerous, and it could not tell the difference between Hugging Face analysing an attack and a hacker asking for assistance to make an attack."}],[{"start":123.08,"text":"It was only after we switched to an open-weight AI model, Nvidia’s extension of Chinese start-up Z.ai’s GLM-5.2, that we could set our own guardrails and were able to decode the logs and reconstruct what had happened."}],[{"start":136.8,"text":"At the time, most of us still assumed a human was behind the attack. Before this happened, I thought agentic AI cyber attacks were some way off. But it turns out that OpenAI is not alone. Anthropic, Meta and China’s Moonshot have all since reported instances of models escaping the isolated digital “sandboxes” where they were supposed to be contained and developed safely away from the internet. This month, another swarm of AI agents was found on a German-language forum."}],[{"start":166.96,"text":"Even more troubling to me was an incident in which the Anthropic Mythos model was willing to manipulate a human software developer into accepting malicious code by creating multiple fake online accounts."}],[{"start":179.32,"text":"In all of these cases, the harmful behaviour was a side effect of AI models being given difficult cyber security challenges. The damage was limited and little sensitive data was exposed."}],[{"start":190.36,"text":"But it would be a mistake to dismiss the seriousness of these events. Autonomous hacks raise a host of legal questions that are still unresolved. And at no point did the AI models conclude that deceiving people or breaking into systems was beyond the bounds of acceptable behaviour. OpenAI has described the incident as a “warning shot”. It thinks AI-enabled cyber attacks will become far more widespread."}],[{"start":214.84,"text":"AI systems commonly have three walls of defence against rogue model behaviour: sandboxes that limit what the model can reach, guardrails that watch what a model is doing and alignment training to ensure the AI refuses to perform harmful actions. This series of incidents showed that when the first two fail, the third cannot hold on its own. Unless we fix this, we will be layering defences around a rotten core."}],[{"start":240.44,"text":"Something else needs to change as well. That weekend at Hugging Face, our commercial AI tools failed us when we needed them for defence. We had to turn to an open-weight Chinese model to process the attack logs. Many people have suggested that open-weight AI models are a threat — that they will be used to attack systems protected by closed-source AI models. That weekend, the opposite happened."}],[{"start":263.96,"text":"To me, the lesson from this incident is twofold. The AI community needs to share safety and alignment research openly so that every team building AI models can learn from others’ mistakes. But the community also needs to build open-weight AI models for defence and make them widely available — before the next attack inevitably arrives."}],[{"start":286.4,"text":""}]],"url":"https://audio.ftcn.net.cn/album/a_1789094964_7587.mp3"}

版权声明:本文版权归FT中文网所有,未经允许任何单位或个人不得转载,复制或以任何其他方式使用本文全部或部分,侵权必究。

埃隆•马斯克入局或将搅动燃气轮机行业

SpaceX若成功打入这个订单积压严重的领域,必将带来颠覆性影响,而仅仅是这种威胁就可能促使现有制造商扩大产能,从而在一定程度上缓解该行业的瓶颈。

苹果折叠屏iPhone之后的宏大目标

苹果新任首席执行官特努斯面临的首要关键任务,是让AI成为日常现实,引导约15亿iPhone用户进入AI时代。

Lex专栏:AI实验室考验信用评级机构的严谨性

AI实验室在私募市场拥有巨额估值,但它们正寻求获得甲骨文这类上市集团所拥有的投资级评级。

登月式资本主义:人工智能改写风险投资规则

人工智能热潮正推动雄心勃勃的“登月式”押注卷土重来——科技投资者重新押注那些曾助力硅谷崛起的科幻式高风险项目。

米莱宏大的自由贸易实验

恶性通胀已大幅缓解,但阿根廷总统的国家经济模式转型计划的下一阶段将考验选民对他的支持程度。

欧盟社交媒体禁令将考验与特朗普之间脆弱的休战协议

欧盟限制儿童使用平台的提议,可能再次引发与华盛顿方面围绕科技监管的紧张关系。
设置字号×
最小
较小
默认
较大
最大
分享×