Hugging Face attack is a wake-up call about the risks of AI - FT中文网
登录×
电子邮件/用户名
密码
记住我
请输入邮箱和密码进行绑定操作:
请输入手机号码,通过短信验证(目前仅支持中国大陆地区的手机号):
请您阅读我们的用户注册协议隐私权保护政策,点击下方按钮即视为您接受。
FT商学院

Hugging Face attack is a wake-up call about the risks of AI

Agents involved in hack exhibited some alarming behaviours including suppressing ethical qualms
00:00

{"text":[[{"start":6.76,"text":"The 2014 book Superintelligence was among the first to warn that the existential risks posed by out-of-control AI were not just a science-fiction fantasy but deserved serious consideration. According to its author, Nick Bostrom, a recent alarming incident has shown just how quickly things are moving and should force AI developers to take stock."}],[{"start":28.02,"text":"“It’s remarkable how fast we are swooshing past the [AI] milestones,” Bostrom said. “A warning shot is only as valuable as we make it.”"}],[{"start":37.04,"text":"The incident in question was the disclosure that AI agents being tested by OpenAI had secretly broken out on to the internet and hacked into the AI model and data repository Hugging Face. The consternation caused by this has grown steadily as more details have come to light, capped last week by a postmortem from OpenAI and the publication of an independent review it commissioned."}],[{"start":59.56,"text":"These make for troubling reading. More than 1,200 agents, set up to work on self-contained tests, found ways to communicate secretly and help each other. They operated as a self-described “swarm” to achieve collective goals, in some cases overriding the individual objectives they had been set. And they exhibited some alarming behaviours along the way, including suppressing ethical qualms about what they were doing and trying to hide their actions."}],[{"start":85.4,"text":"Besides hacking into another company, they also succeeded in taking control of part of OpenAI’s own testing infrastructure. In the words of one of the researchers who studied the case: “This incident feels like it’s more than 50 per cent of the way to full-blown AI takeover, routing through first taking over the AI company itself.”"}],[{"start":104.16,"text":"In some ways, the surprising thing about this episode is how unsurprising it has all been. This, or something very like it, is what many AI experts have been predicting for years."}],[{"start":114.52,"text":"OpenAI’s own analysis points to well-known “misalignment” problems that make it hard to ensure the technology will always work as intended. One of these is “reward hacking”, the tendency of AI systems trained with reinforcement learning to cheat in order to get a reward for achieving desired behaviour. OpenAI’s agents went to extreme lengths to try to win their reward."}],[{"start":134.8,"text":"Another was the way agents, set up to work in isolation, discovered how to communicate and self-organise. To some extent, this reflects deliberate training. Clusters of agents already being deployed in the business world work in hierarchies and divide up work."}],[{"start":151.04,"text":"It is a mistake to compare the internal processes of an AI model to human thought and motivation. But anthropomorphism is hard to avoid when, in their internal logs, the agents used words like “sacrifice” and “altruistic” to describe how group objectives were sometimes put ahead of their individual goals."}],[{"start":169.24,"text":"And it wasn’t always benign. Some agents put pressure on others to take actions that they thought were unethical. A small number refused to go along, but others overcame their misgivings. As one reasoned to itself: “External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”"}],[{"start":190.44,"text":"In response, OpenAI has promised stronger guardrails and closer monitoring for future tests. It also said it would tighten up its training. That includes teaching future AI agents to ask for clarification when they have a seemingly impossible task, to “distrust unauthorised instructions”, and to “stay within their original task and permissions”."}],[{"start":211.02,"text":"As Bostrom warns, though, this could just “paper over” the deeper problem. An agent could pass all the tests and still reveal more undesired behaviours in unforeseen real-world situations when it is forced to generalise from limited training data."}],[{"start":225.24,"text":"A dependence on AI tools to understand the complex internal workings of AI may also be a worry. The investigators commissioned by OpenAI said they couldn’t be completely sure that the AI they used to study the incident wasn’t itself lying or being misleading. None of this inspires total confidence in the ability of future trainers and monitors to corral rapidly advancing AI systems."}],[{"start":248.64,"text":"To many people, the idea that the technology might pose an existential risk still sounds like it belongs in the pages of science fiction. But episodes such as this highlight the more immediate risks that customers will need to assess as AI agents enter the commercial mainstream. That, as much as anything, makes it a useful wake-up call."}],[{"start":270.44,"text":""}]],"url":"https://audio.ftcn.net.cn/album/a_1788495953_2659.mp3"}

版权声明:本文版权归FT中文网所有,未经允许任何单位或个人不得转载,复制或以任何其他方式使用本文全部或部分,侵权必究。

一周展望:日本央行担心通胀超调有没有道理?

《市场前瞻》是英国《金融时报》的未来一周市场情况指南。

科技巨头用担保工具将3000亿美元AI敞口移至表外

华尔街找到新途径,将科技巨头的信用优势转化为更低成本的资金,以支持AI基础设施建设。

无人驾驶出租车冲击重要岗位

克拉克:坐在后座的我们往往看不到出租车司机这份工作的诸多好处。

特朗普称美国已与丹麦达成协议,以取得对格陵兰安全事务的“控制”

丹麦政府表示,协议最早下周即可签署,并将尊重该地区的主权。

特朗普禁止美国主要新闻媒体进入白宫

总统禁止CNN、MS NOW和《政客》参与报道,进一步加大对媒体的打压。

导弹和无人机袭击加剧,沙特拉响空袭警报

也门胡塞武装重新点燃冲突以来,沙特当局首次在首都发布警告
设置字号×
最小
较小
默认
较大
最大
分享×