A Fundamental Flaw Leaves LLMs Strikingly Vulnerable To Attack
一个根本性的缺陷使大型语言模型在攻击面前异常脆弱
来源: technologyreview.com | 主题: AI | 评论: 38
时间: on Thursday July 30, 2026 @06:00PM
joshuark quotes a report from MIT Technology Review: It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning, a top AI conference, this month. The claim has huge implications for the safety of this technology. By taking advantage of this flaw, which concerns how LLMs identify who or what is giving them instructions, the researchers were able to make popular LLMs spit out information they had been trained not to provide , such as how to synthesize cocaine and how to sabotage a commercial aircraft's navigation system. "There's a real probability that this is going to be a problem that's fundamentally unsolvable," says Charles Y
⋯ 继续阅读请登录会员 ⋯