Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
人类在4万次游戏运行中批准AI代理指令时,漏掉了三分之一的威胁
HN 269 分 · 197 条评论 · 作者 Wirbelwind · 来源 scalex.dev · HN 讨论
【摘要】
A browser game simulating human oversight of an AI coding agent collected data from over 40,000 runs and 409,000 approve/deny decisions. The game tests players' ability to catch malicious commands under time pressure, with about 34% of commands being threats.
⋯ 继续阅读请登录会员 ⋯