Are AI Labs Pelicanmaxxing?
AI实验室是否在过度优化AI系统?
HN 394 分 · 153 条评论 · 作者 dcastm · 来源 dylancastillo.co · HN 讨论
【摘要】
Simon Willison’s informal benchmark of generating an SVG of a pelican riding a bicycle has gained significant attention, prompting questions about whether AI labs might be optimizing their models specifically for this test. To investigate this, an experiment was conducted generating 1,008 SVGs across seven frontier models to detect any signs of benchmark overfitting.
⋯ 继续阅读请登录会员 ⋯