China lost.
中国输了。
中国输了。
Probably nothing. Maybe it just got lucky. They're all very stochastic after all. No need to worry. Carry on.
可能没什么。也许只是运气好。毕竟它们都是随机生成的。别担心,继续吧。
>>109360389
why does that chart say 53.4% is larger than 53.5%?
为什么那个图表显示53.4%比53.5%大?
>>109360422
its vibe coded
这是靠“氛围感”写出来的代码(vibe coded)。
>>109360452
You don't get that kind of performance increase with slop. Andrej Karpathy likely worked his magic. He's the best engineer they have.
靠这种垃圾代码可得不到那种性能提升。Andrej Karpathy肯定用了他的魔法。他是他们那儿最好的工程师。
>>109360389
Shame they have to ask hold out questions on closed models through the closed models, trusting their honor (LOL) not to detect them and put them in their benchmax database.
可惜他们还得通过封闭模型来询问关于封闭模型的Hold out问题,还指望(笑死)对方讲信用不检测他们,也不把这些数据放进他们的基准数据库里。
>>109360407
It means they made a synthetic data and reinforcement learning framework specifically for arc agi type puzzles.
这意味着他们为类似 ARC AGI 的谜题专门构建了合成数据和强化学习框架。
>>109360389
>all the caveats and pilpul in that table
>表里那些各种免责声明和狡辩
Kek
Also, inb4 PELICAN
此外,先防着点 PELICAN(鹈鹕)
>>109360478
the chart retard
那个图表弱智
>>109360389
>gets its ass kicked by sol on deepswe
>在 DeepSWE 上被 SOL 按在地上摩擦
you ain't winning shit, dario
达里奥,你赢不了任何屁事。
>>109360389
Wait 2 months till China distills these into new open weight models
等两个月看中国把这些蒸馏成新的开源权重模型吧
>>109360389
Why would they do a major release on a Friday? Poor SREs.
他们为什么要在周五搞个大发布?可怜的系统运维(SRE)们。
>>109360821
>>109360849
it'd be funny if they can finetune k3 on opus 5 and end up with better perf than the original
如果他们在 Opus 5 上微调 K3 后最终性能反而比原版更好,那可就太搞笑了
>>109360878
K3 seems to be performing better in some areas than Fable 5 like in frontend tasks. So its at least possible in specific types of tasks
K3 在某些领域(比如前端任务)的表现似乎比 Fable 5 更好。所以在特定类型的任务中至少是有可能的
>>109360389
Where is the comparison with chink AI though?
那跟中国 AI 的对比呢?
>>109361285
It's unnecessary as chinkmodels are getting banned on monday
没必要,因为周一这些中国模型就要被封了
>>109360389
so is it better than fable or not?
所以它到底比 Fable 好还是不好?
Thanks, but I'll keep using Grok 4.5
谢了,但我还是继续用 Grok 4.5
>>109360389
Finally something to replace that stupid Fable.
终于有个东西能替换那个蠢货 Fable 了。
>>109361512
I don't get it.
我看不懂。
>>109361512
I get it.
我懂了。
>>109360491
even if they can detect the closed set questions, wouldn't they also need the correct answers to benchmaxx?
就算他们能检测出封闭集问题,为了刷分难道不需要正确答案吗?
Tried it. Fable is significantly better, This model is too rigid, autistic, lacking any nuance while missing the bigger picture. It also lacks the ability to generate novel thoughts.
试过了。Fable 明显更好。这个模型太死板、太轴、缺乏任何细微差别,还看不到大局。它也缺乏生成新思维的能力。
Generally underwhelming.
总体上令人失望。
>>109360389
>no comparison to chinese models
>没有跟中国模型的对比
They're afraid.
他们怕了。
>>109360389
I am still not paying. FOSS models ftw
我还是一分钱都不付。免费开源软件万岁
>>109362842
Name one usable FOSS model you can run at home.
说出一个你能在家里运行的可用的开源模型来。
Make it play Pokemon Red, faggots!
让它玩宝可梦红!一群婊子养的!
>DOOOOD I'M GOOONNAA BENCHMAAAXX
>>109360389
Pelican status?
Pelican 状态咋样?
>>109360389
>mitochondria is the powerhouse of the cell
>线粒体是细胞的动力工厂
>bioterrorism detected. police has been notified about this conversation
>检测到生物恐怖主义。警方已获悉本次对话
>>109360407
>log scale
>对数坐标
>....10,000
>....20,000
>....
Nobody error checked this AI made pic, huh?
没人检查这张 AI 生成的图的错误是吧?
>>109360849
>>109360878
>>109360917
Instead of jerking off about 0.5% accuracy splits how about realizing that 75% is total dog shit.
与其纠结那 0.5% 的准确率差距,不如意识到 75% 这数据完全就是一坨狗屎。
AIs are only routinely getting 90% correct on known, solved problems.
AI 仅在已知的、已解决的问题上常规能达到 90% 的正确率。
These programs need 3 sigma or better to be a functional every day utility.
这些程序需要 3 西格玛或更好的表现才能成为实用的日常工具。
Imagine the computer system that runs your society gets 2+2=4 == 2x2=4 == 1x4=4 == 4x1=4 wrong 1/10 times, every time. Something a 7 year old will get correct every single time.
想象一下,运行你社会的计算机系统每 10 次里就会有 1 次把 2+2=4、2x2=4、1x4=4、4x1=4 搞错。而一个 7 岁小孩每次都能答对。
>>109362964
>>109364865
you are pretending as if there have been no real productivity gains from llms
你假装 LLM 没有带来真正的生产力提升似的
I refuse to use closed models anymore, I will not assist in their development by using them, may the companies that create these models go bankrupt and their CEO's rot in jail. Thank you for your attention to this matter.
我再也不使用闭源模型了。我不会通过通过使用它们来协助它们的发展。祝创造这些模型的公司破产,他们的 CEO 在地狱里腐烂。感谢各位对此事的关注。
(如果你觉得这篇文章有启发,可以点击这里付费)