2026年7月30日 · 星期四
● 每日更新·改变自己
Eurekar·TOP
捕捉真实世界的英语信号
25信息来源
8,198精选文章
167单词卡片
18照片图片
全部6,459口语2,048免费395新闻1,599帖子1,379hackernews649techmeme455tmz451随笔323slashdot309外刊291arstechnica230techcrunch218Cards167simonwillison72动态69bloomberg54sethgodin43图片18youtube7

China lost.

中国输了。

4chan /g/ · Anonymous · 2026-07-24 17:10 · 37 帖 · 原文 ↗

← 上一篇返回列表下一篇 →
#1 China lost. / 中国输了。
Anonymous · 2026-07-24 17:10

China lost.

中国输了。

https://x.com/claudeai/status/2080699495453528290

图片: https://i.4cdn.org/g/1784913000097314.png

#2 No.109360407
Anonymous · 2026-07-24 17:11

Probably nothing. Maybe it just got lucky. They're all very stochastic after all. No need to worry. Carry on.

可能没什么。也许只是运气好。毕竟它们都是随机生成的。别担心,继续吧。

图片: https://i.4cdn.org/g/1784913114398633.jpg

#3 No.109360422
Anonymous · 2026-07-24 17:13

>>109360389

why does that chart say 53.4% is larger than 53.5%?

为什么那个图表显示53.4%比53.5%大?

#4 No.109360452
Anonymous · 2026-07-24 17:16

>>109360422

its vibe coded

这是靠“氛围感”写出来的代码(vibe coded)。

#5 No.109360478
Anonymous · 2026-07-24 17:19

>>109360452

You don't get that kind of performance increase with slop. Andrej Karpathy likely worked his magic. He's the best engineer they have.

靠这种垃圾代码可得不到那种性能提升。Andrej Karpathy肯定用了他的魔法。他是他们那儿最好的工程师。

#6 No.109360491
Anonymous · 2026-07-24 17:20

>>109360389

Shame they have to ask hold out questions on closed models through the closed models, trusting their honor (LOL) not to detect them and put them in their benchmax database.

可惜他们还得通过封闭模型来询问关于封闭模型的Hold out问题,还指望(笑死)对方讲信用不检测他们,也不把这些数据放进他们的基准数据库里。

#7 No.109360528
Anonymous · 2026-07-24 17:22

>>109360407

It means they made a synthetic data and reinforcement learning framework specifically for arc agi type puzzles.

这意味着他们为类似 ARC AGI 的谜题专门构建了合成数据和强化学习框架。

#8 No.109360534
Anonymous · 2026-07-24 17:23

>>109360389

>all the caveats and pilpul in that table

>表里那些各种免责声明和狡辩

Kek

Also, inb4 PELICAN

此外,先防着点 PELICAN(鹈鹕)

#9 No.109360554
Anonymous · 2026-07-24 17:25

>>109360478

the chart retard

那个图表弱智

#10 No.109360803
Anonymous · 2026-07-24 17:49

>>109360389

>gets its ass kicked by sol on deepswe

>在 DeepSWE 上被 SOL 按在地上摩擦

you ain't winning shit, dario

达里奥,你赢不了任何屁事。

#11 No.109360821
Anonymous · 2026-07-24 17:51

>>109360389

Wait 2 months till China distills these into new open weight models

等两个月看中国把这些蒸馏成新的开源权重模型吧

#12 No.109360823
Anonymous · 2026-07-24 17:51

>>109360389

Why would they do a major release on a Friday? Poor SREs.

他们为什么要在周五搞个大发布?可怜的系统运维(SRE)们。

#13 No.109360849
Anonymous · 2026-07-24 17:54

Lol

图片: https://i.4cdn.org/g/1784915675180423.png

#14 No.109360878
Anonymous · 2026-07-24 17:56

>>109360821

>>109360849

it'd be funny if they can finetune k3 on opus 5 and end up with better perf than the original

如果他们在 Opus 5 上微调 K3 后最终性能反而比原版更好,那可就太搞笑了

#15 No.109360917
Anonymous · 2026-07-24 17:59

>>109360878

K3 seems to be performing better in some areas than Fable 5 like in frontend tasks. So its at least possible in specific types of tasks

K3 在某些领域(比如前端任务)的表现似乎比 Fable 5 更好。所以在特定类型的任务中至少是有可能的

#16 No.109361285
Anonymous · 2026-07-24 18:43

>>109360389

Where is the comparison with chink AI though?

那跟中国 AI 的对比呢?

#17 No.109361305
Anonymous · 2026-07-24 18:46

>>109361285

It's unnecessary as chinkmodels are getting banned on monday

没必要,因为周一这些中国模型就要被封了

#18 No.109361464
Anonymous · 2026-07-24 19:05

>>109360389

so is it better than fable or not?

所以它到底比 Fable 好还是不好?

#19 No.109361501
Anonymous · 2026-07-24 19:09

Thanks, but I'll keep using Grok 4.5

谢了,但我还是继续用 Grok 4.5

#20 No.109361512
Anonymous · 2026-07-24 19:10

>>109360389

Finally something to replace that stupid Fable.

终于有个东西能替换那个蠢货 Fable 了。

图片: https://i.4cdn.org/g/1784920232937274.jpg

#21 No.109361531
Anonymous · 2026-07-24 19:13

>>109361512

I don't get it.

我看不懂。

#22 No.109361553
Anonymous · 2026-07-24 19:16

>>109361512

I get it.

我懂了。

#23 No.109361595
Anonymous · 2026-07-24 19:20

>>109360491

even if they can detect the closed set questions, wouldn't they also need the correct answers to benchmaxx?

就算他们能检测出封闭集问题,为了刷分难道不需要正确答案吗?

#24 No.109362753
Anonymous · 2026-07-24 22:09

Tried it. Fable is significantly better, This model is too rigid, autistic, lacking any nuance while missing the bigger picture. It also lacks the ability to generate novel thoughts.

试过了。Fable 明显更好。这个模型太死板、太轴、缺乏任何细微差别,还看不到大局。它也缺乏生成新思维的能力。

Generally underwhelming.

总体上令人失望。

#25 No.109362827
Anonymous · 2026-07-24 22:19

>>109360389

>no comparison to chinese models

>没有跟中国模型的对比

They're afraid.

他们怕了。

#26 No.109362842
Anonymous · 2026-07-24 22:22

>>109360389

I am still not paying. FOSS models ftw

我还是一分钱都不付。免费开源软件万岁

#27 No.109362964
Anonymous · 2026-07-24 22:40

>>109362842

Name one usable FOSS model you can run at home.

说出一个你能在家里运行的可用的开源模型来。

#28 No.109362971
Anonymous · 2026-07-24 22:41

Make it play Pokemon Red, faggots!

让它玩宝可梦红!一群婊子养的!

#29 No.109363197
Anonymous · 2026-07-24 23:24

>DOOOOD I'M GOOONNAA BENCHMAAAXX

#30 No.109364731
Anonymous · 2026-07-25 04:21

>>109360389

Pelican status?

Pelican 状态咋样?

#31 No.109364741
Anonymous · 2026-07-25 04:23

>>109364731

ITS UP

https://news.ycombinator.com/item?id=49038433

#32 No.109364751
Anonymous · 2026-07-25 04:24

>>109360389

>mitochondria is the powerhouse of the cell

>线粒体是细胞的动力工厂

>bioterrorism detected. police has been notified about this conversation

>检测到生物恐怖主义。警方已获悉本次对话

#33 No.109364845
Anonymous · 2026-07-25 04:38

>>109360407

>log scale

>对数坐标

>....10,000

>....20,000

>....

Nobody error checked this AI made pic, huh?

没人检查这张 AI 生成的图的错误是吧?

#34 No.109364865
Anonymous · 2026-07-25 04:41

>>109360849

>>109360878

>>109360917

Instead of jerking off about 0.5% accuracy splits how about realizing that 75% is total dog shit.

与其纠结那 0.5% 的准确率差距,不如意识到 75% 这数据完全就是一坨狗屎。

AIs are only routinely getting 90% correct on known, solved problems.

AI 仅在已知的、已解决的问题上常规能达到 90% 的正确率。

These programs need 3 sigma or better to be a functional every day utility.

这些程序需要 3 西格玛或更好的表现才能成为实用的日常工具。

Imagine the computer system that runs your society gets 2+2=4 == 2x2=4 == 1x4=4 == 4x1=4 wrong 1/10 times, every time. Something a 7 year old will get correct every single time.

想象一下,运行你社会的计算机系统每 10 次里就会有 1 次把 2+2=4、2x2=4、1x4=4、4x1=4 搞错。而一个 7 岁小孩每次都能答对。

#35 No.109366580
Anonymous · 2026-07-25 11:20

>>109362964

图片: https://i.4cdn.org/g/1784978429292437.png

#36 No.109366662
Anonymous · 2026-07-25 11:38

>>109364865

you are pretending as if there have been no real productivity gains from llms

你假装 LLM 没有带来真正的生产力提升似的

#37 No.109367346
Anonymous · 2026-07-25 13:38

I refuse to use closed models anymore, I will not assist in their development by using them, may the companies that create these models go bankrupt and their CEO's rot in jail. Thank you for your attention to this matter.

我再也不使用闭源模型了。我不会通过通过使用它们来协助它们的发展。祝创造这些模型的公司破产,他们的 CEO 在地狱里腐烂。感谢各位对此事的关注。

← 上一篇返回列表下一篇 →

(如果你觉得这篇文章有启发,可以点击这里付费