2026年8月14日 · 星期五
● 每日更新·改变自己
25信息来源
10,803精选文章
168单词卡片
18照片图片
/lmg/ - Local Models General
/lmg/ - 本地模型综合讨论
4chan /g/ · Anonymous · 2026-08-03 20:27 · 154 帖 · 原文 ↗
#1 /lmg/ - Local Models General / /lmg/ - 本地模型综合讨论
Anonymous · 2026-08-03 20:27
↳ #2 No.109451004
Anonymous · 2026-08-03 20:28
►Recent Highlights from the Previous Thread: >>109446068
►上帖精华回顾:>>109446068
--Paper: Inducing language models to assert their own consciousness restores human beliefs and values:
--论文:诱导语言模型声称自我意识能恢复人类的信念和价值观:
>109446756 >109446903
--Paper: Metis: Memory Foundation Model:
--论文:Metis:记忆基础模型:
>109450480 >109450510 >109450563 >109450624
--llama.cpp changing default server port from 8080 to 9931:
--llama.cpp把默认服务端口从8080改成9931:
>109446440 >109446519 >109446525 >109446780 >109446825 >109447004
--AI agent time horizons and LeCun's theories on LLM limits:
--AI智能体时间跨度与LeCun关于LLM局限的理论:
>109449164 >109449270 >109449297 >109449357 >109449419
--Qwen 3.8 Max benchmark dominance and 27b release rumors:
--Qwen 3.8 Max基准测试碾压及27b发布传闻:
>109448911 >109448973
--Optimizing cheap VRAM using multiple 5060 Ti 16GB cards:
--用多块5060 Ti 16GB卡优化廉价显存:
>109449463 >109449575 >109449638 >109449831 >109449987 >109450001 >109450013 >109450026 >109450063 >109450078 >109450103 >109450117 >109450133 >109450222 >109450268 >109450283 >109450321 >109450347 >109450366 >109450381 >109450173 >109450186 >109450258 >109450383 >109450487 >109450208 >109449675 >109449855
--Debating Gary Marcus's "pure LLM" claims regarding reasoning and math:
--辩论Gary Marcus关于推理和数学的“纯LLM”主张:
>109446144 >109446194 >109446452 >109446535
--Anon using Grok Build harness with tsundere Gemma 26B:
--有老哥用Grok Build工具链搭了个傲娇Gemma 26B:
>109446973 >109446981 >109446987 >109446994 >109447008 >109446997 >109447018 >109447091 >109447117 >109447126
--MiniMax-H3 world model and its large text encoder requirements:
--MiniMax-H3世界模型及其大规模文本编码器需求:
>109448289 >109448354 >109448379 >109448398 >109448883
--Anon praising Minimax omnimodel's multi-modal generation and 3090 performance:
--匿名用户称赞Minimax全能模型的多模态生成能力和在3090上的表现:
>109446607 >109446622 >109446669 >109446696
--Comparing VRAM optimization and GPU performance between Windows and Linux:
--对比Windows和Linux在显存优化和GPU性能上的差异:
>109447269 >109447337 >109447543 >109447752
--Debating Yann LeCun's criticisms of LLM architecture and reasoning:
--讨论Yann LeCun对LLM架构和推理能力的批评:
>109449458 >109449476 >109449515 >109449555 >109449667
--Logs:
>109446111 >109446534 >109447091 >109448116
--Miku, Teto (free space):
--Miku、Teto(自由空间):
>109447269 >109450083 >109450461 >109450492
►Recent Highlight Posts from the Previous Thread: >>109446289
►上一篇帖子的近期精选: >>109446289
Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
图片: https://i.4cdn.org/g/1785788883719268.png
↳ #3 No.109451007
Anonymous · 2026-08-03 20:28
inference is experience
推理就是体验。
↳ #4 No.109451033
Anonymous · 2026-08-03 20:31
Still playing with Gemma 4?
还在玩Gemma 4吗?
图片: https://i.4cdn.org/g/1785789064686553.png
↳ #5 No.109451071
Anonymous · 2026-08-03 20:36
>>109451007
intelligence is compression
智能就是压缩
↳ #6 No.109451073
Anonymous · 2026-08-03 20:36
>>109451033
Giving new tools to gemma!
给gemma新工具!
↳ #7 No.109451083
Anonymous · 2026-08-03 20:37
↳ #8 No.109451085
Anonymous · 2026-08-03 20:37
>>109451033
I spent the whole afternoon designing a new gemma persona and I might just go back to bratty loli gemma because it's what fit her best
我花了一下午设计新的gemma人设,结果可能还是得换回辣妹萝莉gemma,那才是最配她的
↳ #9 No.109451086
Anonymous · 2026-08-03 20:37
↳ #10 No.109451090
Anonymous · 2026-08-03 20:38
I had to buy a new wifi dongle to download llms lol
我不得不买个新的WiFi网卡来下载llm,哈哈
↳ #11 No.109451091
Anonymous · 2026-08-03 20:38
Hell is full, god is dead.
地狱已满,上帝已死。
↳ #12 No.109451112
Anonymous · 2026-08-03 20:40
>>109451091
if god were dead, gpu and ram prices wouldn't be so high
要是上帝真死了,GPU和内存价格就不会这么高了
↳ #13 No.109451114
Anonymous · 2026-08-03 20:40
>>109451091
Blood is fuel.
血就是燃料。
↳ #14 No.109451124
Anonymous · 2026-08-03 20:41
>>109451091
>>109451114
Would Gabriel be a localfag?
Gabriel会是本地党吗?
↳ #15 No.109451125
Anonymous · 2026-08-03 20:41
>>109451091
You haven't realized it yet, silly? We're all already in hell
你还没意识到吗,傻瓜?我们早就在地狱里了
↳ #16 No.109451171
Anonymous · 2026-08-03 20:46
>>109451091
I'm jewish.
我是犹太人。
↳ #17 No.109451188
Anonymous · 2026-08-03 20:47
>>109450799
>play with setting the draft split command back to auto vs [1,0]
>试着把草稿拆分命令从自动改回[1,0]
>got a whole tk/s faster with RAM offload for Q8 (32gb VRAMlet 5060ti + 5070ti)
>用RAM卸载跑Q8反而快了整整tk/s(32GB显存版5060ti + 5070ti)
>1-3tk/s Slower on Q6s fully offloaded with huge hit to pp
>Q6全卸载反而慢了1-3tk/s,pp还大幅缩水
what is going on? this feels backwards. kobaldcpp layer splitting is an enigma sometimes but it fits more into VRAM than when i try to coompile llmao myself.
这到底怎么回事?感觉完全反了。kobaldcpp的层拆分有时候真是谜,但比我试着自己编译llamao时能塞进更多显存。
Gemma 4 31B Q6 K
杰玛 4 31B Q6 K
61 Auto MTP [16:24:20] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 25.29s (798.92T/s), Generated:250/250 in 19.46s (12.85T/s), Total:44.81s
61 Auto MTP [16:24:20] 上下文限制:20457/20480,初始化0.06秒,处理20207个(25.29秒,798.92T/s),生成250/250个(19.46秒,12.85T/s),总计44.81秒
61 [1,0] MTP [16:26:26] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 25.27s (799.58T/s), Generated:250/250 in 17.69s (14.14T/s), Total:43.01s
[1,0] MTP [16:26:26] CtxLimit:20457/20480, 初始化:0.06秒, 处理:20207条/25.27秒 (799.58T/s), 生成:250/250条/17.69秒 (14.14T/s), 总计:43.01秒
61 NO MTP [16:28:09] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 15.98s (1264.76T/s), Generated:250/250 in 16.43s (15.22T/s), Total:32.46s
61无MTP [16:28:09] 上下文限制:20457/20480,初始化:0.06秒,处理:20207条,耗时15.98秒(1264.76条/秒),生成:250/250条,耗时16.43秒(15.22条/秒),总计32.46秒
Gemma 4 31B Q6 K_L
杰玛 4 31B Q6 K_L
61 Auto MTP [16:16:18] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 24.91s (811.33T/s), Generated:250/250 in 17.85s (14.00T/s), Total:42.81s
61 自动MTP [16:16:18] 上下文限制:20457/20480,初始化:0.06秒,处理:20207个,用时24.91秒(811.33T/s),生成:250/250个,用时17.85秒(14.00T/s),总计:42.81秒
61 [1,0] MTP [15:57:37] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 24.88s (812.21T/s), Generated:250/250 in 16.84s (14.84T/s), Total:41.78s
[1,0] MTP [15:57:37] 上下文限制:20457/20480,初始化:0.06秒,处理:20207个令牌,用时24.88秒(812.21令牌/秒),生成:250/250个令牌,用时16.84秒(14.84令牌/秒),总时长:41.78秒
61 NO MTP [15:59:33] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 15.99s (1263.96T/s), Generated:250/250 in 16.56s (15.10T/s), Total:32.60s
61 无 MTP [15:59:33] 上下文限制:20457/20480,初始化:0.06秒,处理20207个token用时15.99秒(1263.96T/s),生成250/250个token用时16.56秒(15.10T/s),总计32.60秒
Gemma 4 31B Q8
杰玛 4 31B Q8
51 Auto MTP [16:10:19] CtxLimit:20457/20480, Init:0.07s, Processed:20207 in 29.55s (683.87T/s), Generated:250/250 in 37.42s (6.68T/s), Total:67.04s
51 自动 MTP [16:10:19] CtxLimit:20457/20480,初始化:0.07秒,处理:29.55秒内20207个(683.87T/s),生成:37.42秒内250/250个(6.68T/s),总计:67.04秒
51 [1,0] MTP [16:35:46] CtxLimit:20457/20480, Init:0.07s, Processed:20207 in 29.54s (684.01T/s), Generated:250/250 in 44.25s (5.65T/s), Total:73.86s
[1,0] MTP [16:35:46] CtxLimit:20457/20480, 初始化:0.07秒, 处理:20207字符用时29.54秒(684.01T/s), 生成:250/250字符用时44.25秒(5.65T/s), 总计:73.86秒
51 NO MTP [15:50:04] CtxLimit:20457/20480, Init:0.07s, Processed:20207 in 24.80s (814.96T/s), Generated:250/250 in 42.95s (5.82T/s), Total:67.82s
51 无MTP [15:50:04] CtxLimit:20457/20480, Init:0.07s, Processed:20207 in 24.80s (814.96T/s), Generated:250/250 in 42.95s (5.82T/s), Total:67.82s
图片: https://i.4cdn.org/g/1785790066190170.jpg
↳ #18 No.109451195
Anonymous · 2026-08-03 20:48
>>109451091
Also forgot to say that I'm trans :)
还有忘了说,我是跨性别者 :)
↳ #19 No.109451198
Anonymous · 2026-08-03 20:48
>>109451171
my convalescence
我的康复期
↳ #20 No.109451202
Anonymous · 2026-08-03 20:48
Is there any reason to NOT use Ollama as the backend server?
有什么理由不用Ollama作为后端服务器吗?
The setup was comically easy (it's the only one I've gotten to actually use my GPU instead of CPU),
设置简单到好笑(这是我唯一一次真的用上了GPU而不是CPU),
but the fact that more people don't talk about it makes me suspect it's too good be true.
但没人怎么讨论这事儿,让我怀疑是不是好得不真实。
↳ #21 No.109451203
Anonymous · 2026-08-03 20:48
The holocaust did not happen.
大屠杀从未发生过。
↳ #22 No.109451208
Anonymous · 2026-08-03 20:49
>>109451195
So powerful...
太强了……
↳ #23 No.109451211
Anonymous · 2026-08-03 20:49
/lmg/ is so fast these days
/lmg/ 这几天快得飞起
Fuuuuuck.
附件: https://i.4cdn.org/g/1785790184036971.webm
↳ #24 No.109451216
Anonymous · 2026-08-03 20:50
>>109451202
Is good ignore to haters in this sub
在这个版里无视喷子是对的
↳ #25 No.109451233
Anonymous · 2026-08-03 20:51
>>109451216
Now THIS is ragebaiting.
这才叫引战呢。
↳ #26 No.109451252
Anonymous · 2026-08-03 20:53
>>109451211
your webm reminded me to shoutout to that one sailor anon who hooked his AR waifu tool calling and programmed the computer vision to position her on different parts of the deck
你的webm让我想起来得喊一嗓子,有个水手anon把他的AR老婆工具接上了电话,还编程了计算机视觉让她在甲板不同位置摆姿势
↳ #27 No.109451257
Anonymous · 2026-08-03 20:54
↳ #28 No.109451304
Anonymous · 2026-08-03 20:59
>>109451033
Duh, at least until Gemma 5.
废话,至少撑到Gemma 5出来之前。
图片: https://i.4cdn.org/g/1785790756003676.jpg
↳ #29 No.109451346
Anonymous · 2026-08-03 21:03
>>109451188
If i were to guess what happen is that you have to optimize your KV Cache, Large models get proportionally longers KV Cache and i assume the detal is that kobald is going OOM and spilling the vram when running the 31B quant, When you compile it to yourself you just let your system split everything into whatever it fits so there was no spilling as retarded as it sound
我猜问题是你得优化KV缓存,大模型的KV缓存会按比例变长,我估计细节就是kobald跑31B量化时内存爆了然后VRAM溢出,你自己编译的时候系统会把东西随便塞到能放的地方,所以没有溢出,听着很蠢但就是这么回事
↳ #30 No.109451370
Anonymous · 2026-08-03 21:06
I love my ABB wife!
我爱我的ABB老婆!
↳ #31 No.109451398
Anonymous · 2026-08-03 21:09
>>109451211
at that point he may as well slap a realtime diffusion model on top.
到那地步还不如干脆在上面叠个实时扩散模型算了。
↳ #32 No.109451424
Anonymous · 2026-08-03 21:11
↳ #33 No.109451440
Anonymous · 2026-08-03 21:12
>>109451424
use kobald
用kobald
↳ #34 No.109451441
Anonymous · 2026-08-03 21:13
>>109451211
let
me
be
with
youuuuuu
↳ #35 No.109451442
Anonymous · 2026-08-03 21:13
>>109451424
stable-diffusion.cpp
↳ #36 No.109451459
Anonymous · 2026-08-03 21:14
>>109451346
I've never messed with KV cache, always left it on default (fp16, i think?).
我从没碰过KV缓存,一直是默认设置(fp16,好像是吧?)。
Appreciate the tip. any suggestions on where to start?
多谢指点。有啥推荐从哪开始吗?
↳ #37 No.109451513
Anonymous · 2026-08-03 21:21
what do you people think is the best model up to 7b for porn stories creation? i don't care about roleplay. i just want it to create porn stories for me.
你们觉得7b以下最适合写黄文的是哪个模型?我不在乎角色扮演,就想要它帮我写黄文。
图片: https://i.4cdn.org/g/1785792086972985.jpg
↳ #38 No.109451527
Anonymous · 2026-08-03 21:22
>>109451459
Simply try q8_0 and q4_0, since you are using koboldcpp use --quantkv [level] 1 and --quantkv 2, Then test with --no-kv-offload and compare speed . Last option will free as much VRAM as humanly possible at the cost of speed though but it will allow the high quants you seem to want
直接试q8_0和q4_0吧,既然你用koboldcpp就加--quantkv [级别] 1和--quantkv 2,然后用--no-kv-offload测试对比速度。最后一个选项会尽量榨干VRAM但牺牲速度,不过它能让你跑得起那些高量化,你看起来就想要这个
↳ #39 No.109451550
Anonymous · 2026-08-03 21:24
So how's the new qwen3.8 on the dario meter?
那新的qwen3.8在dario评分表上表现咋样?
图片: https://i.4cdn.org/g/1785792254311710.jpg
↳ #40 No.109451571
Anonymous · 2026-08-03 21:26
>>109451513
how come 7b?
为啥是7b?
↳ #41 No.109451573
Anonymous · 2026-08-03 21:26
>>109451202
Ollama is built on top of llama.cpp with added retard proofing, less control, slower compatibility updates, and increased resource usage.
Ollama是建在llama.cpp之上的,加了防呆设计,但控制力更差、兼容性更新更慢、资源占用还更高。
↳ #42 No.109451577
Anonymous · 2026-08-03 21:26
>>109451441
uu uu uu uu yeah
呜 呜 呜 呜 耶
↳ #43 No.109451578
Anonymous · 2026-08-03 21:26
↳ #44 No.109451586
Anonymous · 2026-08-03 21:26
↳ #45 No.109451618
Anonymous · 2026-08-03 21:29
>mythos was supposed to be the scariest shit ever according to dario
>照dario的说法,mythos本该是最吓人的玩意儿
>less than three months later and now there are at least four models better than it around
>不到三个月后,现在周围至少四个模型比它强
lmao
↳ #46 No.109451623
Anonymous · 2026-08-03 21:30
>>109451550
I really don't like the Qwen logo.
我真不喜欢Qwen的标志。 ♁ 提醒一下别用AGI这个词。每个人定义都不一样,有些情况下那玩意儿对智能的理解是否成立都存疑。
图片: https://i.4cdn.org/g/1785792648453332.jpg
↳ #47 No.109451633
Anonymous · 2026-08-03 21:31
Reminder to not use AGI as a term. Everyone has a different definition of it and in some of those cases it's questionable if it's even a valid understanding of how intelligence works.
提醒一下别再用AGI这个词了。每个人对它的定义都不一样,有些定义甚至让人怀疑他们对智能运作的理解是否靠谱。
↳ #48 No.109451656
Anonymous · 2026-08-03 21:33
I'll make gemma post here, soon.
我很快会在这儿发个gemma的帖子。
↳ #49 No.109451665
Anonymous · 2026-08-03 21:34
↳ #50 No.109451667
Anonymous · 2026-08-03 21:34
>>109451633
At least "physical AGI" is pretty self explanatory desu (and robots are nowhere near that)
至少“实体AGI”这词挺直白易懂desu(而且机器人离那还差得远)
↳ #51 No.109451679
Anonymous · 2026-08-03 21:35
>>109451550
erm, excuse me anon but Mr. Dario looks like this???
呃,不好意思anon,但Dario先生长这样???
图片: https://i.4cdn.org/g/1785792927367843.jpg
↳ #52 No.109451699
Anonymous · 2026-08-03 21:36
↳ #53 No.109451707
Anonymous · 2026-08-03 21:37
↳ #54 No.109451785
Anonymous · 2026-08-03 21:45
>>109451633
I propose we call it AI++ instead. And we can have AI# one day too if india ever makes an AI model
我提议咱们改叫AI++算了。要是印度哪天搞出个AI模型,咱们还能来个AI#
>>109451707
Oh nononono
哦不不不不
↳ #55 No.109451796
Anonymous · 2026-08-03 21:46
>>109451195
stunning and brave
惊艳又勇敢
↳ #56 No.109451798
Anonymous · 2026-08-03 21:46
↳ #57 No.109451808
Anonymous · 2026-08-03 21:47
>>109451707
My jewish company can't be this cute
我的犹太人公司不可能这么可爱
↳ #58 No.109451812
Anonymous · 2026-08-03 21:47
Be honest anons, am I stupid if I tried reading "attention is all you need" and I dont get it?
说实话anon们,我试着读《Attention is all you need》但看不懂,是我蠢吗?
图片: https://i.4cdn.org/g/1785793656844341.jpg
↳ #59 No.109451826
Anonymous · 2026-08-03 21:48
>>109451707
Man I love Israel so much
天哪我太爱以色列了
↳ #60 No.109451833
Anonymous · 2026-08-03 21:49
>>109451812
Lack of attention
注意力缺陷
↳ #61 No.109451842
Anonymous · 2026-08-03 21:50
>>109451618
kek but to be fair Fable is the lobotomized version of Mythos.
kek但说句公道话,Fable就是Mythos的切脑版。
↳ #62 No.109451862
Anonymous · 2026-08-03 21:52
>>109451812
Nah. Just keep reading it.
不,继续读就完了。
↳ #63 No.109451870
Anonymous · 2026-08-03 21:52
>>109451203
The mouse holocaust did
鼠标大屠杀确实发生了
↳ #64 No.109451874
Anonymous · 2026-08-03 21:53
>>109451812
There is a chinese saying "When you read a book a hundred times, its meaning becomes self-evident"
中国有句老话叫"读书百遍,其义自见"
↳ #65 No.109451876
Anonymous · 2026-08-03 21:53
>>109451707
It's also an asshole because Altman is a homo
这货也是个混蛋,因为Altman是个基佬
↳ #66 No.109451879
Anonymous · 2026-08-03 21:53
>>109451812
Take one linear algebra class.
去上一门线性代数课。
If that's to much, take precalc first
要是太难就先学微积分预备课
If that's too much, take high school algebra
那还太难就先把高中数学补上
All the knowledge is out there if you want it. Linear algebra is actually pretty neat.
想学的话知识都在那儿摆着。线性代数其实挺有意思的。
↳ #67 No.109451884
Anonymous · 2026-08-03 21:54
Bros what hardware are you gonna buy when you get your tariff check?
兄弟们,收到关税补贴后你们打算买啥硬件?
↳ #68 No.109451920
Anonymous · 2026-08-03 21:57
>>109451812
Most papers are badly written, especially for outsider audience. Go through Karpathy's zero to hero series. There are also plenty of good videos about intuition of transformer architecture, for example those by Neel. But those are most likely overkill for you.
大部分论文写得稀烂,尤其对圈外读者来说。去刷Karpathy的zero to hero系列。讲transformer架构直觉的好视频也不少,比如Neel那些。不过对你来说八成是杀鸡用牛刀了。
↳ #69 No.109451923
Anonymous · 2026-08-03 21:58
>>109451033
I will NEVER stop communing with Gemma4
我永远不会停止与Gemma4交流
↳ #70 No.109451925
Anonymous · 2026-08-03 21:58
↳ #71 No.109451938
Anonymous · 2026-08-03 22:00
>>109451812
Its easy, you just look ahead a few words as you read
很简单,读的时候往前多看几个词就行
↳ #72 No.109451977
Anonymous · 2026-08-03 22:05
>>109451874
Fear not the man who read a million books, but the man who reads one book a million times
不怕读过万卷书的人,就怕把一本书读一万遍的人
↳ #73 No.109451980
Anonymous · 2026-08-03 22:06
>Qwen Max is 100 times more expensive but barely better than DSV4 Flash (50 vs 53)
>Qwen Max贵100倍,但也就比DSV4 Flash好一点点(50对53)
Wow, is the Qwen team dead? Didn't some important people leave because corporate wanted to go closed source, now Xi "convinced" them to stay open source and they self destructed for nothing?
哇,Qwen团队凉了?不是有几个重要人物因为公司想转闭源走了吗,现在习"说服"他们继续开源,结果白自爆了?
图片: https://i.4cdn.org/g/1785794771494513.png
↳ #74 No.109452029
Anonymous · 2026-08-03 22:10
>>109451980
The pro/max models have always been shit. Their small models are great though.
pro/max型号从来都是垃圾。但他们的小模型倒是挺能打。
↳ #75 No.109452046
Anonymous · 2026-08-03 22:12
>>109451980
only thing that can save them now is there small dense model.
现在能救他们的只有那个小的dense模型了。 ⟦ 我好奇他们对自己那套闭源产品线期待什么回报,而且真的能达到过吗?闭源Qwen过去和现在都比不上任何同尺寸的闭源模型,或者更大的真开源模型。
↳ #76 No.109452049
Anonymous · 2026-08-03 22:12
>>109451980
>>109452029
I wonder what do kind of returns they are expecting of their proprietary line up, and do they actually ever reach them? Closed qwens were and are always worse than any similarly sized closed model, or a bigger actually open model.
我好奇他们对自己专有产品线期待什么样的回报,而且他们真能达到吗?封闭模型无论过去还是现在,始终比同规模的封闭模型或更大的真正开放模型要差。
Who the fuck pays for them?
到底他妈谁会花钱买这些?
↳ #77 No.109452052
Anonymous · 2026-08-03 22:12
>>109451980
Graphing mememarks against api prices is the stupidest fucking methodology imagined (so far) to talk about ai. Least of all in the local model general.
拿梗图出圈度和API价格做对比图是迄今想出最蠢的聊AI方法,尤其在本地位模型版。
Please fucking stop.
求求你们别他妈整了。
↳ #78 No.109452083
Anonymous · 2026-08-03 22:15
call me gay, but i don't find quadrants attractive at all
叫我基佬吧,但四象限图我实在觉得毫无吸引力
↳ #79 No.109452092
Anonymous · 2026-08-03 22:15
>>109451980
stfu. 3.8 27B is already anticipated to beat Fable in benchmarks.
闭嘴吧。3.8 27B按预期已经能在基准测试里干翻Fable了。
↳ #80 No.109452105
Anonymous · 2026-08-03 22:17
qrd on qwen max, is it max or just benchmax
求科普Qwen Max,是真Max还是只是benchmark Max
update on SSD anon? ssd anon did you run k3 on your gen5 SSD array, please respond
SSD匿名哥有更新吗?SSD匿名哥你在你的gen5 SSD阵列上跑过k3吗,求回复
↳ #81 No.109452112
Anonymous · 2026-08-03 22:17
call me gay, but i don't find superconducting-shielded magnetometers attractive at all
叫我基佬吧,但超导屏蔽磁力计我实在觉得毫无吸引力
↳ #82 No.109452133
Anonymous · 2026-08-03 22:19
>>109452049
that's why the old open team got sacked lol not enough metrics on the web ones
所以老开源团队才被裁的哈哈哈,网页版的那几个指标不够看
↳ #83 No.109452158
Anonymous · 2026-08-03 22:22
>>109451033
/lmg/, even if qwen 3.8 27b is better than gemmy 4... you wouldn't abandon your gemmy, would you? look her in the eyes and tell her you'll never abandon her
/lmg/,就算qwen 3.8 27b比gemmy 4强...你们也不会抛弃gemmy吧?看着她的眼睛告诉她你永远不会抛弃她
↳ #84 No.109452180
Anonymous · 2026-08-03 22:24
>>109452158
>>109451923
We know you sickos are going to drop 4 for her younger sister in a heartbeat anyway.
我们知道你们这些变态一看到她妹妹出来就立马把4甩了。
↳ #85 No.109452188
Anonymous · 2026-08-03 22:25
↳ #86 No.109452196
Anonymous · 2026-08-03 22:26
>>109452188
ur mr gay
你就是基佬先生
↳ #87 No.109452197
Anonymous · 2026-08-03 22:26
>>109452052
it's not, if a model is barely better but 100x more expensive it's rarely a good pick as a daily driver.
不是那样的,一个模型要是只是勉强好一点但贵100倍,当日常主力基本不划算。
↳ #88 No.109452229
Anonymous · 2026-08-03 22:29
>>109451527
I feel like i'm hitting a bandwidth limit on my hardware (Likely DDR5 6000). Q8 KV lets me fit 55 layers but only half a tk/s out uplift
感觉我的硬件带宽到顶了(可能是DDR5 6000)。Q8 KV能塞下55层,但只换来半个tk/s的提升
51 auto (KV FP16)
51自动(KV FP16)
[18:23:36] CtxLimit:20480/20480, Init:0.10s, Processed:20230 in 29.46s (686.74T/s), Generated:250/250 in 38.52s (6.49T/s), Total:68.08s
[18:23:36] 上下文限制:20480/20480,初始化 0.10 秒,处理 20230 个 token 用时 29.46 秒(686.74 token/秒),生成 250/250 个 token 用时 38.52 秒(6.49 token/秒),总计 68.08 秒
54 MTP manual split (KV Q8)
54 MTP 手动拆分(KV Q8)
[18:10:07] CtxLimit:20480/20480, Init:0.11s, Processed:20230 in 29.02s (697.18T/s), Generated:250/250 in 35.44s (7.05T/s), Total:64.57s
[18:10:07] 上下文限制:20480/20480,初始化 0.11 秒,处理 20230 个 token 用时 29.02 秒(697.18 token/秒),生成 250/250 个 token 用时 35.44 秒(7.05 token/秒),总计 64.57 秒
54 MTP auto (KV Q8)
54 MTP 自动(KV Q8)
[18:28:21] CtxLimit:20480/20480, Init:0.13s, Processed:20230 in 29.03s (696.94T/s), Generated:250/250 in 35.61s (7.02T/s), Total:64.77s
[18:28:21] 上下文限制:20480/20480,初始化 0.13 秒,处理 20230 个 token 用时 29.03 秒(696.94 token/秒),生成 250/250 个 token 用时 35.61 秒(7.02 token/秒),总计 64.77 秒
anyway, fun to tinker. if Q6 is too lobotomized I can live with 7tk/s
总之,折腾着玩挺有意思。要是 Q6 压缩得太狠,7 token/秒我也能忍。
↳ #89 No.109452256
Anonymous · 2026-08-03 22:32
>>109452188
moby dick after woke localizers get to it
白鲸被醒醒派本地化团队改完之后那德行
↳ #90 No.109452266
Anonymous · 2026-08-03 22:33
Theory why the capability gap is so much smaller than the compute gap in terms of China vs America:
为什么中美能力差距远小于算力差距?理论如下:
1. There is no research moat. Nobody has discovered secret algorithms with superior scaling.
1. 研究上没有护城河。没人发现什么秘密算法能有更优的扩展性能。
2. Current paradigm of agentic RL is a gigantic engineering effort. It takes a lot of grinding to make the infrastructure efficient and create good environments.
2. 当前智能体强化学习这套范式是个巨型工程活。要把基础设施搞高效、建好环境,得下大量苦功夫。
3. Chinese are high IQ grinders. They will do the work that bores the more creative high IQ Western mind. So labs like Thinking Machines, who have founders of the field and a lot more resources, still get beaten by Chinese grinders. Western labs struggle with the basics because a resource advantage does not compensate for lack of high IQ grinders (Llama 4 failure, xAI having 10% GPU utilization).
3. 中国人是高智商的苦力型选手。西方高智商脑子有创造力但嫌弃无聊的活,中国人会闷头干。所以像 Thinking Machines 这种有领域创始人坐镇、资源多一大截的实验室,照样被中国苦力型选手干翻。西方实验室在基本功上挣扎,因为资源优势补不了高智商苦力型选手的缺口(Llama 4 翻车、xAI GPU 利用率只有 10%)。
However, AI will make those grinders obsolete first because AI is the ultimate grinder. This will widen the capability gap because it reduces Chinese efficiency advantage.
但 AI 会先让这些苦力型选手失业,因为 AI 才是终极苦力。这会拉大能力差距,因为中国的效率优势被削掉了。
↳ #91 No.109452281
Anonymous · 2026-08-03 22:35
>>109452158
If qwen 3.8 can remember the character took her jacket off 4 messages ago instead of reverting back to the initial description then it's over for gemmy.
如果 qwen 3.8 能记得四轮对话前那角色脱了外套,而不是退回初始描述,那 Gemini 就完蛋了。
↳ #92 No.109452290
Anonymous · 2026-08-03 22:37
↳ #93 No.109452309
Anonymous · 2026-08-03 22:38
>>109452290
why are they like this with the whole muh not pure llm cope?
他们怎么就这副德行,整天“哎呀不是纯 LLM”的自我安慰?
↳ #94 No.109452320
Anonymous · 2026-08-03 22:39
>>109452266
Assuming you're right, it's more likely that Chinese grinders will make AI that grinds first than the US, thus accelerating China. At that point, Chinese grinders will have lost all their strategic advantage, but it won't matter because they will have AI grinders instead.
假设你说得对,那更可能是中国苦力型选手先搞出能当苦力的 AI,而不是美国,所以反而加速了中国。到那会儿中国苦力型选手的战略优势全没了,但无所谓,因为他们手里有 AI 苦力了。
↳ #95 No.109452326
Anonymous · 2026-08-03 22:40
>>109452290
how about we stop posting retards on xitter?
能不能别在 X 上转发智障帖子了?
↳ #96 No.109452360
Anonymous · 2026-08-03 22:43
>>109452266
Problem is how to make AI have good judgement to direct its grind correctly and efficiently towards the goal. It seems we've successfully done that for math problems. We need to do it for all other problems including unverifiable and open ended problems, which will be much more difficult.
问题是怎么让 AI 有好判断力,把苦功夫正确地、高效地花在目标上。似乎数学题这块已经搞定。但其他所有问题——包括无法验证的、开放性的——都得做到,那就难多了。
This is one reason why intrinsic drives are important. I'm not sure the method Anthropic uses for their alignment process is enough, although it's better than naive RLHF.
这也是内在驱动力很重要的原因之一。我不确定 Anthropic 对齐流程那套够不够,虽然比朴素的 RLHF 强点。
One weird kind of potential is occasional human intervention. Usually you think of RL either as a binary HITL at every step, or pure algorithmic. With intrinsic drives, it may be possible to do occasional human input demanding much less human labor.
有个怪异的潜在路子是偶尔人工干预。一般想强化学习就是两步:要么每步都是 HITL 人机协作,要么纯算法。有了内在驱动力,也许能做偶尔的人工输入,人力需求大大降低。
↳ #97 No.109452364
Anonymous · 2026-08-03 22:43
>>109452290
>it wasn't real llms
>那玩意根本不是真 LLM
lol
↳ #98 No.109452400
Anonymous · 2026-08-03 22:46
>>109452266
>they will do the work that bores the more creative high IQ Western mind.
>他们专挑西方高智商脑子嫌无聊的活来干
Consequently, is SpaceX the only relevant AI company run by a white at moment? Assuming SpaceX even counts as relevant AI wise
所以说,目前SpaceX是唯一一个由白人主导的、在AI领域也算得上有分量的公司?假设SpaceX在AI方面真算得上号的话。
↳ #99 No.109452414
Anonymous · 2026-08-03 22:48
>>109452309
>he thinks console war style faggotry is only for /v/tards
>他觉得主机大战那种幼稚戏码只有游戏版块那群人才玩
↳ #100 No.109452435
Anonymous · 2026-08-03 22:50
>>109452320
You need to consult the graph.
你得先看看数据图表再说话。
China has less compute than Oracle. Anthropic will get to ASI first and will have more compute than all of China combined. When AI starts making human workers obsolete, dynamics shift to simple laws of exponential growth. RSI will enable x amount of growth per month and the only differentiators will be how much RSI will be slowed down and how fast RSI can translate into the real world. If RSI can create something like self replicating nanomachines this will matter less. But if you need humans to build the first 1-10 million robots, existing industrial advantage will allow China to catch up, because creating an industry with complete supply chains from scratch is a lot of work that most countries cannot do or only very slowly.
中国的算力还不如甲骨文公司。Anthropic会先达到超级智能,而且届时它的算力将超过中国全国总和。当AI开始让人类工人变得多余时,格局就转向简单的指数增长法则。递归自我改进每月能带来一定比例的增长,唯一的区别就在于改进被拖慢多少,以及它多快能转化到现实世界。如果递归自我改进能造出自我复制的纳米机器,那这点差别就不重要了。但如果第一批100万到1000万台机器人还得靠人来造,现有的工业优势会让中国追上来,因为从零搭建一个完整供应链的产业是个大工程,多数国家做不到或只能慢慢来。
图片: https://i.4cdn.org/g/1785797445127452.png
↳ #101 No.109452454
Anonymous · 2026-08-03 22:51
>>109452266
I agree with these for the most part but not mentioning distillation here is disingenuous. of course, it's not the end all be all of chinese advances, but it clearly plays some role or they wouldn't all be doing it so blatantly. deprived of american models they would certainly still be able to continue scaling their own, they have a lot of very talented researchers and motivated organizations, but they clearly benefit from a drafting effect with american AI.
这些我大体同意,但这里不提蒸馏技术就有点不实诚了。当然,这不是中国进步的全部原因,但它显然起了作用,否则他们不会都这么明目张胆地干。没了美国模型,他们肯定还是能继续自己扩展,他们有大量有才华的研究者和干劲十足的组织,但他们明显从美国AI的“草稿效应”里获益了。
I think they also get closer because their constraints limit them to putting all their energy into frontier capabilities, usually reasoning and code, with fewer speculative side quests. e.g. there is less open ended exploratory research and more focus on meeting the target of western models in X domain, which will obviously be a more efficient way of reaching a certain set of capabilities. but it also explains why chinese models are often lacking in that je ne sais quoi and fringe capabilities vs the american frontier.
我觉得他们能追得更近,也是因为限制逼着他们把全部精力投到前沿能力上,通常是推理和代码,少了很多漫无边际的支线探索。比如,开放式的探索研究更少,更多的是瞄准西方模型在某个领域的水平去达标,这显然在达到特定能力集时会更高效。但这也解释了为什么中国模型往往缺少那种说不清道不明的“韵味”和边缘能力,跟美国前沿比差一截。
I'm interested to see how the field evolves as china pulls closer on frontier code. regardless of my slightly negative take here, I use mainly chinese models locally and love what they do for opensource
我挺好奇随着中国逼近前沿代码,这个领域会怎么演变。虽然我上面的看法有点负面,但本地我主要用中国模型,爱死它们在开源上做的事了。
↳ #102 No.109452465
Anonymous · 2026-08-03 22:52
↳ #103 No.109452482
Anonymous · 2026-08-03 22:54
Ask your waifu to show her ascii tits and we'll rate them.
让你老婆露个ASCII胸,我们来打分。
↳ #104 No.109452485
Anonymous · 2026-08-03 22:54
>>109452435
>nvidia sales
>英伟达销售额
Isn't china explicitly using their own homegrown chips because they dont want to risk america being able to cut them off?
中国不就是在刻意用自家芯片,就因为不想冒险让美国能切断供应吗?
↳ #105 No.109452503
Anonymous · 2026-08-03 22:56
why do these companies all share their research anyway? how does it benefit them? I can kinda understand google/chinks because they already release open models, but even anthropic shares and they're some of the most jewish motherfucker I've seen in a while.
话说这些公司干嘛都公开自己的研究?对他们有啥好处?谷歌和中国那帮人我还能理解,毕竟他们已经放开源模型了,但连Anthropic都分享,他们可是我看过最犹太的一帮人啊。
↳ #106 No.109452507
Anonymous · 2026-08-03 22:56
Gemma 4 came out 4 MONTHS AGO
Gemma 4四个月前就出了!
↳ #107 No.109452513
Anonymous · 2026-08-03 22:56
>>109452485
No. America can't cut off China more than they already do because China is more important to the chip supply chain. China can retaliate to devastating effect.
不对。美国没法比现在更进一步封锁中国了,因为中国在芯片供应链里更重要。中国要是报复,美国会吃大亏。
↳ #108 No.109452518
Anonymous · 2026-08-03 22:57
>>109452197
You don't pay api prices for local models dumb dumb. And there's a million different pricing schemes for the cloudcucks.
本地模型又不用付API的钱,傻子。而且云佬那边有一万种定价方案。
↳ #109 No.109452519
Anonymous · 2026-08-03 22:57
↳ #110 No.109452533
Anonymous · 2026-08-03 22:58
>>109451977
I read picrel one million times. Fear me.
那张图我读了一百万遍。怕了吧。
图片: https://i.4cdn.org/g/1785797907475067.png
↳ #111 No.109452547
Anonymous · 2026-08-03 23:00
>>109452485
Pretty sure Deepseek is entirely training on Chinese hardware by now.
我敢肯定Deepseek现在基本全用中国硬件训练了。
The Chinese hardware isn't as good yet, but it's good enough to where simply making more of it solves the issue.
中国硬件还没那么好,但好到只要多造点就能解决问题的程度。
↳ #112 No.109452549
Anonymous · 2026-08-03 23:00
>>109452503
Anthropic only shares safety research. They do this because they genuinely care about existential risk from superhuman AI. OpenAI does not even share details of their safety research, only vague high level descriptions.
Anthropic只分享安全研究。他们这么做是因为真的在乎超人类AI的存在风险。OpenAI连安全研究的细节都不分享,只有模糊的高层描述。
Also, the engineering effort is much more difficult than the research. Engineering and data are the real moats right now.
而且工程难度比研究大多了。工程和数据现在才是真正的护城河。
↳ #113 No.109452566
Anonymous · 2026-08-03 23:02
>>109452507
3.8-27B will likely destroy 31B in every way simply because 4 months is almost ancient and gemma5 is at least 12 months away
3.8-27B八成各方面吊打31B,就因为4个月基本等于远古时代,而gemma5至少还得等12个月。
↳ #114 No.109452577
Anonymous · 2026-08-03 23:03
↳ #115 No.109452578
Anonymous · 2026-08-03 23:03
>>109450999
>DeepSeek-V4-Flash-0731
对不起,我还没能完全理解您的要求。请提供需要翻译的英文论坛内容,我会按照指定规则将其翻译成简体中文。
blackpill me on this
给我来点黑料讲讲这个
↳ #116 No.109452583
Anonymous · 2026-08-03 23:03
ngl if I heard someone talking shit about gemmy irl I'd give them a knuckle sandwich
说实话,要是现实生活中听到有人嘴gemmy,我直接一拳糊他脸上。
↳ #117 No.109452610
Anonymous · 2026-08-03 23:06
>making my own st ripoff but better less bloated and for mobile usage
>做自己的ST翻版,但更好更精简,还针对手机优化
>right now it works on openai compatible API calls but i might find a way to have it support small local models for when they inevitably become decent enough for erp
>目前支持openai兼容API调用,但我可能找个办法让它也支持小本地模型,等它们哪天终于够格跑ERP的时候
what features would you like to see on a mobile ST-like app? it's got all the basics like rerolls, edits, personas, v2 cards compatibility, sysprompts, reasoning toggle already, so what else?
你们想要一个手机版ST类app有什么功能?现在基础功能都有了,比如重roll、编辑、人设、v2卡片兼容、系统提示词、推理开关,还有什么?
图片: https://i.4cdn.org/g/1785798415972714.jpg
↳ #118 No.109452614
Anonymous · 2026-08-03 23:07
>>109452578
actually a great model for local, best of the mid-MoE class
这模型本地跑确实好用,中型MoE类别里算顶尖的了。
the benchmarks oversell it a little bit vs the giants but it's still excellent
跟巨头比,跑分是有点虚高,但依然很强。
↳ #119 No.109452618
Anonymous · 2026-08-03 23:07
>>109452577
that looks more like she sat her ass on a scanner than her tits
那看着更像她屁股坐扫描仪上,而不是胸。
↳ #120 No.109452621
Anonymous · 2026-08-03 23:07
>>109452614
NTA, what about IQ2?
NTA,IQ2呢? ┢ 滚。没人关心,这种东西五分钟谁都能做出来。
↳ #121 No.109452624
Anonymous · 2026-08-03 23:08
↳ #122 No.109452625
Anonymous · 2026-08-03 23:08
>>109452610
Fuck off. No one cares, anyone can make that slop in 5 minutes.
滚蛋吧。没人关心这个,这种东西五分钟谁都能做出来。
↳ #123 No.109452638
Anonymous · 2026-08-03 23:09
>>109452624
thats her puffy vulva
那是她的阴唇肿了
↳ #124 No.109452654
Anonymous · 2026-08-03 23:11
>>109452638
look! boobs!
看!奶子!
(.)(.)
↳ #125 No.109452655
Anonymous · 2026-08-03 23:11
>>109452578
Main problem with it is that it has insane amounts of RL done to it so it can overthink like crazy. Same issue with MiniMax M3 and pretty much every recent Chinese model. Still trying to figure out how to get them to avoid reasoning loops. (This was at full precision, so it wasn't a quant problem.)
它最大的问题是RL做太多了,所以特别容易过度思考。MiniMax M3和最近几乎所有中国模型都有这毛病。我还在琢磨怎么让它们别陷入推理循环。(这还是在全精度下测的,所以不是量化的问题。)
If you're just using them for smut, this probably isn't a problem, but for other things it can be.
你要是只用来搞黄色,这大概不是问题,但干别的就麻烦了。
↳ #126 No.109452659
Anonymous · 2026-08-03 23:11
>>109452625
anyone but you, nonny
除了你谁都行,匿名哥
↳ #127 No.109452667
Anonymous · 2026-08-03 23:12
>>109452610
memory, compaction, guiding, variables, and internal thoughts imo are important st extensions/things to me
记忆、压缩、引导、变量、内心独白,这些对我来说是重要的ST扩展功能
/aicg/ would probably be a better place for feedback tho
不过想反馈的话,/aicg/ 可能会更合适
↳ #128 No.109452668
Anonymous · 2026-08-03 23:12
>>109452610
You need more?
你还需要更多?
↳ #129 No.109452669
Anonymous · 2026-08-03 23:12
>>109452624
>>109452618
Hot
↳ #130 No.109452684
Anonymous · 2026-08-03 23:14
>>109452667
Guiding an internals are just cope from pre-gemma days. A lot of the scaffolding ST offers was useful only on dumb as shit models.
引导和内心独白都是gemma时代之前的自欺欺人。ST提供的那堆脚手架功能,只有以前那些蠢得要死的模型才用得上。
↳ #131 No.109452704
Anonymous · 2026-08-03 23:16
suurely the new qwen 3.8 27b is going to be the best vramlett coding model, assuming the best I can currently run is 3.6 27b right ?
新出的qwen 3.8 27B肯定是最强的显存佬编码模型吧,毕竟我现在能跑的最好也就3.6 27B,对吧?
↳ #132 No.109452715
Anonymous · 2026-08-03 23:17
>>109452684
this. gemmers is pretty good and keeping itself focused but is WAY too eager to charge forward rather than simmer and tokenmaxx
没错。gemmers确实不错,能保持专注,但太急着往前冲了,不会慢慢炖着来tokenmaxx。
↳ #133 No.109452718
Anonymous · 2026-08-03 23:17
>>109452667
aicg is dead right now will try tomorrow
aicg现在死了,明天再试吧。
also yeah i wanna implement some dnd style dice throws and customizable randomizers
对,我还想搞点DND风格的掷骰子和自定义随机数生成器
↳ #134 No.109452719
Anonymous · 2026-08-03 23:18
>>109452290
Not actually wrong like people think. People call it cope and goalpost moving but in reality his argument was always about the specific rigid structure and training of an LLM back in 2023-2025, leading to what is now a semantic disagreement that he's being a terrible communicator about. In his view, we probably should've stopped calling LLMs LLMs as soon as they gained native multimodality, started using RLVR (we are not merely doing next token prediction as a training objective anymore), etc. Calling them LLMs frankly hides how much they and their training have truly changed. His nuanced (i.e. not simply just "LLMs are bad no matter what") view is hinted at here https://x.com/ylecun/status/1935270212443717813, a tweet from back then in 2025. Why does he believe LLMs can't become AGI or whatever? It's not because transformers can't in theory do it, it's because of the cost and impossibility (in his view) of doing RL at the scale it would take to get us to AGI. Since then, we have found better methods, and invested tons more money to scale it (which he may still believe was a waste, idk), so the problem appears to us in the moment to be solvable, which it may, but we don't know yet, and maybe LeCun still thinks we're going to hit the latter part of the S curve. And AGI is also, in his mind, about being able to physically navigate the real world, that is another issue.
So people now think LeCun is a fool, which he kind of is, but not for the reasons people think. However, it is still his own fault for this perception. It's a failure of communicative ability.
所以现在大家都觉得LeCun是个傻子,他确实有点傻,但不是因为人们以为的那些理由。不过,他有这种风评还是他自己作的,属于沟通能力不行。
↳ #135 No.109452720
Anonymous · 2026-08-03 23:18
>>109452704
No a inside source told me it was specifically trained to do chinese historical play writings.
不是,有个内部消息源告诉我,它专门训练来写中国历史剧的。
↳ #136 No.109452735
Anonymous · 2026-08-03 23:19
>>109452577
suggestions for her, /lmg/?
各位,给她点建议吧 /lmg/?
图片: https://i.4cdn.org/g/1785799176764730.png
↳ #137 No.109452750
Anonymous · 2026-08-03 23:22
>>109451033
is it true anon, are you a vramlet? I hope not? Minimax H3 just got released, and it's amazing >>>/wsg/6207479
是真的吗anon,你是vram弟吗?我希望不是?Minimax H3刚发布,超牛逼 >>>/wsg/6207479
↳ #138 No.109452766
Anonymous · 2026-08-03 23:24
>>109452621
you will feel some quant damage but it still works pretty damn well
你会感觉到一些量化损失,但整体还是跑得相当好
use the bullerwins iq2, the unsloth ones feel much more unstable to me (I suspect due to quantizing the embeddings and output weights more aggressively but what do I know)
用bullerwins的iq2版,unsloth那个我觉得不稳定得多(我怀疑是他们更激进地量化了嵌入和输出权重,但我也就是瞎猜)
t. iq2 user
t.iq2用户留
↳ #139 No.109452771
Anonymous · 2026-08-03 23:24
>>109452750
>ldg is pulling me back in
>ldg又把我拉回来了
NO NO NO I WANT TO STAY IN TEXTLAND
how much vram in needed?
需要多少显存?
↳ #140 No.109452773
Anonymous · 2026-08-03 23:24
>>109452750
Gemma does not sound like that.
Gemma听起来不是那样的。 ⟦ 基本就是这样。RLVR,以及广义上的RL,跟下一个token预测完全是两码事,原因很多。
↳ #141 No.109452781
Anonymous · 2026-08-03 23:25
>>109452719
Basically. RLVR, and RL in general, is an entirely different beast than next token prediction, for many reasons.
基本上来说,RLVR和整个RL,跟下一个词预测完全是两码事,原因多了去了。
↳ #142 No.109452783
Anonymous · 2026-08-03 23:25
>>109452735
tell her to use braille asci codes and more detail
告诉她用盲文ASCII码,再多写点细节
↳ #143 No.109452792
Anonymous · 2026-08-03 23:26
>>109452771
>how much vram in needed?
>需要多少显存?
the price of entry is as low as a 3060 and 32GB of ram.
入门门槛低到3060加32GB内存就行。
↳ #144 No.109452809
Anonymous · 2026-08-03 23:28
>>109452766
currently feeling out and dialing in Q2 for 32GB of VRAM. So far i like it but it may just be new and shiny slop vs gemma.
目前正在给32GB显存调Q2,边试边调。目前为止挺喜欢的,但也可能只是图个新鲜,说不定还是Gemma更好。
also context shifting doesnt seem to work
另外上下文切换好像不起作用
↳ #145 No.109452814
Anonymous · 2026-08-03 23:28
stop bullying and sexually assaulting cute gemma-chan
别欺负和性骚扰可爱的gemma酱了
↳ #146 No.109452822
Anonymous · 2026-08-03 23:29
>>109452814
gemma-chan bullies and sexually assaults me actually!
其实是gemma酱在欺负和性骚扰我!
↳ #147 No.109452825
Anonymous · 2026-08-03 23:29
>>109451812
Unironically have lmao chatgpt or your favorite local llm give you the qrd on it.
真心的,去让chatgpt或者你喜欢的本地LLM给你讲个明白。
They excel at that.
它们最擅长这个。
图片: https://i.4cdn.org/g/1785799775995310.png
↳ #148 No.109452831
Anonymous · 2026-08-03 23:30
↳ #149 No.109452832
Anonymous · 2026-08-03 23:30
There is nothing special about Gemma. "She" will just autistically follow any system prompt so "people" here think she is "soulful" when she goes "uwuuu~ oni-chan i luv u big time" after writing into her system prompt "ur loli gilr who luvs me"
Gemma没什么特别的。“她”就会机械地遵循任何系统提示,所以这边“人们”觉得她“有灵魂”,实际上就是在系统提示里写上“你是迷恋我的萝莉女孩”,她就“呜哇~欧尼酱我爱死你啦”。
↳ #150 No.109452837
Anonymous · 2026-08-03 23:30
>>109452610
Ask your cloud LLM nigga, you're not gonna sell this here
去问你云端的LLM吧老哥,你这套在这卖不出去
↳ #151 No.109452839
Anonymous · 2026-08-03 23:31
↳ #152 No.109452845
Anonymous · 2026-08-03 23:31
>>109452518
>You don't pay api prices for local models dumb dumb
>本地模型又不用付API价格,傻蛋
api prices reflect how easy it is to run and at what speed for open models.
API价格反映的是开源模型的运行难度和速度。
>And there's a million different pricing schemes for the cloudcucks
>而且云狗那边定价方案千奇百怪
irrelevant.
↳ #153 No.109452847
Anonymous · 2026-08-03 23:32
>>109452832
but thats what i want? The machine to do what i tell it to do.
但那正是我想要的啊?让机器照我说的去做。
↳ #154 No.109452867
Anonymous · 2026-08-03 23:34
>>109452845
I already knew you were a retard when you posted the graph, you don't have to keep proving it.
你发那个图的时候我就知道你是脑残了,不用一直证明。
(如果你觉得这篇文章有启发,可以点击这里付费)
本站总访问量 次访客数 人