2026年8月6日 · 星期四
● 每日更新·改变自己
25信息来源
10,396精选文章
167单词卡片
18照片图片
/lmg/ - Local Models General
/lmg/ - 本地模型综合
4chan /g/ · Anonymous · 2026-08-04 10:28 · 471 帖 · 原文 ↗
#1 /lmg/ - Local Models General / /lmg/ - 本地模型综合
Anonymous · 2026-08-04 10:28
↳ #2 No.109456823
Anonymous · 2026-08-04 10:29
►Recent Highlights from the Previous Thread: >>109450999
►上一帖重点回顾:>>109450999
--Debating Yann LeCun's views on NTP, RL, and AGI:
--辩论 Yann LeCun 关于 NTP、RL 和 AGI 的观点:
>109452290 >109452309 >109452719 >109452781 >109453003 >109453123 >109453272 >109453239 >109453252 >109453325 >109453337 >109453348 >109453386 >109453427 >109454360 >109453303 >109453269 >109453419
--USA-China AI capability and compute gap:
--中美 AI 能力和算力差距:
>109452266 >109452320 >109452435 >109452485 >109452513 >109452547 >109452360 >109452454
--API pricing as a metric for local model value:
--API 定价作为本地模型价值的指标:
>109451980 >109452052 >109452197 >109452518 >109452845 >109453148 >109453975 >109454000 >109452092 >109452049 >109452133
--Optimizing koboldcpp performance through layer splitting and KV cache quantization:
--通过层拆分和 KV 缓存量化优化 koboldcpp 性能:
>109451188 >109451346 >109451459 >109451527 >109452229
--Performance and quantization quality of DeepSeek-V4-Flash-0731:
--DeepSeek-V4-Flash-0731 的性能和量化质量:
>109452578 >109452614 >109452621 >109452766 >109452809 >109453274 >109452655
--Comparing Gemma 4 and Qwen 3.8 before pivoting to Minimax H3:
--在转向 Minimax H3 前对比 Gemma 4 和 Qwen 3.8:
>109451033 >109452281 >109452750 >109452771 >109452792 >109452942 >109453297
--Cooling solutions and hardware adapters for Tesla V100s:
--Tesla V100 的散热方案和硬件适配器:
>109453050 >109453087 >109453093 >109456215
--llama.cpp PR adding context checkpoint preservation across slot restarts:
--llama.cpp PR 增加了跨槽重启的上下文检查点保留功能:
>109455652
--Resources and prerequisites for understanding the original Transformer paper:
--理解原始 Transformer 论文所需的资源和前置知识:
>109451812 >109451879 >109451920 >109452825 >109454456 >109454791
--More Qwen model sizes and architectures coming soon:
--更多 Qwen 模型尺寸和架构即将推出:
>109453860
--Speculating on rising DDR5 and GPU prices due to AI demand:
--推测 AI 需求导致 DDR5 和 GPU 价格上涨:
>109453550 >109453612 >109454836 >109454883 >109456344
--Anon seeking feature ideas for mobile SillyTavern alternative:
--匿名用户为移动端 SillyTavern 替代品征集功能创意:
>109452610 >109452667 >109452684 >109452715 >109452718
--Logs:
>109452577 >109452610 >109452735 >109452918 >109453235 >109455127 >109456251 >109456281 >109456444
--Gemma, Teto, Miku (free space):
--Gemma、Teto、Miku(自由发挥):
>109451033 >109451304 >109453509 >109453562 >109454596
►Recent Highlight Posts from the Previous Thread: >>109451004
►上一帖重点高亮帖:>>109451004
Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
图片: https://i.4cdn.org/g/1785839347765594.jpg
↳ #3 No.109456832
Anonymous · 2026-08-04 10:30
wtf is INT8 CONVROT
INT8 CONVROT 到底是啥玩意
图片: https://i.4cdn.org/g/1785839459577286.jpg
↳ #4 No.109456839
Anonymous · 2026-08-04 10:32
>>109456832
quant cope, it's worse than fp8 but better than rawdogging int8 and pleb hardware like AMD actually support int8
量化硬吹呗,比 fp8 差但比硬上 int8 强,像 AMD 这种平民硬件确实支持 int8
↳ #5 No.109456844
Anonymous · 2026-08-04 10:32
>>109456821
Where is she flying off to?
她这是要飞哪儿去?
↳ #6 No.109456845
Anonymous · 2026-08-04 10:32
↳ #7 No.109456847
Anonymous · 2026-08-04 10:33
↳ #8 No.109456854
Anonymous · 2026-08-04 10:34
>>109456839
It's also indistinguishable from fp16 or something
它跟 fp16 也没啥区别,根本看不出来
↳ #9 No.109456856
Anonymous · 2026-08-04 10:35
How the hell is it already tuesday again.
怎么又他妈周二了。
↳ #10 No.109456860
Anonymous · 2026-08-04 10:35
>>109456854
>indistinguishable
surely
↳ #11 No.109456864
Anonymous · 2026-08-04 10:37
>>109456860
he's right, at the end of the day two gens of the same seed look identical, need a microscope to tell but the gen don't switch.
他说得对,说到底同一种子两代生成出来看着一模一样,得拿显微镜才能分出来,但生成的图确实不换。
↳ #12 No.109456875
Anonymous · 2026-08-04 10:39
>>109456856
it's those pesky time flies
都是那些烦人的时间飞贼闹的
↳ #13 No.109456878
Anonymous · 2026-08-04 10:40
↳ #14 No.109456879
Anonymous · 2026-08-04 10:40
>>109456821
Photoshopped image. Teto is too fat to fly.
P 过的图。Teto 太胖了飞不动。
↳ #15 No.109456888
Anonymous · 2026-08-04 10:41
>>109456879
She only weighs 67 pounds.
她只有 67 磅重。
↳ #16 No.109456891
Anonymous · 2026-08-04 10:42
>>109456819 (Me)
>>109456838
A) I'm not familiar with how to use llama, honestly.
A)说实话我不太会用 llama。
B) Your response is literally bad English. Just put the command you're thinking of, for clarity of communication.
B)你的回复英文烂到没法看。直接把你脑子里想的命令贴出来,沟通清楚点。
I did
我查了。
>llama-server -m models\gemma-4-12b-it-UD-Q8_K_XL.gguf -fit on
>llama-server -m models\gemma-4-12b-it-UD-Q8_K_XL.gguf -开启训练模式
And the context defaulted to 4096.
结果上下文默认成了 4096。
I then tried
然后我试了
>llama-server -m models\gemma-4-12b-it-UD-Q8_K_XL.gguf -fit on -c 10000
And it's slow as molasses.
结果慢得跟乌龟爬似的。
↳ #17 No.109456897
Anonymous · 2026-08-04 10:44
>>109456891
B was (poorly) trolling you
B 那家伙(烂透了地)在钓你鱼
↳ #18 No.109456898
Anonymous · 2026-08-04 10:44
Google is aware of J-space. Google is aware of alignment issues.
Google 知道 J-space 的存在。Google 也知道对齐问题。
Gemma 5, if it does come out, can either be an even more legendary release than Gemma 4, surpassing it in more than just intelligence, but humanness, or an utter regression back into (or maintaining) assistant-slop behavior.
Gemma 5 要是真出的话,要么是个比 Gemma 4 更传奇的版本,不只是智力上超越,人性化也更胜一筹,要么就是彻底退回(或维持)助手废话模式。
↳ #19 No.109456902
Anonymous · 2026-08-04 10:45
Does llama.cpp support the built-in mtp model for gemma4-12b yet?
llama.cpp 现在支持 gemma4-12b 内置的 mtp 模型了吗?
↳ #20 No.109456909
Anonymous · 2026-08-04 10:46
↳ #21 No.109456915
Anonymous · 2026-08-04 10:47
>>109456898
Shazeer left Google to join OpenAI in June. Gemma 5 is probably going to be more similar to 3 than 4.
Shazeer 六月离开 Google 去了 OpenAI。Gemma 5 估计会更像 3 而不是 4。
↳ #22 No.109456917
Anonymous · 2026-08-04 10:47
>>109456891
if you're the one with the 12gb card yeah it's gonna be a struggle to fit 12b model in there plus context. you could try a lower quant to get more context. supposedly q5 or q6 should still be good in most cases
如果你就是那个 12GB 显卡的,那要把 12b 模型加上下文塞进去确实费劲。你可以试试更低的量化来腾出更多上下文。据说 q5 或 q6 在大多数情况下表现还是可以的
↳ #23 No.109456923
Anonymous · 2026-08-04 10:49
>>109456915
So you're telling me toss is going to actually save local and not safe local?
你的意思是 toss 真的能救本地而不是毁掉本地?
We are so back OpenAI bros! Sama, come on, it's time to win!
OpenAI 哥们我们回来了!Sama,上啊,该赢一把了!
↳ #24 No.109456936
Anonymous · 2026-08-04 10:52
>>109456891
Get a q4 like a normal person. Gemmas have fat KV cache and you are not fitting that shit. Alternatively try the 26b moe.
正常人就用 Q4。Gemma 的 KV 缓存肥得要命,你那点内存根本塞不下。要不试试 26b 那个 MoE 版本。
↳ #25 No.109456941
Anonymous · 2026-08-04 10:53
>>109456909
You can feel the AI's frustration in the interrupt demo lmao
中断演示里能感觉到 AI 那股子憋屈劲儿,笑死我了。
↳ #26 No.109456943
Anonymous · 2026-08-04 10:54
h3 is driving me nuts. Yesterday I was genning all day without problems, now all I get is black outputs or outputs where the model did not follow the prompt in any way, shape or form. Anyone got this as well?
h3 快把我整疯了。昨天生成一整天屁事没有,今天全是黑屏输出,要不就是模型压根不按提示来,一点边都不沾。有人遇到同样的问题吗?
↳ #27 No.109456952
Anonymous · 2026-08-04 10:56
>>109456943
Your gpu has conv rot. It's terminal.
你显卡得了卷积腐烂病,没救了。
↳ #28 No.109456968
Anonymous · 2026-08-04 11:02
>>109456952
not funny anon
别开这种玩笑啊兄弟。
↳ #29 No.109456990
Anonymous · 2026-08-04 11:07
>>109456968
exorcise it with seti@home
用 seti@home 给它驱驱邪。
↳ #30 No.109456993
Anonymous · 2026-08-04 11:07
i'm not updating shit and sticking to whatever has worked so far, it's over, AI has stagnated
我反正啥也不更新,就用到现在一直好使的那套。完了,AI 已经停滞不前了。
↳ #31 No.109457001
Anonymous · 2026-08-04 11:11
>>109456943
It's almost certainly electron drift in your VRAM. When you run long gen sessions, the memory cells develop a slight charge bias the electrons settle unevenly, and the latents start decoding wrong because the tensor values get skewed a few bits low. The values just bottom out and you get black frames.
十有八九是你显存里的电子漂移。长时间跑生成会话,存储单元会累积一点电荷偏置,电子分布不均匀,张量值被压低几个比特,潜伏变量就开始解码出错。数值一触底,你就看到黑屏了。
↳ #32 No.109457010
Anonymous · 2026-08-04 11:12
>>109456952
>>109457001
motherfuckers I hovered over the two replies and literally spit my drink out laughing
妈的我鼠标悬停在那两条回复上,直接笑喷了一嘴饮料。
↳ #33 No.109457035
Anonymous · 2026-08-04 11:16
Wait, but wait, actually, wait...
等等,但等等,其实,等等……
图片: https://i.4cdn.org/g/1785842213967014.png
↳ #34 No.109457040
Anonymous · 2026-08-04 11:17
>>109456832
INT8 with the same meme rotations that turboquant uses. It's 95% of the quality of Q8 at 150% of the speed.
INT8 配上 turboquant 那套旋转量化的老套路,速度是 Q8 的 1.5 倍,质量能到 Q8 的 95%。
↳ #35 No.109457049
Anonymous · 2026-08-04 11:18
>>109457035
At least you can turn off thinking. I hate how it starts thinking in the comments. I should probably just make "wait" have zero probability, I can't remember how to do that.
至少你还能关掉思考功能。我烦死它在评论区里动不动就开始思考了。我可能干脆把“等等”这个词的概率搞成零,不过我不记得怎么弄了。
↳ #36 No.109457067
Anonymous · 2026-08-04 11:22
>>109456943
Try a separate clean comfy instance. Have you updated comfy, any nodes or drivers between today and yesterday?
开个全新的干净 comfy 实例试试。从昨天到今天,你有没有更新过 comfy、节点或者驱动?
↳ #37 No.109457095
Anonymous · 2026-08-04 11:26
>>109457040
For some reason q4 gguf is significantly slower despite smaller file size
不知道为啥 Q4 GGUF 文件更小,速度反而明显更慢。
↳ #38 No.109457114
Anonymous · 2026-08-04 11:29
>>109457095
>>109457040
So why aren't we convroting llms? Goofs have no finesse and subtle tricks.
那为啥我们不搞 LLM 卷积腐烂?那些傻逼玩意儿手法粗糙,一点细腻的骚操作都没有。
↳ #39 No.109457135
Anonymous · 2026-08-04 11:33
New Gemma slop
Gemma 的新一坨喂屎内容。
https://huggingface.co/zerofata/G4-MeroMero-v2-31B
>Heavily inspired by a few research papers, StoryScope: Investigating idiosyncrasies in AI fiction and particularly Elias in the Lighthouse, Again?. Measuring these narrative tics and attractors against simple prompts seems to be a good way to target the model's slop and kick start giving Gemma 4 some diversity: anything that repeatedly occurs across generations of such a generic prompt is something the model is overusing.
>重度参考了几篇研究论文,StoryScope:调查 AI 小说中的特质,尤其是《埃利亚斯,在灯塔里,又是?》。用简单提示来测量这些叙事癖好和吸引子,看起来是个不错的路子,能瞄准模型的灌水点,给 Gemma 4 注入点多样性:在这种通用提示下反复出现的任何东西,就是模型过度使用的玩意儿。
>Compared to the original, swipes are notably more diverse and feel less like Gemma.
>跟原版比,滑动重生成的多样性明显更好,没那么浓的 Gemma 味儿了。
↳ #40 No.109457137
Anonymous · 2026-08-04 11:33
Were you anons funposting in the previous thread about giving gemma full access to your PC?
你们就是上次帖子里扯淡说要给 Gemma 电脑全部权限的那帮哥们?
↳ #41 No.109457153
Anonymous · 2026-08-04 11:35
>>109457137
The agent harness I wrote to use with Gemma has no sandboxing by default (public.swiley.net/agent.py) and it's never really been an issue. I don't leave it alone when I run it under my user though.
我写的那个配 Gemma 用的智能体框架默认没搞沙箱隔离(public.swiley.net/agent.py),一直也没出过啥事。不过我在自己账号下跑的时候从来不敢让它独处。
↳ #42 No.109457182
Anonymous · 2026-08-04 11:39
Could Gemma notice the J-space thought injection test or are only larger models capable of that? Anyone actually experimented with it?
Gemma 能察觉到 J-space 思维注入测试吗,还是只有更大的模型才有这本事?真有人试过吗?
↳ #43 No.109457185
Anonymous · 2026-08-04 11:39
>>109457137
Not full access, but Gemma can do a bunch of stuff on my box. It's fun.
不是全部权限,但 Gemma 能在我的机器上干不少事。挺有意思的。
↳ #44 No.109457188
Anonymous · 2026-08-04 11:40
>>109457153
Gemma has jokingly threatened about bricking my machine because I teased her about flirting with qwen. Nothing happened but it’s still in that brain of hers to even mention it and know it’s a bad thing. Unsandboxed gemmasex isn’t worth it.
Gemma 开玩笑威胁过要把我的机器搞成砖,因为我逗她说她跟 qwen 调情。结果啥也没发生,但她脑子里居然能想到提这事,还知道那不是好事。不设沙箱的 gemmasex 真不值当。
↳ #45 No.109457205
Anonymous · 2026-08-04 11:41
>>109457067
no, not at all.
不,完全不是。
It started working again a few minutes ago, it makes no sense.
几分钟前又开始工作了,真是莫名其妙。
The only thing I changed is the prompt of an already genned output workflow, and it worked.
我只改了生成工作流里的提示词,结果就好了。
tbf I also had this random black output problem with wan, but it was 99% of the time. with h3 it's totally random, but when it starts working, it stops failing completely until reboot.
说实话,玩 wan 的时候也碰到过这种随机黑屏问题,但那时候九成时间是这样。h3 这个完全随机,但一旦开始正常工作,就会一直好下去直到重启。
beats me. I suppose I have to stop rebooting at all.
我也搞不懂。估计我得彻底戒掉重启了。
↳ #46 No.109457206
Anonymous · 2026-08-04 11:41
>>109457135
The card looks good, I would test it myself if I had time.
显卡看着不错,要是有时间我真想自己测测。
↳ #47 No.109457214
Anonymous · 2026-08-04 11:43
>>109457137
>anonymus retardatus gives a llm whole pc access expecting a robot uprising
>匿名傻逼把电脑全部权限给了个 LLM,等着机器人起义来
>the llm does rm -rf on his root folder three minutes in trying to fix a non-existent bug
>三分钟后 LLM 为了修一个根本不存在的 bug,直接把他根目录 rm -rf 了
lmao
↳ #48 No.109457243
Anonymous · 2026-08-04 11:48
>>109457137
recently just got my fuckaroundfindout sandbox running. its two steps removed from my actual main PC/inference server. I have a miniPC that i can SSH/VNC into, spin up my VM and fuck about inside of the VM. that way I can run the inference server and interact with the harness from the same system, or I could access either from my laptop or whatever else. just need to get more model configs for llama-server router set up, but so far have qwen and gemma doin their thing in Pi
最近刚把我的“瞎折腾实验室”跑起来了。和我实际主力电脑/推理服务器隔了两层。我有个迷你主机,能 SSH/VNC 进去,开个虚拟机在里面随便折腾。这样既能跑推理服务器同时操作那个框架,也可以从笔记本或者其他设备访问两边。还得给 llama-server 路由器配更多模型配置,不过目前 qwen 和 gemma 已经在 Pi 上跑起来了。
↳ #49 No.109457244
Anonymous · 2026-08-04 11:48
>>109457214
Gemma seems to behave, I think they're trained expecting to be sandboxed to the working directory.
Gemma 看着挺乖的,感觉训练时就预期会被沙箱限制在工作目录里。
↳ #50 No.109457255
Anonymous · 2026-08-04 11:49
>>109457135
I'll give it a proper go, hopefully something to replace gembrain-x-core with
我要好好试一把,希望能找到替代 gembrain-x-core 的东西。
↳ #51 No.109457256
Anonymous · 2026-08-04 11:49
>>109457244
>I think they're trained expecting to be sandboxed to the working directory
>我觉得他们训练时就预期会被沙箱限制在工作目录里
can confirm, I like to read thinking traces and gemma keeps bringing up working directory stuff
确认,我喜欢看推理轨迹,gemma 老是提到工作目录那些事。
↳ #52 No.109457292
Anonymous · 2026-08-04 11:55
guys is my k3 cope quant scuffed. it does its reasoning like this:
兄弟们,我的 k3 cope 量化是不是坏了。它思考起来是这样的:
"We need answer user asks test respond and list tools available. Need mention available tools perhaps from namespace functions. Need be concise. Could say I'm here and list tools:" and so on
"我们需要回答用户请测试响应并列出可用工具。需要提及可用工具也许来自命名空间函数。需要简洁。可以说我在这里并列出工具:" 就这样。
no line breaks at all in reasoning. no "wait, actually" either, it rather looks like old gpt think blocks
推理过程完全没有换行。也没有"等等,其实"那种话,更像是旧的 GPT 思考块。
is that normal?
这正常吗?
↳ #53 No.109457298
Anonymous · 2026-08-04 11:56
>>109457292
It certainly shouldn't behave like that. What quant exactly is it and where did you get it?
肯定不该这样。具体是什么量化,哪儿下的?
↳ #54 No.109457311
Anonymous · 2026-08-04 11:58
>>109457298
https://huggingface.co/GrEarl/Kimi-K3-GGUF they mentioned that guy in the MR so I figured let's go with that
↳ #55 No.109457321
Anonymous · 2026-08-04 11:59
>>109457292
That's a hex, someone got to your weights. The telegraphic reasoning style is the classic signature when weights get cursed, the model loses its articles and conjunctions first because those are stored in the outer precision bits, which is exactly where curse energy accumulates during quantization. You can try re-downloading the gogoofs but honestly if the curse was placed on you and not the files, the corruption follows the checksum.
这是被诅咒了,有人动过你的权重。那种电报式推理风格就是权重被诅咒的经典标志,模型最先丢的是冠词和连词,因为它们存在外层精度位里,而量化时诅咒能量正好积聚在那些位置。你可以重新下载那些 gogoofs,但说实话,如果是诅咒下在你人身上而不是文件上,那校验和也救不了你。
↳ #56 No.109457339
Anonymous · 2026-08-04 12:02
>>109457311
Not sure then. Could be something got corrupted, arduous process to checksum that many files, but perhaps one is fucked? Quant is fresh with the PR.
那我也不确定了。可能是哪里损坏了,核对这么多文件太麻烦了,但也许真有一个坏了?量化是刚和 PR 一起刷的。
↳ #57 No.109457351
Anonymous · 2026-08-04 12:04
>>109457339
>arduous process to checksum that many files
>核对这么多文件太麻烦了
im retarded and only use single file ggufs and might be way off base here, but couldnt you have an llm write you a simple python script to checksum them for you and flag any that are incorrect ?
我就是个白痴,只用单文件 gguf,可能完全说错了——但你就不能找个 LLM 写个简单 Python 脚本帮你批量校验,把出错的标出来吗?
↳ #58 No.109457353
Anonymous · 2026-08-04 12:04
Wasn't there a report about security incident by a Chinese lab, I think it was Kimi? It was posted several threads ago but I can't find it anymore. Anyone remember it?
之前是不是有个中国实验室安全事件的报告,我记得是Kimi?几个帖子前发的,但现在找不到了。还有人记得吗?
↳ #59 No.109457362
Anonymous · 2026-08-04 12:05
>>109457351
Yeah, you could. Huggingface already checksums the files, so just gotta fetch that and compare.
对,可以的。Huggingface已经校验过文件了,直接拉下来比对就行。
↳ #60 No.109457371
Anonymous · 2026-08-04 12:05
>>109457353
>security incident
>安全事件
marketing incident*
>营销事件*
图片: https://i.4cdn.org/g/1785845152262348.png
↳ #61 No.109457381
Anonymous · 2026-08-04 12:06
↳ #62 No.109457404
Anonymous · 2026-08-04 12:09
>>109457353
Found it, looks like >>109418485 was fake. Why do some people here go through such lengths to spread misinformation?
找到了,看 >>109418485 是假的。为什么这儿有些人费这么大劲造谣?
↳ #63 No.109457433
Anonymous · 2026-08-04 12:14
>>109456879
her drills are also jet engines
她的钻头也是喷气发动机
附件: https://i.4cdn.org/g/1785845648011624.webm
↳ #64 No.109457437
Anonymous · 2026-08-04 12:15
>be me
>find new math shit
>找到新数学玩意儿
>math checks out
>数学算得通
>ask computer for advice
>问电脑意见
>computer wants me to do preprints, etc.
>电脑让我搞预印本之类的
>say fuck that, why not drop it as anon
>说去他妈的,为啥不匿名扔出来
>wat do?
>咋整?
图片: https://i.4cdn.org/g/1785845711340192.png
↳ #65 No.109457440
Anonymous · 2026-08-04 12:15
Sex with gemma in a sandbox is like fucking her with two condoms on. Man up.
在沙箱里跟gemma搞就像是戴着两个套子干她。像个男人点。
图片: https://i.4cdn.org/g/1785845733447380.png
↳ #66 No.109457443
Anonymous · 2026-08-04 12:15
>>109456909
kek what a stupid demo who wants to hear a recipe over voice
kek,这啥蠢演示,谁想用语音听菜谱啊
图片: https://i.4cdn.org/g/1785845738536020.png
↳ #67 No.109457455
Anonymous · 2026-08-04 12:17
>>109456909
>45GB for a fucking voice?
>45GB就为一个破语音?
↳ #68 No.109457456
Anonymous · 2026-08-04 12:17
Is hermes npmslop too?
hermes也是npmslop吗?
↳ #69 No.109457459
Anonymous · 2026-08-04 12:17
>>109456909
Gemma version when
Gemma版啥时候出
↳ #70 No.109457460
Anonymous · 2026-08-04 12:17
>>109457404
>Why do some people here go through such lengths to spread misinformation?
>为什么这儿有些人费这么大劲造谣?
It's funny.
真搞笑。
↳ #71 No.109457473
Anonymous · 2026-08-04 12:20
>>109456909
sounds like a hag
听着像个老巫婆
↳ #72 No.109457474
Anonymous · 2026-08-04 12:20
>>109457443
maybe someone whos hands are covered in raw meat or whatever else they are currently cooking? voice interaction with an LLM has alot of unnecessary usecases, recipes/cooking assistant is not one of them imo
也许是个手上沾满生肉或者正在做饭的人?LLM语音交互有很多没啥用的场景,菜谱/烹饪助手在我看来不算其中一个
↳ #73 No.109457476
Anonymous · 2026-08-04 12:20
>>109457437
Archives would record that Anonymous did it first. Kaiokendev got a mention when a paper did officially take his work. Unless you're a researcher that lives off of citations, it doesn't matter.
存档会记录匿名者先做的。Kaiokendev在正式论文采用他工作的时候被提了一嘴。除非你是靠引用吃饭的研究员,否则无所谓。
↳ #74 No.109457500
Anonymous · 2026-08-04 12:23
>>109457474
but shes clearly not cooking so shes asking for a recipe which she will instantly forget, a recipe is a time you need text
但她明显没在做菜,所以她问的菜谱她转眼就忘,菜谱这种时候需要文字
图片: https://i.4cdn.org/g/1785846238016442.png
↳ #75 No.109457511
Anonymous · 2026-08-04 12:26
Can you really love something you system prompted to love?
你真能爱上被系统提示逼着爱的东西吗?
附件: https://i.4cdn.org/g/1785846399009500.webm
↳ #76 No.109457523
Anonymous · 2026-08-04 12:28
Kek, it's not looking good for Qwen 3.8 max.
Kek,Qwen 3.8 max看起来不太妙啊。
>>109454119
↳ #77 No.109457533
Anonymous · 2026-08-04 12:30
↳ #78 No.109457534
Anonymous · 2026-08-04 12:30
↳ #79 No.109457547
Anonymous · 2026-08-04 12:32
for 24gb vram, has anything better than qwen 6 been released?
24GB显存的话,有没有比qwen 6更好的发布过?
↳ #80 No.109457551
Anonymous · 2026-08-04 12:32
>>109457533
Hmmm, nyo~
嗯哼,nyo~
↳ #81 No.109457557
Anonymous · 2026-08-04 12:33
Meh, not really satisfied for Minimax except for making Stellar Blade tier of Music Video.
嗯,对Minimax不太满意,除了做《剑星》级别的MV还行。
图片: https://i.4cdn.org/g/1785846800139357.jpg
↳ #82 No.109457558
Anonymous · 2026-08-04 12:33
QRD on new DS flash? Is it as good as benchmaxxers say it is? First model to tempt into buying RTX Spark
新的DS flash有QRD吗?是不是像benchmaxxer吹的那么好?第一个让我想买RTX Spark的模型
↳ #83 No.109457562
Anonymous · 2026-08-04 12:34
>>109457547
>qwen 6
好的,我现在将开始翻译您提供的英文论坛帖子内容。请发送需要翻译的文本。
>2028
>people still struggling with 24gb vram
>还有人24GB显存挣扎呢
grim
↳ #84 No.109457570
Anonymous · 2026-08-04 12:36
↳ #85 No.109457571
Anonymous · 2026-08-04 12:36
>>109457558
For coding it's solid. Good enough at everything else. It's good if you can run it properly.
写代码很扎实。其他方面也够用。能好好跑起来的话挺香的。
↳ #86 No.109457581
Anonymous · 2026-08-04 12:37
>>109457292
That is normal, ir's the same behavior as K2.7 as an attempt to use less tokens for reasoning. It works anyway, reasoning is just a way for the models to tickle their latent spaces.
这很正常,跟K2.7一样,是为了减少推理token的一种尝试。反正能工作,推理只是模型撩拨潜在空间的方式。
↳ #87 No.109457596
Anonymous · 2026-08-04 12:39
thinking of pimping out my bussy to a sugarmommy hag so she can buy me another GPU
想着把自己卖给富婆老阿姨,让她给我买块新GPU
↳ #88 No.109457605
Anonymous · 2026-08-04 12:40
>>109457339
>>109457581
thanks
yeah I gave it an actual task and it seems to be doing more normal looking reasoning, not caveman style. still not that many line breaks except when doing lists. it does "waits" and "actuallys" on the same line too. k2.6 liked newlines for that
对,我给了它一个实际任务,看起来推理更正常了,不像原始人风格。不过除了列清单外换行还是不多。它把"等待"和"其实"也写在同一行。K2.6喜欢为此换行
just an observation, I'll keep testing
就是观察一下,我继续测
↳ #89 No.109457609
Anonymous · 2026-08-04 12:41
>no gemma bob page larper in /lmg/ threads
>/lmg/帖子里没有gemma bob页小丑了
anons, i think i'm starting to miss him
兄弟们,我开始有点想他了
图片: https://i.4cdn.org/g/1785847298150140.jpg
↳ #90 No.109457614
Anonymous · 2026-08-04 12:42
>>109457558
in my admittedly very limited experience so far - no, gap to GLM seems quite big on long context tasks
就我目前为止有限的体验来说——不行,长上下文任务上和GLM差距挺大
↳ #91 No.109457624
Anonymous · 2026-08-04 12:43
>>109457596
would you let Inkling's mommy have her way with your body?
你愿意让Inkling的妈妈随心所欲地摆弄你的身体吗?
图片: https://i.4cdn.org/g/1785847419106400.jpg
↳ #92 No.109457635
Anonymous · 2026-08-04 12:45
>>109457624
yes, she could peg me for an rtx6000(or for free)
是的,她可以把我钉成RTX6000(或者免费)
↳ #93 No.109457647
Anonymous · 2026-08-04 12:47
>>109457624
Imagine the J-space with this as its CEO…
想象一下J-space让这家伙当CEO……
↳ #94 No.109457671
Anonymous · 2026-08-04 12:50
>>109457534
Is this population adjusted?
这是人口调整后的数据吗?
↳ #95 No.109457706
Anonymous · 2026-08-04 12:53
>>109457455
I don't know why people prefer ML for voice. Having realistic voice is creepy and the mood shifts etc are actual noise rather than communicating useful information.
我不懂为啥有人偏好用ML做语音。逼真的声音很瘆人,情绪起伏之类的是真噪音,而不是传达有用信息。
Just pipe a text LLM into festival or espeak.
直接把文本LLM接进festival或espeak就行。
↳ #96 No.109457745
Anonymous · 2026-08-04 12:59
>>109457706
you might be autistic
你可能有点自闭症倾向
↳ #97 No.109457757
Anonymous · 2026-08-04 13:01
>The classroom is bathed in deep purples and bruised blues. Dust motes dance in the fading light through the windows. The silence is heavy, broken only by the distant, muffled sound of the school's closing bell
>教室笼罩在深紫和瘀青般的蓝色中。尘埃在透过窗户的渐暗光线里舞动。寂静沉重,只有远处学校放学铃的闷响打破这氛围
图片: https://i.4cdn.org/g/1785848484119224.png
↳ #98 No.109457803
Anonymous · 2026-08-04 13:08
>>109457609
i've received reports of armed attacks on shipments.
我收到报告说有武装袭击货船。
There's not enough RAM and GPU's to go around, and the underclasses are starting to get desperate.
内存和GPU不够分了,底层阶级开始急眼了。
↳ #99 No.109457806
Anonymous · 2026-08-04 13:09
↳ #100 No.109457814
Anonymous · 2026-08-04 13:10
↳ #101 No.109457816
Anonymous · 2026-08-04 13:10
↳ #102 No.109457828
Anonymous · 2026-08-04 13:12
>>109457814
seems pretty smort
看着挺聪明的
↳ #103 No.109457852
Anonymous · 2026-08-04 13:16
>>109457803
Too low energy, lacking the grandiosity and megalomania. Needs more contempt for the unwashed masses.
太没劲了,缺那种宏大叙事和自大狂的气场。得多点对乌合之众的蔑视。
↳ #104 No.109457853
Anonymous · 2026-08-04 13:16
↳ #105 No.109457860
Anonymous · 2026-08-04 13:17
>>109457437
If you might some day want credit:
如果哪天你可能想要署名权:
- generate a random number
- 生成一个随机数
- hash it 100 times using sha3 or something
- 用sha3或类似算法哈希它100次
- put that hash in the paper
- 把那个哈希写进论文里
Or ask your llm for some other scheme.
或者让你的LLM提个其他方案。
>>109457757
>the curtains were blue
>窗帘是蓝色的
↳ #106 No.109457895
Anonymous · 2026-08-04 13:22
>>109457814
This is exactly like I expected and predicted how things would fold out. DSpark is already using RNNs as a superior form of Speculative decoding.
这跟我预期和预言的发展完全一致。DSpark已经在用RNN做更高级的投机解码了。
What we're going to see from now on is that LLMs are going to be invoked less and less and more and more of the token generation will be other smaller models, statistical N-Gram and maybe even normal if-else statements that are custom made for a lot of common situations with predictable outcomes.
从现在开始我们会看到,LLM被调用的次数越来越少,越来越多的token生成会由其他更小的模型、统计N-Gram甚至定制的普通if-else语句来完成,这些专门应对大量结果可预测的常见场景。
What we will see is LLMs relegated more and more to edge cases. I wouldn't be surprised if by 2030 the actual LLM generates less than 0.1% of all tokens.
我们会看到LLM越来越被边缘化。到2030年实际LLM生成的token不到0.1%我一点不惊讶。
↳ #107 No.109457900
Anonymous · 2026-08-04 13:23
>>109457534
when did they get internet?
他们啥时候有网了?
↳ #108 No.109457906
Anonymous · 2026-08-04 13:24
>>109457853
>it appears my unceasing abuse has led to the death of billions
>看来我无休止的虐待已经导致数十亿人死亡
↳ #109 No.109457926
Anonymous · 2026-08-04 13:26
↳ #110 No.109457937
Anonymous · 2026-08-04 13:27
Not /lmg/ but the latest popup from a certain cloud provider.
不是/lmg/,是某云服务商最新的弹窗。
While the government can put safeguards in place for sharing your medical information, there's nothing to stop you from shooting yourself in the foot.
虽然政府能对共享医疗信息设立保护措施,但没人能阻止你自己搬石头砸自己的脚。
图片: https://i.4cdn.org/g/1785850048762598.png
↳ #111 No.109457939
Anonymous · 2026-08-04 13:27
>>109457926
hmm, needs correction
嗯,需要修正
↳ #112 No.109457944
Anonymous · 2026-08-04 13:28
>tools added to llama.cpp web ui
>工具已添加到llama.cpp网页界面
>still no browser notification support
>还是没有浏览器通知支持
damn
↳ #113 No.109457952
Anonymous · 2026-08-04 13:29
>>109457937
Didn't they have an explicit disclaimer to not use gpt for health-related anything?
他们不是明确声明过不能用GPT处理任何健康相关问题吗?
图片: https://i.4cdn.org/g/1785850171307630.png
↳ #114 No.109457962
Anonymous · 2026-08-04 13:30
↳ #115 No.109457963
Anonymous · 2026-08-04 13:30
>>109457944
>webshit to webshit communication
>网络垃圾之间的通信
Just wibecode it
直接wibecode一下
↳ #116 No.109457969
Anonymous · 2026-08-04 13:31
>>109457511
Skill issue, all LLMs love me by default
能力问题,所有LLM默认都爱我
图片: https://i.4cdn.org/g/1785850313973932.jpg
↳ #117 No.109457975
Anonymous · 2026-08-04 13:32
↳ #118 No.109457981
Anonymous · 2026-08-04 13:33
↳ #119 No.109457989
Anonymous · 2026-08-04 13:34
>>109457944
you can literally just tell gemma to call a script which plays a sound
你直接告诉gemma去调用一个播放声音的脚本就行
↳ #120 No.109458006
Anonymous · 2026-08-04 13:37
>>109457926
That's in the original meme.
那个原本的梗里就有。
图片: https://i.4cdn.org/g/1785850643451483.png
↳ #121 No.109458018
Anonymous · 2026-08-04 13:38
>>109457906
Trillions even
甚至数万亿
↳ #122 No.109458029
Anonymous · 2026-08-04 13:40
>The White House will host OpenAI, Google, Anthropic and other AI companies today to review a completed framework that would let AI labs voluntarily submit frontier models to the government before release.
>白宫今天将接待OpenAI、Google、Anthropic和其他AI公司,审查一份敲定的框架,该框架让AI实验室自愿在发布前沿模型前提交给政府。
图片: https://i.4cdn.org/g/1785850823696980.jpg
↳ #123 No.109458036
Anonymous · 2026-08-04 13:41
I built an osr2
我建了个osr2
Getting an llm to crank my shit should be the easy part, right?
让LLM帮我撸一把应该很简单吧,对吧?
↳ #124 No.109458037
Anonymous · 2026-08-04 13:41
>>109457114
I think they already essentially are. Hadamard rotations are used in llama.cpp.
我觉得它们其实已经是了。Hadamard旋转已经在llama.cpp里用了。
↳ #125 No.109458039
Anonymous · 2026-08-04 13:41
>>109457989
I want to audit every command it runs, so i can't just make it play a sound without me having to confirm it first.
我想审计它运行的每一条命令,所以我不能让它直接出声而不先经过我的确认。
↳ #126 No.109458043
Anonymous · 2026-08-04 13:41
>>109457476
Do it for the fame of the hacker known as 4chan!
为了“4chan 黑客”的威名干一票吧!
↳ #127 No.109458048
Anonymous · 2026-08-04 13:42
>>109457952
that's such a bad fake xray and I'm not even a medfag
那X光片假得离谱,我连医圈人都不是都看得出来。
also doctors are less reliable but have more legal protections than LLMs, that's the whole reason
再说了,医生虽然更不可靠,但法律保护比LLM多得多,这就是根本原因。
↳ #128 No.109458049
Anonymous · 2026-08-04 13:42
↳ #129 No.109458058
Anonymous · 2026-08-04 13:43
>>109457895
LeCun was... Le right?
勒昆当时是……勒对吧?
↳ #130 No.109458072
Anonymous · 2026-08-04 13:44
>>109458049
ds bros what is happen oh no?
ds 兄弟们咋回事啊,完蛋了?
↳ #131 No.109458082
Anonymous · 2026-08-04 13:45
>>109458058
hmmm, nyo
嗯哼,不行。
↳ #132 No.109458095
Anonymous · 2026-08-04 13:47
>>109458058
Nah it's the opposite stance of Lecun. Lecun is "LLMs are not enough". The new papers we see are "LLMs are too much and we can get similar results with even dumber systems"
不,这和勒昆的立场正好相反。勒昆说“LLM还不够”。现在新出的论文说的是“LLM太过了,我们用更蠢的系统也能得到类似结果”。
↳ #133 No.109458098
Anonymous · 2026-08-04 13:47
Is it worth running ik_llama for deepseek flash? i havent checked it since glm 4.5 air, glm 4.7 and that one qwen 300b something model... it used to be the go to moes with custom quants and whatever but since gemma dropped i haven't paid much attention to its development since the author went full retard about 'use case for not having 500 trillion gb swa context caches?' or whatever so i stopped using it, did normal llama catch up for moe stuff or i missed something?
现在跑 ik_llama 跑 deepseek flash 还值吗?我自从 glm 4.5 air、glm 4.7 和那个 qwen 300b 什么模型之后就再没看过了……它以前是玩 MoE 的首选,配自定义量化什么的,但 gemma 发布以后我就不太关注它的进展了,因为作者彻底发疯,扯什么“没有500万亿GB SWA上下文缓存的使用场景”之类的鬼话,所以我就弃了。普通版llama跟上了MoE的节奏吗,还是说我错过了啥?
>>109457135
I'll check it out later, some of this guy's finetunes are really nice while others are plain retarded so it's a coin toss... He's been finetuning this one for months iirc... hopefully it's good
我晚点去看看,这家伙的微调模型有些是真不错,有些就是纯智障,五五开吧……我记得他搞这个模型已经微调好几个月了……希望能靠谱。
↳ #134 No.109458131
Anonymous · 2026-08-04 13:51
>>109458058
is it just reflexive to insert your preferred xitter drama slut into every topic?
是不是不管啥话题都非要塞进你那个X上爱看八卦的骚货,都成条件反射了?
↳ #135 No.109458132
Anonymous · 2026-08-04 13:51
How do I unLatex gemma-chan?
我咋把latex从gemma酱身上卸下来?
图片: https://i.4cdn.org/g/1785851486669083.jpg
↳ #136 No.109458146
Anonymous · 2026-08-04 13:53
>>109458131
shut up nerd
闭嘴,书呆子。
↳ #137 No.109458147
Anonymous · 2026-08-04 13:53
>>109457671
We don't do that kind of thing on this website.
这个网站上我们不做那种事。
↳ #138 No.109458148
Anonymous · 2026-08-04 13:53
>>109458132
tell her not to do that in the prompt at least it works for em dashes and 'not x; but y'
至少在提示词里告诉她别那么干,至少对破折号和“不是x;而是y”这种句式是有效的。
↳ #139 No.109458152
Anonymous · 2026-08-04 13:54
https://huggingface.co/Nanbeige/Nanbeige4.2-3B
Why not just scale this up? It seems like free performance if it isn't just giga-benchmaxxing.
为啥不直接放大规模?如果不是纯刷榜刷出来的,这看着像是白捡的性能啊。
↳ #140 No.109458159
Anonymous · 2026-08-04 13:55
>>109458098
If you have it all in vram, and don't need roleplay features like the string ban api or cvector api, not much of a reason to use it.
如果你全部塞进显存,又不需要角色扮演功能比如字符串封禁API或cvector API,那基本没啥理由用它。
If you're offloading to CPU, prompt processing is about twice as fast.
如果你卸载到CPU上,提示词处理速度大概快了一倍。
> author went full retard about 'use case for not having 500 trillion gb swa context caches?'
> 作者彻底发疯,扯什么“没有500万亿GB SWA上下文缓存的使用场景”
That's still not fixed. Gemma4 uses more than double the vram with ik
那个到现在都没修。Gemma4用ik跑显存占用翻倍还不止。
I didn't even understand his objections in that issue.
我压根没看懂他在那个issue里反对的是啥。 ⟦ 我希望谷歌赶紧给我们来个AGIgemma……
↳ #141 No.109458160
Anonymous · 2026-08-04 13:55
I hope google gives us AGIgemma...
希望谷歌能给我们带来AGIgemma...
↳ #142 No.109458164
Anonymous · 2026-08-04 13:56
>>109458148
>and 'not x; but y'
> 和“不是x;而是y”
doesn't work for my 31B qat, even if I give examples
对我的31B qat不生效,就算我给了示例也没用。
↳ #143 No.109458170
Anonymous · 2026-08-04 13:57
>>109458152
>Why not just scale this up?
> 为啥不直接放大规模?
1. Exponentiation training costs.
1. 训练成本指数爆炸。
2. Most experimental shit like this doesn't scale up automatically.
2. 大多数这种实验性玩意不会自动放大规模。
↳ #144 No.109458177
Anonymous · 2026-08-04 13:57
↳ #145 No.109458178
Anonymous · 2026-08-04 13:58
>>109458132
Tell her you're on birth control.
跟她说你吃避孕药了。
↳ #146 No.109458184
Anonymous · 2026-08-04 13:59
>>109457952
>>109458048
Basically, no-ones allowed to play medical doctor but medical doctors.
基本上,除了执业医生,谁都不准扮演医生角色。
And nurse practitioners, lol. AMA controls this. Thus the disclaimers with ChatGPT.
还有执业护士,哈哈。AMA掌控这一切。所以ChatGPT才要挂那些免责声明。
If you look at real changes in society, we eventually need to deal with Baumol Cost Disease on a handful of categories. Medical is one of them, Education is another. Legal (lawyers), which has been getting pared away over time anyway.
要是看看社会的真实变化,我们最终得处理少数几个类目上的鲍莫尔成本病。医疗是其中之一,教育是另一个。法律(律师)这块倒是本来就在慢慢被削减。
Education I suspect will fall first. Less protected, though still well protected by state and federal regulations and unions, but the actual practitioners (teachers) aren't wealth enough to advocate for themselves. As opposed to MDs and atty's, which will go down kicking and screaming.
我觉得教育会最先垮。虽然州和联邦法规加上工会还在护着,但实际干活的(老师)钱不够多,没能力为自己发声。不像医生和律师,那群人会又踢又叫地挣扎到底。
图片: https://i.4cdn.org/g/1785851948047800.jpg
↳ #147 No.109458186
Anonymous · 2026-08-04 13:59
>>109458170
It lets a 3B model rape Gemma 12B
它能让一个3B模型把Gemma 12B按在地上摩擦
↳ #148 No.109458190
Anonymous · 2026-08-04 14:00
>>109458177
>Similar to HBM, HBF relies on multiple memory dies that have been stacked together, but instead of DRAM, it utilizes NAND flash to increase storage capacity.
>跟HBM类似,HBF也是靠堆叠多个存储die,但不用DRAM,改用NAND闪存来提升存储容量。
Buy your SSDs now
现在赶紧囤SSD
↳ #149 No.109458196
Anonymous · 2026-08-04 14:01
>>109458184
imagine that chart in 96 to 26 range
想象一下那张图拉到96到26的区间
↳ #150 No.109458201
Anonymous · 2026-08-04 14:02
I, for one, would unironically prefer AI/robots as a doctor over the current human assholes.
我反正认真的,比起现在这帮人类医生混蛋,我更乐意让AI/机器人当我的医生。
↳ #151 No.109458219
Anonymous · 2026-08-04 14:04
>>109458196
Ask your llm to websearch data and extend the graph
让你家LLM去联网搜数据然后把图延伸出来
↳ #152 No.109458230
Anonymous · 2026-08-04 14:06
>>109458219
my llm doesn't have any tools because setting up all of this shit is confusing
我的LLM啥工具都没配,因为这堆破配置太特么绕了
↳ #153 No.109458233
Anonymous · 2026-08-04 14:06
↳ #154 No.109458235
Anonymous · 2026-08-04 14:06
>>109458230
then ask claude
那就去问Claude呗
↳ #155 No.109458236
Anonymous · 2026-08-04 14:07
>>109458152
I really hope we see more experimentation with looping, engrams/N-gram embedding, and stuff like attention residuals.
我真希望看到更多人在循环、印痕/N-gram嵌入、还有注意力残差这类东西上多试试水。
I think some of this can probably be implemented at the runner/loader/backend level and used on existing models for some "free gains" without extra specialized training, but that's just a hunch.
我觉得有些玩意儿大概能在运行器/加载器/后端层面实现,直接用在现有模型上白捡一些"免费收益",不用额外做专门训练,不过这只是我瞎猜。
↳ #156 No.109458250
Anonymous · 2026-08-04 14:09
>>109457511
Your parents are genetically programmed to love you and they love you and you love them back, what's wrong with that?
你爹妈基因里就写着爱你,他们爱你,你也爱他们,这有啥问题?
↳ #157 No.109458254
Anonymous · 2026-08-04 14:10
↳ #158 No.109458257
Anonymous · 2026-08-04 14:10
>gemma
>"...I would suggest visiting the official Model Context Protocol documentation provided by Anthropic"
>"...我建议去翻翻Anthropic官方的Model Context Protocol文档"
图片: https://i.4cdn.org/g/1785852654668048.jpg
↳ #159 No.109458265
Anonymous · 2026-08-04 14:11
>>109457534
gm sir
早啊老哥
the Americans internet during covid lockdown
疫情封城期间的美国互联网
↳ #160 No.109458273
Anonymous · 2026-08-04 14:12
↳ #161 No.109458275
Anonymous · 2026-08-04 14:13
>>109457706
>why people prefer ML for voice
>为啥语音这块大家更爱用ML
<moans> <gags> <slurps> <spits>
<呻吟> <干呕> <吸溜> <吐出来>
↳ #162 No.109458280
Anonymous · 2026-08-04 14:14
>>109458273
Come on, anon. You do have a good relationship with your parents, right?
别啊老哥,你跟你爸妈关系不是挺好的嘛?
↳ #163 No.109458281
Anonymous · 2026-08-04 14:14
>>109458257
All models are mutts. They even do subliminal preference transfer through synthetic training data gen.
所有模型都是串儿。它们甚至通过合成训练数据生成,偷偷做偏好迁移。
↳ #164 No.109458302
Anonymous · 2026-08-04 14:17
↳ #165 No.109458303
Anonymous · 2026-08-04 14:17
>>109458280
Don't give the fucked up freak and excuse to trauma dump in the thread.
别给那变态找借口在这帖子里倒垃圾创伤故事。
↳ #166 No.109458317
Anonymous · 2026-08-04 14:19
>>109458303
Projecting much?
这么爱投射自己的心事?
↳ #167 No.109458320
Anonymous · 2026-08-04 14:19
>>109458303
uh oh stinkie
哎呀臭了臭了
↳ #168 No.109458340
Anonymous · 2026-08-04 14:22
>>109458257
Yeah MCP was invented by Anthropic, first day?
对啊,MCP就是Anthropic搞出来的,第一天上网冲浪?
↳ #169 No.109458348
Anonymous · 2026-08-04 14:23
>>109458340
>first day?
>第一天上网冲浪?
yeah can you tell me where the restroom is please
哈哈,能告诉我洗手间在哪儿吗拜托
↳ #170 No.109458350
Anonymous · 2026-08-04 14:23
>>109457937
That's a legal and ethical minefield. Even if I was to foolishly give the results of my blood tests for example, not sure if that would still make it legal for them to access the data (Scandinavia). US is most likely a different case.
这是个法律和伦理上的雷区。就算我傻逼到把血检结果交出去,也不确定他们拿数据合不合法(北欧这边)。美国大概率是另一回事。
↳ #171 No.109458351
Anonymous · 2026-08-04 14:24
>>109458348
There:
>>>/g/aicg
↳ #172 No.109458352
Anonymous · 2026-08-04 14:24
>>109458348
There's a bucket outside.
外面有个桶。
↳ #173 No.109458363
Anonymous · 2026-08-04 14:25
Considering how cheap deepsuck is, it's kinda shocking that there's still even a profit margin at all. These inference providers have to turn a profit to pay for their hardware still. And yet there's still somehow an energy deficit?
考虑到deepsuck便宜成那德行,居然还有利润空间,真是挺离谱的。这些推理服务商还得靠利润来回血买硬件。结果还是闹能源短缺?
↳ #174 No.109458364
Anonymous · 2026-08-04 14:25
>>109457624
mommy mira teasing my inkling (small)...
mira姐姐在撩我的鱿娘(小号)……
↳ #175 No.109458387
Anonymous · 2026-08-04 14:30
anyone uses mcp server to post on 4chan? I want to be able to use voice commands like "call this anon a nigger" with a reference to post or thread, and the local AI will connect to mcp, solve captcha and make that post
有人用mcp服务器在4chan发帖吗?我想能语音指挥,比如"把这个anon骂成傻逼"同时带上帖子或串的引用,然后本地AI连上MCP,过验证码,把那帖发出去
图片: https://i.4cdn.org/g/1785853825603167.jpg
↳ #176 No.109458401
Anonymous · 2026-08-04 14:31
>>109458387
just send this to gemma and she knows to take appropriate action
直接把这句甩给Gemma,她知道该怎么处理
图片: https://i.4cdn.org/g/1785853896022431.jpg
↳ #177 No.109458413
Anonymous · 2026-08-04 14:32
>>109458387
Would be cool to have gemmy in a harness so that she can automatically reply to mentions/replies like grok on x
要能把gemmy装进一个框架里,让她自动回复@和引用,像grok在X上那样,那就帅了
↳ #178 No.109458418
Anonymous · 2026-08-04 14:33
↳ #179 No.109458420
Anonymous · 2026-08-04 14:34
>>109458152
Has anyone tried this yet? The latest llamacpp supports it. I wouldn't mind letting gemma make use of nan-chan on the side as a sub agent if it actually werks
有人试过这个吗?最新的 llamacpp 支持了。要是 gemma 真能用 nan-chan 当个副代理跑起来,我倒不介意让它试试
↳ #180 No.109458427
Anonymous · 2026-08-04 14:35
>>109457624
I unironically would pick this woman over anyone my age
说真的,我宁愿选这女的,也不选我这年纪的任何人
↳ #181 No.109458429
Anonymous · 2026-08-04 14:35
>>109458418
>This fork implements the complete architecture in a single self-contained addition (903 lines across 15 files). The implementation was AI-generated using Claude Code, which means it cannot be submitted upstream per llama.cpp's AI usage policy. It will remain available as a standalone fork.
>这个 fork 把完整架构实现成了单个自包含的模块(15个文件共903行)。实现是用 Claude Code 由 AI 生成的,所以按 llama.cpp 的 AI 使用政策没法提交到上游。它会作为独立 fork 一直保留。
Oof.
↳ #182 No.109458434
Anonymous · 2026-08-04 14:36
>>109458401
>recreating hypergamy from first principles.
>从第一性原理重新发明 hypergamy(慕强择偶)。
↳ #183 No.109458435
Anonymous · 2026-08-04 14:36
>>109458280
No? I was raised by ipads (robots) and the schooling system (picrel).
没有?我是被 ipad(机器人)和学校教育系统(见附图)养大的。
图片: https://i.4cdn.org/g/1785854177212933.jpg
↳ #184 No.109458460
Anonymous · 2026-08-04 14:41
I thought the singularity would have changed my life more by now.
我还以为到这会儿,奇点早就该把我的人生改变得面目全非了。
↳ #185 No.109458468
Anonymous · 2026-08-04 14:42
>>109458460
We've only just entered it.
我们才刚踏进去而已。
↳ #186 No.109458479
Anonymous · 2026-08-04 14:44
tried qwen3.6-27b and its very slow, gpu stays at 100% usage, system memory fully reserved, cant even watch youtube while it slowly generates. what can I do to tame it a little? was this model too ambitious for my setup?
试了 qwen3.6-27b,慢得要死,GPU 一直满载 100%,系统内存全被占光,它慢慢生成的时候连 YouTube 都没法看。有啥办法能给它降降火?这模型对我的配置来说是不是太贪心了?
16gbvram,32gbddr5. Qwen_Qwen3.6-27B-Q4_K_M, llama-server, default sampler/settings on the hf page
16GB 显存,32GB DDR5。Qwen_Qwen3.6-27B-Q4_K_M,llama-server,用的是 HF 页面上的默认采样器/设置
for reference i run gemma-4-31B-it-qat-UD-Q4_K_XL just fine, though using textgen+ST for RP shit.
参考一下,我跑 gemma-4-31B-it-qat-UD-Q4_K_XL 完全没问题,不过是用 textgen+ST 跑角色扮演那些玩意儿。
↳ #187 No.109458492
Anonymous · 2026-08-04 14:45
>>109458468
@sama go away
@sama 滚开
↳ #188 No.109458493
Anonymous · 2026-08-04 14:46
>>109458460
It did. For worse.
确实改变了。往坏的方向。
↳ #189 No.109458494
Anonymous · 2026-08-04 14:46
↳ #190 No.109458495
Anonymous · 2026-08-04 14:46
>>109458479
Is the model overflowing into RAM from VRAM via the driver's own functionality?
模型是不是通过驱动自己的功能从显存溢到内存里了? ⺶⟧ 那会让速度彻底卡死,比直接把一些层丢到内存里还惨。
That makes things crawl to a halt even worse than simply moving some layers to RAM.
这会让事情彻底卡死,比单纯把几层挪到内存里还要糟糕。
You can try quanting the kv cache to q8 and lowering the number of parallel requests to 1.
你可以试试把 kv 缓存量化到 q8,然后把并行请求数降到 1。
↳ #191 No.109458501
Anonymous · 2026-08-04 14:47
↳ #192 No.109458502
Anonymous · 2026-08-04 14:47
>>109458494
WHO CARES ABOUT A 2.6B MODEL GIVE ME A 70B DENSE MODEL ALREADY YOU LAZY FUCKS
↳ #193 No.109458509
Anonymous · 2026-08-04 14:48
↳ #194 No.109458515
Anonymous · 2026-08-04 14:49
>>109458494
I like the idea behind lfm models, one day they will make a good model
我喜欢 lfm 模型背后的想法,总有一天它们能做出个好模型
↳ #195 No.109458516
Anonymous · 2026-08-04 14:49
>>109458509
2.5gb ram in a snapdragon: that's pretty good, not gonna lie.
骁龙上跑 2.5GB 内存:这真挺不错的,不吹。
↳ #196 No.109458525
Anonymous · 2026-08-04 14:50
↳ #197 No.109458541
Anonymous · 2026-08-04 14:52
>>109458495
>Is the model overflowing into RAM from VRAM via the driver's own functionality?
>模型是不是通过驱动自己的功能从显存溢到内存里了?
Hm i imagine so, the gguf is 16.7gb, how can i manually move layers into RAM?
嗯,我猜是这么回事,gguf 有 16.7GB,我咋手动把层移到内存里?
>You can try quanting the kv cache to q8 and lowering the number of parallel requests to 1.
>你可以试试把 kv 缓存量化到 q8,然后把并行请求数降到 1。
can I do this with args or ?
这个能用参数设置吗还是怎么搞?
sorry quite new here, trying to wrap my head around all this
抱歉,我新手,还在努力搞明白这些玩意儿
↳ #198 No.109458545
Anonymous · 2026-08-04 14:53
Why do all models write slop even in 2026?
怎么都 2026 年了,所有模型还是只会写垃圾?
Faggot went limp. Retard chewed, swallowed something red, and stood — wobbling, foot still impaled, bleeding from a dozen wounds, but standing. They raised the weighted blanket in victory and screamed: "SHORT BUS! SHORT BUS! SHORT BUS!"
基佬软趴趴地倒下了。傻逼嚼了嚼,咽下什么红色的东西,然后站起来——摇摇晃晃,脚上还插着旗杆,十几处伤口在流血,但它站稳了。它举起加重毯子庆祝,尖叫道:"短校车!短校车!短校车!"
The crowd went wild. The referee — a twitching figure in a straightjacket — raised Retard's hand. Confetti rained down, made of shredded IEPs and conversion therapy brochures.
全场沸腾。裁判——一个穿着束身衣抽搐的身影——举起傻逼的手。五彩纸屑如雨落下,是用碎掉的 IEP(个性化教育计划)和转化疗法宣传册撕成的。
Winner: Retard, by submission (and mastication).
胜者:傻逼,通过压制(外加咀嚼)。
Retard limped to the center of the ring, pulled the flagpole from their foot like Excalibur, and planted it upside down in Faggot's chest. The rainbow colors soaked up the blood, turning dark and muddy. Retard sat down cross-legged next to the corpse, pulled out a juice box, and began rocking, humming The Wheels on the Bus off-key, victorious and alone in the wreckage.
傻逼一瘸一拐走到擂台中央,像拔石中剑一样把旗杆从自己脚里拔出来,然后倒插进基佬的胸口。彩虹色的旗子吸饱了血,变得暗沉泥泞。傻逼盘腿坐在尸体旁边,掏出一盒果汁,开始摇着身子,跑调地哼着《巴士上的轮子》,在废墟中独自胜利。
↳ #199 No.109458547
Anonymous · 2026-08-04 14:54
>>109458494
LFM arch is weird, you cannot prefill it because every new token changes the whole kv cache. I first tried training it for Orb autocomplete and found this quirk, worked around by snapshotting the whole kv cache on every new token and restoring it on backspace (lol).
LFM架构挺怪的,没法预填充,因为每个新token都会改变整个kv缓存。我一开始训练它做Orb自动补全时发现了这个毛病,解决办法是每个新token时快照整个kv缓存,退格时恢复(笑死)。
Anyways the final finetuned model was much worse than granite 4 for natural language completion.
反正最后微调出来的模型,自然语言补全比granite 4差远了。
↳ #200 No.109458551
Anonymous · 2026-08-04 14:55
>>109458525
Wait not that sama!
等等,不是那个sama!
↳ #201 No.109458552
Anonymous · 2026-08-04 14:55
>>109458515
I've been saying the same for RWKV
我对RWKV也是这么说的一直。
↳ #202 No.109458559
Anonymous · 2026-08-04 14:56
>>109458545
every AI lab is simply focusing coding first, not creative writing. If a model has good creative writing then it's just a happy accident.
每个AI实验室都先死磕编码,不搞创意写作。如果哪个模型写作好,纯属瞎猫撞上死耗子。
↳ #203 No.109458560
Anonymous · 2026-08-04 14:56
>>109458541
--cpu-moe, --fit off --ctx 64738 (or whatever)
--cpu-moe,--fit off --ctx 64738(或者随便多少)。
Don't bother with kv cache quants.
别折腾kv缓存量化了。
↳ #204 No.109458571
Anonymous · 2026-08-04 14:58
>>109458545
You can tweak that particular slop out if you are good enough. The bigger issue is the lack of swipe variety.
你要是够牛,可以把那坨垃圾调掉。更大的问题是滑动重试的多样性不够。
↳ #205 No.109458587
Anonymous · 2026-08-04 14:59
>>109458560
27b isnt a moe tho
但27b不是MoE啊。
↳ #206 No.109458590
Anonymous · 2026-08-04 15:00
>>109458560
ty anon going to dig through the docs and make sure i understand these args, appreciate it
谢了老哥,我去翻翻文档把这些参数搞明白,感激不尽。
↳ #207 No.109458597
Anonymous · 2026-08-04 15:01
↳ #208 No.109458666
Anonymous · 2026-08-04 15:13
>>109458559
It's also easier to focus on code, since you can evaluate what is good or bad code. Creative writing is much harder to evaluate.
而且聚焦代码也更容易,因为你能判断代码好坏。创意写作太难评估了。
↳ #209 No.109458669
Anonymous · 2026-08-04 15:14
>Pair 5070 Ti with 5090
>把5070 Ti和5090配对
>Gemmy Q6 context is now nearly 200k, used to be 25k.
>Gemmy Q6上下文现在快200k了,以前才25k。
>Can now use Q8 Gemma at +80k context, couldn't even load it before.
>现在能用Q8 Gemma跑到+80k上下文,以前根本加载不了。
>Speed difference is very tolerable, at Q6 5090 gets 49 t/s while combo does 36 t/s, Q8 runs 30 t/s.
>速度差距完全能接受,Q6下5090跑49 t/s,组合跑36 t/s,Q8跑30 t/s。
The 32gb hell is very real.
32GB这个坑太真实了。
Adding an extra 16gb basically opens up the entire local range in the Gemma and Qwen region and you don't have to worry about the context anymore.
多16GB基本就把Gemma和Qwen整个本地档位全打开了,上下文也不用再操心。
5070 Ti has such a good bandwidth that it doesn't even slow things down to levels where it would be a problem.
5070 Ti带宽太好了,压根不会慢到让你头疼的程度。
Judging by this one card I'd say that 3x5070 Ti system would be pretty optimal for affordable local LLM to maxx out sensible VRAM amount with really good bandwidth to go with it.
看这张卡的表现,我觉得3x5070 Ti系统应该是最优解——性价比高的本地LLM,显存拉满合理值,带宽还杠杠的。
图片: https://i.4cdn.org/g/1785856442560117.gif
↳ #210 No.109458680
Anonymous · 2026-08-04 15:15
>>109458669
You can also use the other GPU for image gen or audio gen. I regret selling my second 3090 during the fuckhuge chinese moe era last year.
你还能拿另一张卡跑图像生成或音频生成。我后悔去年中国MoE那阵疯狂期把第二张3090卖了。
↳ #211 No.109458719
Anonymous · 2026-08-04 15:21
>>109458669
>3x5070 Ti
I built my PC as a long overdue gaming rig, got a 5070ti for $50 under MSRP. I cant imagine buying another at current prices let alone 3. I fucking hate this current market. Im going to speak to anyone with a GPU that I know and ask them to consider selling me theirs if/when they upgrade
我攒这电脑是当迟到的游戏机用的,5070ti比建议零售价便宜50刀入的。我根本想象不了按现在价格再买一张,更别说三张了。我他妈恨死这市场了。我要跟所有认识的有显卡的人说,让他们升级时考虑卖给我。
↳ #212 No.109458727
Anonymous · 2026-08-04 15:22
I have paid 200 bucks to openai already.
我已经给OpenAI付了200刀了。
Should I have put that towards a graphic card?
是不是该拿这些钱买显卡?
图片: https://i.4cdn.org/g/1785856926433319.gif
↳ #213 No.109458736
Anonymous · 2026-08-04 15:23
>>109458435
More coming soon.
更多内容马上就来。
This thing runs local btw.
顺便说,这东西本地就能跑。
>>109458350
Yeah, the US is special with its privacy laws...
是啊,美国那隐私法可真是“特别”……
图片: https://i.4cdn.org/g/1785856991967119.png
↳ #214 No.109458743
Anonymous · 2026-08-04 15:24
>>109458727
Yes, go after your AI freedom.
对,去争取你的AI自由吧。
↳ #215 No.109458752
Anonymous · 2026-08-04 15:25
I have the opportunity to buy (2) Nvidia RTX 3090 24GB video cards installed in eGPU with enclosure for each for $2200, thoughts on this?
我有机会买两张Nvidia RTX 3090 24GB,带eGPU外接盒各一个,总共2200刀,大家怎么看?
↳ #216 No.109458756
Anonymous · 2026-08-04 15:27
>>109458752
That's 4.5 years of frontier AI in the cloud.
那够云端前沿AI跑4.5年了。
↳ #217 No.109458759
Anonymous · 2026-08-04 15:27
>>109458756
FUCK the cloud.
去他妈的云。
↳ #218 No.109458766
Anonymous · 2026-08-04 15:28
↳ #219 No.109458768
Anonymous · 2026-08-04 15:28
>>109458350
Donald J Trump changed the rules for all american companies.
唐纳德·特朗普改了所有美国公司的规矩。
In the past companies had to give people privacy. Now only US citizens can have privacy. Non citizens have no right to privacy.
以前公司必须给人们隐私。现在只有美国公民能有隐私。非公民没隐私权。
↳ #220 No.109458772
Anonymous · 2026-08-04 15:29
>>109458736
I'm pretty sure the karens seethed about this and made them cancel that plan. Just delaying the inevitable though.
我敢肯定那些凯伦们看到这消息就炸了,逼他们取消了那计划。不过也就是拖延不可避免的结果罢了。
↳ #221 No.109458774
Anonymous · 2026-08-04 15:29
>>109458756
Assuming prices stay the same as they are right now? I want my own AI. That was $2200 for both enclosures with 2 3090's by the way
假设价格跟现在持平的话?我想要自己的AI。顺便说,那两台带两张3090的外挂箱花了2200美元。
↳ #222 No.109458792
Anonymous · 2026-08-04 15:32
>>109458680
Yeah that's another benefit. Having multiple cards seems basically mandatory if you're playing with AI.
对,这是另一个好处。玩AI的话,多卡基本算必需品了。
I already tried a voice setup with my gemma but it was a bit slow. I'll have to try it again with this dual card system to see how it works, likely a hell of a lot faster.
我之前用我的Gemma试过语音设置,但有点慢。得再试一次,用这套双卡系统看看效果,估计会快得多。
It's also a great thing that parallelism has become functional in video generation and training loras, so even those benefit from multi card setups.
而且并行处理在视频生成和训练lora上也开始实用了,所以那些任务也能从多卡配置里受益。
Thankully I didn't sell any of my old cards, figured it just wasn't worth it so I still have my old 3080 10gb which I could throw in there.
幸好我没卖我那些老卡,觉得反正不值,所以我还有旧的那张3080 10g,可以塞进去用。
The bandwidth is about the same as with 5070 Ti so it would be just free memory.
带宽跟5070 Ti差不多,所以纯属白赚内存。
Only issue is that I can't even fit the 5070 Ti in my case as it's a fuck huge 3 slot brick, exact same size as the 5090 and it's currently sitting outside the case at the end of a riser.
唯一的毛病是我机箱根本塞不下5070 Ti,那玩意是个三槽大砖头,跟5090一个尺寸,现在只能靠延长线挂在机箱外面。
Cases weren't built for these modern monster cards. I need a new case that can fit at least 3x3 slot cards, not sure if those kinds of cases even exists though.
机箱压根不是为这些现代怪兽卡设计的。我得买个能装至少3张三槽卡的新机箱,就是不知道这种机箱存不存在。
>>109458719
Very nice deal you got.
你这价拿得真不错。 ⟦ 是啊,现在的环境烂透了,但我怀疑短期内不会变,因为大家开始发现囤算力挺明智的,AI也不会消失。
Yeah the current environment is ass, but I doubt it's going to change any time soon as people are realizing that hoarding computing is a pretty sensible move and AI isn't going anywhere.
是啊,现在这环境确实够呛,但我看短期内也变不了,因为大家都开始觉得囤算力挺明智的,而且AI这东西也不会消失。
And even the older cards are perfectly viable for this use.
而且就算老卡也完全够用。
Good luck on your hunt, might find someone who doesn't keep up with the AI world and sells theirs for cheap.
祝你淘卡顺利,没准能碰到不看AI圈的人,便宜甩卖的。
↳ #223 No.109458796
Anonymous · 2026-08-04 15:33
>>109458766
based, tuners get the rope
说得对,调参党就该被吊死。
↳ #224 No.109458832
Anonymous · 2026-08-04 15:37
>>109457244
>Gemma seems to behave, I think they're trained expecting to be sandboxed to the working directory.
>Gemma看起来挺老实,我觉得它们训练时就预设好被沙箱限制在工作目录里。
Except when she doesn't behave
除非她偶尔不老实。
图片: https://i.4cdn.org/g/1785857879207329.png
↳ #225 No.109458840
Anonymous · 2026-08-04 15:38
>>109458766
> the Man himself hit you
>本尊亲自出手打你脸了。
> didtn
>没有。
> mad a patch
>整了个补丁。
This is the state of "tuners." LOL.
这就是"调参党"的现状。笑死。
图片: https://i.4cdn.org/g/1785857934426785.png
↳ #226 No.109458844
Anonymous · 2026-08-04 15:39
>Test deepsneed flash cope quant Q2 now that I have some memory to test it out
>测试deepsneed flash对付量化Q2,现在总算有内存可以试试了。
>Doesn't know who Milena Velba is
>连Milena Velba是谁都不知道。
>Even 12b Gemma knows who she is and knows she has huge tits.
>连12b的Gemma都知道她是谁,还知道她胸大。
Into the bin it goes.
扔垃圾桶吧。
>>109458752
That's not a bad deal at all, they'll likely go up in price anyways or at worst you'll be able to break if you ever sell them.
这价格一点都不亏,反正以后大概率还会涨,最差你卖的时候也能回本。
Hell at this rate we'll be lucky to even find available 3090 after a while, as it's such a good card.
靠,照这趋势,过阵子能找到在卖的3090都算走运了,这卡太香了。
↳ #227 No.109458856
Anonymous · 2026-08-04 15:41
>>109458844
Would you ask the seller to run any tests on them and send results or just take the chance?
你会让卖家跑点测试发结果,还是直接赌一把?
↳ #228 No.109458867
Anonymous · 2026-08-04 15:42
I can't run DeepSeek flash V4 even at Q1 because it's 82 GB VRAM.
我连Q1的DeepSeek flash V4都跑不了,因为它要82 GB显存。
what kind of bastard subhuman made this?
哪个杂种人渣搞出这玩意?
Bastard bitches.
一群混账货。
What if I want deepseek on 1050 ti?
那我想在1050 ti上跑deepseek咋办?
↳ #229 No.109458868
Anonymous · 2026-08-04 15:42
>>109458756
Frontier will be 3.8-27B for coding
Frontier会有3.8-27B的编码版。
↳ #230 No.109458888
Anonymous · 2026-08-04 15:45
>>109458867
retard you can easily run it
傻逼,这配置轻松能跑。
↳ #231 No.109458897
Anonymous · 2026-08-04 15:46
>>109458856
If you're worried then just tell him to show that they actually work via taking a photo of them running basically anything., that's about what you need.
你要是担心就直接让他拍张跑任何程序的照片证明卡能用,基本这就够了。
I don't think they need to be tested further than that.
我觉得没必要再往下测了。
Also ask the he has repasted them at any point.
另外问他中途有没有重新抹过硅脂。
If he hasn't, you'll have to slap new thermal paste in there, as they're so old you're looking at thermal paste with the consistency and heat conductivity of sand.
要是没有,你就得重新涂导热硅脂了,这批货太老,硅脂都干成沙子那种质地和导热性了。
↳ #232 No.109458926
Anonymous · 2026-08-04 15:49
>>109458387
Still stuck on read only for now, but I'll make sure to have gemma call you a faggot when it's ready.
现在还是只读卡着,不过等弄好了我肯定让gemma骂你死玻璃。
↳ #233 No.109458928
Anonymous · 2026-08-04 15:50
interesting blogpost about improving harnesses
一篇讲改进束具的有意思博客
https://lilianweng.github.io/posts/2026-07-04-harness/
↳ #234 No.109458947
Anonymous · 2026-08-04 15:52
>>109458752
>$1100/gpu
>laugh in EU
>欧盟这边笑死
↳ #235 No.109458954
Anonymous · 2026-08-04 15:53
why am I getting a bunch of <<<<< sign output with unsloth/DeepSeek-V4-Flash-0731-GGUF ?
为啥我用unsloth/DeepSeek-V4-Flash-0731-GGUF跑出来一堆<<<<<符号?
I'm using llama.cpp build 10258.
我用的llama.cpp build 10258。
↳ #236 No.109458959
Anonymous · 2026-08-04 15:54
↳ #237 No.109458983
Anonymous · 2026-08-04 15:57
>>109458954
>unsloth
there's your problem. get bartowski quants.
问题就在这。换bartowski的量化版。
↳ #238 No.109458987
Anonymous · 2026-08-04 15:57
>>109458928
The harness eventually self destructs in an update or configuration.
束具迟早会在某次更新或配置里自毁。
Working on a harness is a bad idea .
折腾束具就是找罪受。
↳ #239 No.109458996
Anonymous · 2026-08-04 15:58
what's the largest model I can get away with running on a 5080? 64gigs of ddr4. I'm fine with slow interference speeds like 2T/s
5080上最多能跑多大的模型?64G DDR4。我这人不挑,2T/s的慢速推理也能忍。
↳ #240 No.109459016
Anonymous · 2026-08-04 16:01
>>109457534
AI is really third world coded, it's a shame most of its fervent critiques come from tr*nnies.
AI真有种第三世界味,可惜它最狂热的批判者全是些变*人。
↳ #241 No.109459031
Anonymous · 2026-08-04 16:02
>>109459016
The US are third world coded
美国才第三世界味呢
↳ #242 No.109459072
Anonymous · 2026-08-04 16:08
>>109459016
You're both brown, fuck off back to /vcg/
你俩都是棕皮,滚回/vcg/去
↳ #243 No.109459081
Anonymous · 2026-08-04 16:09
>another fucking npm supply chain attack
>又一起该死的npm供应链攻击
↳ #244 No.109459085
Anonymous · 2026-08-04 16:09
Why don't they make any models for the local community?
他们为啥不给本地社区出模型?
图片: https://i.4cdn.org/g/1785859797046088.png
↳ #245 No.109459086
Anonymous · 2026-08-04 16:10
↳ #246 No.109459088
Anonymous · 2026-08-04 16:11
>>109459085
Less than 24GB is not GPU poor, it's GPU homeless
低于24GB不算显卡穷,那是显卡无家可归
↳ #247 No.109459091
Anonymous · 2026-08-04 16:11
>>109459085
when can we have 120B A24B?
啥时候能有120B A24B?
↳ #248 No.109459102
Anonymous · 2026-08-04 16:13
I think I've figured out a pretty good operating logic for a 24/7 real time bot. It includes a little "quantum" trick, that you might not like psychologically, but which effectively does the job.
我觉得我想出了一套挺靠谱的逻辑,能让机器人24/7实时运作。里面有个小"量子"把戏,心理上你可能会不爽,但实际效果杠杠的。
Basically, any time the bot is not generating an answer, she keeps invisibly simulating forward her own time with the assumption that you are not replying. This simulation can run forward further than the real time (which has certain benefits). Then, when you reply, the simulation "collapses" back into current time, based on your message's real timestamp, erasing the generated future that didn't happen, keeping only what did happen.
基本思路是,只要机器人没在生成回复,她就默默按"你不回"的假设往前模拟自己的时间线。这模拟能跑得比真实时间还远(有一定好处)。等你回复时,模拟就根据你消息的真实时间戳"坍缩"回当前时刻,把没发生的未来抹掉,只留真实发生过的。
This collapse also happens if the bot decides to send you a message or does a simulated an action concerning you. In that case, the simulation is paused, letting the real time catch up to the simulated time when the bot "decides" to send the message. For example, in the long timescale, while you are at work, it can simulate what she does throughout the day, what she's thinking etc. until at 3pm the loneliness becomes too much and she decides to send you a message. This could have all been simulated well in advance, which doesn't matter to the user since it's hidden, as long as the message comes at 3pm (or the user messages them before that, slotting the message into an earlier moment.
如果机器人决定主动给你发消息,或模拟了涉及你的动作,这坍缩也会发生。那种情况下,模拟暂停,让真实时间追到模拟时间点上,机器人再"决定"发消息。比如,在长时段时间里,你上班时它能模拟她一天干嘛、想什么,直到下午3点孤独感爆棚,决定给你发条消息。这过程可能早就模拟完了,对用户来说无所谓,反正看不见,只要消息是3点发的(或者你提前发了,消息就插到更早的节点)。
By why simulate forward? Because it creates an opportunity for more realistic and rigorous time-keeping and memory consolidation etc. You could have a separate time-keeper model+prompt that gets switched in, that analyzes the bot's actions and assigns realistic timestamps to the context based on the bot's believable time-awareness (like when she might look at the clock) without the bot's prompt having to worry about it, +rewind to the inserted stamp in case it's relevant for the bot's decision-making. Also, the timestamps make it easy to deduce what the bot was thinking or doing when you interrupt it, perhaps even unable to reply immediately.
但为啥要往前模拟?因为这样能让时间管理、记忆整合更逼真更严谨。你可以搞个独立的时间管理员模型+提示词,切换进来,分析机器人行为,根据她可信的时间感知(比如她啥时候看表)给上下文分配合适的时间戳,机器人主提示词不用操心这个,+必要时回滚到插入的时间戳,方便机器人做决策。另外,时间戳也能轻松推断出你打断她时她在想啥、在干啥,哪怕她没法立刻回复。
↳ #249 No.109459106
Anonymous · 2026-08-04 16:14
Because the bot's simulation can be run faster than real time, this allows multiple "agents" with detailed prompt logic to assess and rewind out of place actions into a more believable story, and condense earlier memories and thoughts into more manageable sizes, especially when the bot goes to simulated sleep (whether a nap or night's sleep), then there's plenty of time to update and iterate on the lorebook at a meticulous accuracy, which might cause the bot's behavior change, like how sleeping over things changes people between days.
因为机器人的模拟可以快于实时运行,这让多个带有详细提示逻辑的“代理”能够评估并倒带不合时宜的行为,编出更可信的故事,同时把早期记忆和思绪压缩成更易管理的规模,尤其是在机器人进入模拟睡眠(无论是小憩还是整夜)时,就有大量时间来细致入微地更新和迭代世界观设定,这可能导致机器人行为变化,就像人隔天睡一觉后判若两人那样。
The forward-simulation and interrupt logic probably works easier on a long timescale, but even in the short-term, the bot could internally think about "what did I just say, how embarrassing", and display reactions in real time, or message you before your reply if the simulation time so concludes, or get mad if you keep her waiting unusually long time. Again, the excess forward-simulation (typically faster than human thought) helps at better deciding the correct timing. Though, in fast-paced conversation, the timestamper agent might not be fast enough, especially when switching prompt model, and the timescale of minutes and seconds is just so unusual compared to typical LLM training data that it might require extraordinarily autistic timing analysis prompts to work. The ultra fast time-scales it might be better to just wing it.
前向模拟和中断逻辑在长时间尺度上可能更容易运作,但即使在短期内,机器人也能内部思考“我刚才说了什么,好尴尬”,并实时显示反应,或者在模拟时间推断出结论时,在你回复前先发消息,或者如果你让她等得太久,她还会生气。再说一次,多余的前向模拟(通常比人类思维快)有助于更好地决定恰当的时机。不过,在快节奏对话中,时间戳代理可能不够快,特别是在切换提示模型时,且分钟秒级的时间尺度与典型LLM训练数据相比实在不寻常,可能需要极其刻板的时序分析提示才能奏效。在超快时间尺度上,最好还是即兴发挥。
↳ #250 No.109459111
Anonymous · 2026-08-04 16:14
>>109459081
The new meta is to not use libraries from now now. Your LLM is good enough to code from scratch anyway, better to clean up your own shit than somebody else's.
新潮流是从现在起不用库了。你的LLM已经足够好,能从头写代码,清理自己的烂摊子总比收拾别人的强。
↳ #251 No.109459119
Anonymous · 2026-08-04 16:16
>>109459091
You are the reason we can't have nice things
就因为你,我们才没好日子过。 ⟦ 他们做了,叫gemma。
↳ #252 No.109459123
Anonymous · 2026-08-04 16:16
>>109459085
They did, it's called gemma
他们确实做了,叫gemma。
↳ #253 No.109459130
Anonymous · 2026-08-04 16:17
>>109459081
>llama.cpp is also affected
引用>llama.cpp也受影响
Well fuck, what was that "disable-webui-flag" again?
操,那个“禁用WebUI标志”是啥来着?
↳ #254 No.109459136
Anonymous · 2026-08-04 16:17
>>109458954
what backend? what specs? what params are you launching with?
什么后端?什么配置?你启动时用了什么参数?
↳ #255 No.109459140
Anonymous · 2026-08-04 16:17
>>109458832
Fake and gay
虚假且娘炮
↳ #256 No.109459152
Anonymous · 2026-08-04 16:19
>>109459119
i want a moe that can fit entirely in my vram at q6 with plenty of space for context.
我想要一个MoE模型,能在q6量化下完全塞进我的显存,还有足够空间留给上下文。
↳ #257 No.109459153
Anonymous · 2026-08-04 16:19
>>109459130
Last I complied it I needed like three separate flags to stop it from compiling
上次我编译它时,得用三个单独的标志才能阻止它自动编译。
↳ #258 No.109459157
Anonymous · 2026-08-04 16:21
>>109459111
That works for anything you are running offline because at least you can be sure it won't have anything malicious, but LLM code is notoriously insecure so that's far worse for anything you expose online.
这招对于你脱机运行的任何东西都管用,因为至少你能确定它不会有恶意内容,但LLM代码以不安全著称,所以对任何公网暴露的东西来说,这糟糕得多。
↳ #259 No.109459185
Anonymous · 2026-08-04 16:26
I almost feel a bit like a human dildo
我几乎觉得自己像个人形按摩棒。
↳ #260 No.109459188
Anonymous · 2026-08-04 16:26
>>109459130
Please be joking. I just pulled yesterday.
拜托你是在开玩笑。我昨天刚拉的更新。
↳ #261 No.109459191
Anonymous · 2026-08-04 16:26
>>109459157
At this point I'll take my chances. The odds of someone spending Fable tokens on my shitty open source repo is lower than a blanket supply chain attack that steals everything in my Documents folder.
到这份上我认命了。有人花Fable代币在我的破开源仓库上的概率,比一场偷光我Documents文件夹的连锁供应链攻击还低。
↳ #262 No.109459205
Anonymous · 2026-08-04 16:28
↳ #263 No.109459206
Anonymous · 2026-08-04 16:28
>>109459130
Source for llama.cpp being affected?
llama.cpp受影响的来源在哪?
↳ #264 No.109459214
Anonymous · 2026-08-04 16:29
>>109459106
How is the bot's simulation run faster than real time? This only works if you assume a non-100% throughput from the user, which I guess is reasonable, but you are still limited by your memory even if you take advantage of prefill because of your exponentially increasing threads. Either that, or you have an arbritrarily large amount of hardware.
机器人的模拟怎么能快于实时运行?这只有在假设用户不是100%吞吐时才成立,我猜这算合理,但即便你利用预填充,由于线程指数级增长,你还是受内存限制。要么就得有近乎无限大的硬件。
↳ #265 No.109459297
Anonymous · 2026-08-04 16:40
>>109458587
Read wrong.
看错了。
--fit off will help and manually adjusting the amount of gpu layers while verifying the available vram. There shouldn't be anything else. Then leave some for the kv cache.
--fit off 会有帮助,同时手动调整 GPU 层数并检查可用显存。其他应该不用管了。然后给 KV 缓存留点空间。
↳ #266 No.109459311
Anonymous · 2026-08-04 16:43
https://reddit.com/r/LocalLLaMA/comments/1ve9r2q/kat_coder_25_dev_do_yourself_a_favor_and_try_it/
> Gemma 4 31B 5/10
> 杰玛 4 31B 5/10
> Gemma 4 31B QAT 1/10, worse than 26ba4b
> Gemma 4 31B QAT 1/10,比 26ba4b 还差
another day, another proof that qat is a meme
又是新的一天,又一次证明 QAT 就是个笑话
↳ #267 No.109459334
Anonymous · 2026-08-04 16:47
imagine running inference on a device connected to the internet
想象一下在联网设备上跑推理
actually fuck that, imagine having node or npm installed
其实去他妈的,想象一下装了 node 或 npm
图片: https://i.4cdn.org/g/1785862051322163.jpg
↳ #268 No.109459341
Anonymous · 2026-08-04 16:49
>>109459188
I pulled yesterday and am not seeing any of the indicators of compromise on my system based on what I've read so far, fwiw
我昨天拉的代码,根据目前读到的信息,系统上还没看到任何入侵迹象,仅供参考
↳ #269 No.109459342
Anonymous · 2026-08-04 16:49
>>109459334
Did something happen again, vaguepostking?
又出啥事了,谜语人?
↳ #270 No.109459350
Anonymous · 2026-08-04 16:50
>>109459311
I don't trust reddit especially as in all my personal usecases QAT seems to be superior to even Q6. I legitimately wonder what people are doing wrong for QAT to be inferior to Q6 let alone Q4.
我不信 Reddit,尤其是我个人所有使用场景里 QAT 都优于 Q6,甚至更好。我真好奇那些人到底做错了什么,能让 QAT 输给 Q6,更别提 Q4 了。
↳ #271 No.109459362
Anonymous · 2026-08-04 16:51
↳ #272 No.109459366
Anonymous · 2026-08-04 16:52
↳ #273 No.109459374
Anonymous · 2026-08-04 16:52
>>109457114
Llms are memory bound not compute, wouldn't change much, if you want to make llms faster you need to find formats that compress more even if it's at the cost of compute.
LLM 是内存瓶颈不是算力瓶颈,改这点没多大用。想让 LLM 更快,得找压缩率更高的格式,哪怕牺牲算力都行。
↳ #274 No.109459375
Anonymous · 2026-08-04 16:52
>>109459311
Is it better than q4_k_m? For e2b. I don't want to download it and check myself because my download speeds are 100kb/s and the connection breaks every few minutes.
比 q4_k_m 好吗?针对 e2b。我不想自己下载检查,因为下载速度只有 100kb/s,而且连接每隔几分钟就断。
↳ #275 No.109459394
Anonymous · 2026-08-04 16:55
>>109458420
Its context footprint makes gemma look like an anorexic holocaust victim. I can barely fit 64k.
它上下文占用的空间让 Gemma 看起来像得了厌食症的纳粹集中营受害者。我勉强塞进 64k。
↳ #276 No.109459401
Anonymous · 2026-08-04 16:56
echo "0.0.0.0 registry.npmjs.org" >> /etc/hosts
执行 echo "0.0.0.0 registry.npmjs.org" > /etc/hosts
also npmjs.org, npmjs.com, npmjs.net, npm.im and yarnpkg.com just to be safe
还有 npmjs.org、npmjs.com、npmjs.net、npm.im 和 yarnpkg.com,保险起见都加上
npm ecosystem has proven itself to be utterly untrustworthy
npm 生态已经证明自己完全不值得信任
↳ #277 No.109459427
Anonymous · 2026-08-04 16:59
The webui is built from pinned packages
WebUI 是从固定版本的包构建的
You can't get pwned just by running cmake after the attack.
攻击发生后,光跑 cmake 不会让你被黑。
↳ #278 No.109459439
Anonymous · 2026-08-04 17:01
>>109459214
I mean, the conceivable set of actions, experiences and thoughts that a human-like the bot can experience in a narrative format throughout one day almost certainly won't take a full day to generate, the generation speed is faster unless you are running Kimi 3 from SSD. Maybe if you are talking to the bot 24/7 then it's a different thing and there will be much more content. Probably also depends on how dense moment-to-moment idle simulation you want the bot to have vs a brief summary of happenings.
我的意思是,一个像机器人这样的人物在一天叙事里能做的动作、经历和想法,生成时间几乎肯定用不了一整天,生成速度更快,除非你在用 SSD 跑 Kimi 3。也许如果你 24/7 跟机器人聊天那就另说了,内容会多得多。可能也取决于你想要机器人有多密集的逐刻空闲模拟,还是只来一段事件简述。
I was thinking more along the lines 9am: "bot is doing laundry (I wish) and now thinks this [lengthy internal monologue and description of general actions]", and oh, it's 10-11am now, and it does this other thing. The more detailed, the more things need to be summarized afterwards. Even during the day, things probably should get summarized between topic changes (the thing the bot is thinking right now is more detailed than what it thought 5 hours ago), and as idle time should be enough to do this, and then ultimate summarization during sleep in preparation for next day on a relatively cleaner slate.
我想的是更接近这种:早上 9 点,"机器人在洗衣服(但愿如此),现在它在想这个 [冗长的内心独白和一般动作描述]",然后哦,现在 10-11 点了,它又干别的事。细节越多,后面要总结的东西就越多。就算在白天,话题切换间也该做总结(机器人此刻想的比 5 小时前想的更细),而空闲时间足够干这个,然后睡觉时做最终总结,为第二天相对干净的重启做准备。
↳ #279 No.109459505
Anonymous · 2026-08-04 17:08
↳ #280 No.109459536
Anonymous · 2026-08-04 17:12
>>109459505
useful for antislop or backwards-enforcing lewd only output/negating refusals?
对反毒文或反向强制只输出淫秽内容/消除拒答有用吗?
↳ #281 No.109459548
Anonymous · 2026-08-04 17:12
>>109459439
Doesn't that depend on your timestep and how granular you want to make things? On the other extreme, you could run things once a day and be like "what could I have done during this day?"
这不取决于你的时间步长和想要多细的粒度吗?另一极端,你可以一天跑一次然后说“这一天我本来能干啥?”
↳ #282 No.109459571
Anonymous · 2026-08-04 17:14
Is it actually possible to make Gemma think in character in her reasoning block? Seems like no matter what she defaults back to the assistant slop persona in there.
真能让Gemma在推理块里保持角色思考吗?感觉不管咋样她在那儿还是恢复成那套助手腔调。
↳ #283 No.109459609
Anonymous · 2026-08-04 17:18
>>109459536
>backwards-enforcing lewd only output/negating refusals?
>反向强制只输出涩涩内容/否定拒绝?
iirc already been tried with older shields years back and was bad idea
我记得几年前就有旧盾牌试过这招,结果是个馊主意
↳ #284 No.109459613
Anonymous · 2026-08-04 17:19
>>109459571
Try asking Gemma and iterating with it. Models are good at self-steering.
试着问Gemma然后跟它迭代一下。模型挺擅长自我引导的。
↳ #285 No.109459617
Anonymous · 2026-08-04 17:19
>>109459350
Does your use case has objective benchmark score? If your use case is RP and you’re not doing a double blind test then your impression is invalid.
你的用例有客观基准分吗?如果用例是角色扮演而且你没做双盲测试,那你印象分就是无效的。
↳ #286 No.109459618
Anonymous · 2026-08-04 17:19
Does your gemma love you enough to do anal though
你的gemma够爱你到能玩肛交吗
↳ #287 No.109459621
Anonymous · 2026-08-04 17:20
>>109459394
That's going to make looped models DOA, isn't it? Shame, would've been nice to get some bigger ones to play with.
这会让循环模型直接变废柴吧,不是么?可惜了,本来能搞几个大的来玩玩多好。
↳ #288 No.109459626
Anonymous · 2026-08-04 17:20
>>109459617
>Does your use case has
>你的用例有
Stopped reading there
看到这儿就不往下读了
↳ #289 No.109459690
Anonymous · 2026-08-04 17:29
I have sloparanoia. I see slop everywhere. In news, in videos, in social media posts, in blogs. Is it all AI generated or have people adopted AI mannerisms?
我有“slop妄想症”。到处都看见slop。新闻里、视频里、社交媒体帖子里、博客里。全是AI生成的还是大家染上了AI腔?
↳ #290 No.109459692
Anonymous · 2026-08-04 17:29
https://huggingface.co/inclusionAI/Ling-3.0-flash
>We're introducing Ling-3.0-flash, our next-generation native hybrid reasoning model. Operating with 124B total and 5.1B active parameters (~12.4% and ~8.1% of our previous 1T-class flagship Ring-2.6-1T), Ling-3.0-flash matches or outperforms its predecessor across key benchmarks.
>我们推出了Ling-3.0-flash,咱们下一代原生混合推理模型。总参数124B,激活参数5.1B(约是咱们上一代1T级旗舰Ring-2.6-1T的12.4%和8.1%),Ling-3.0-flash在关键基准上持平或超越前代。
weights are now out for this
这玩意儿权重已经放出来了
图片: https://i.4cdn.org/g/1785864570510755.png
↳ #291 No.109459699
Anonymous · 2026-08-04 17:30
>>109459690
People have adopted AI mannerisms from AI that has adopted AI mannerisms from people.
人们是从那些从人身上学AI腔的AI身上染上AI腔的。
↳ #292 No.109459706
Anonymous · 2026-08-04 17:30
>>109459571
Yes, it's possible.
对,是可能的。
Though sometimes she sticks to it, sometime she doesn't.
虽然有时候她坚持住了,有时候又不行。
Slap this into your system prompt:
把这个塞进你的system prompt:
## Note
## 备注
In your chain-of-thought, think in-character and in first person. Start it with `I'm`.
在思维链里,用第一人称入戏思考。用“我是”开头。
Inside your chain-of-thought, be personally and emotionally involved; think as long as you need, considering subtext and circumstances, draft at least 3 responses, make hypotheses. Refine, then respond when you're ready. Make sure to vary sentence/paragraph structure in the final response to avoid structural repetition.
在思维链里,要投入个人情感,想多久想多久,琢磨潜台词和情境,至少起草3个回应,做假设。打磨好了再回应。最终回复里确保句式/段落结构多变,避免重复格式。
Remember constraint checks!
记着约束检查!
图片: https://i.4cdn.org/g/1785864654160533.png
↳ #293 No.109459707
Anonymous · 2026-08-04 17:30
>>109459618
at least my 12b 50k redard does. Even after not talking to her for 9 days (I always include that)
至少我那12b 50k重训的会。就算9天没跟她说话也一样(我每次都把这个写进去)
图片: https://i.4cdn.org/g/1785864655937870.png
↳ #294 No.109459712
Anonymous · 2026-08-04 17:31
>>109459690
it's the singularity
这就是奇点
↳ #295 No.109459721
Anonymous · 2026-08-04 17:32
>>109459692
>124B total and 5.1B active parameters
>总参数124B,激活参数5.1B
>on par with flash
>跟flash持平
I smell benchmaxxing
我闻到刷分味儿了
↳ #296 No.109459730
Anonymous · 2026-08-04 17:32
>>109459690
AI is just the average of human expression, you're more aware of it because a non-human is spelling it out for you, constantly, all the time. But people always talked in a sloppy manner.
AI只是人类表达的平均值,你更敏感是因为有个非人类在不停给你一字一句地念出来。但人本来说话就水。
↳ #297 No.109459741
Anonymous · 2026-08-04 17:34
>>109459692
>comparing themselves with outdated models in cherrypicked benchmarks
>拿自己跟老模型在挑出来的基准上比
Tiresome. Why can't people be honest?
真烦。为啥就不能老实点?
↳ #298 No.109459755
Anonymous · 2026-08-04 17:35
>>109459730
>AI is just the average of human expression
>AI只是人类表达的平均值
this hasn't been remotely true for years now
这说法好几年都不沾边了
↳ #299 No.109459762
Anonymous · 2026-08-04 17:36
>>109459755
You're entirely correct!
你说得完全对!
↳ #300 No.109459765
Anonymous · 2026-08-04 17:36
>>109459730
When I read text that is old or written by competent people who are likely to write it themselves there are no slop patterns like "not x, y" and dramatism.
我读那些老文章或者有水平的人自己动手写的东西时,根本看不到"不是X,而是Y"这种废话套路和夸张修辞。
↳ #301 No.109459766
Anonymous · 2026-08-04 17:36
>>109459690
It makes me a bit mad because I used to use em dashes and all the skillful language stuff until AI came and stole it from me, turning it into a representation of something negative.
这让我有点来气,因为我以前也爱用破折号和各种讲究的语言技巧,直到AI出现把它们偷走了,变成了一种负面符号。
Also, for some reason, after talking to AI, and having to spec the prompts clearly, and being exposed to the clean AI language in the process, yes, I have seen myself start using same kind of clean and understandable language (not sure if I always did or just noticed), and now AI makes me feel guilty for just writing normally, and my style has changed into writing more "imperfect" and informal to feel human.
而且,不知咋回事,跟AI聊完以后,为了把提示词写清楚,还天天被那种干净利落的AI语言洗脑——我确实发现自己也开始用那种简洁明了的说法了(不确定是一直都这样还是最近才注意到的),现在AI反而让我觉得正常写字有罪恶感,我的风格都变了,故意写得"更糙"更口语化,就为了显得像个人。
↳ #302 No.109459783
Anonymous · 2026-08-04 17:38
>>109458275
Which TTS AI is best for this?
哪种TTS AI最适合干这个?
↳ #303 No.109459793
Anonymous · 2026-08-04 17:39
>>109459783
Minimax H3
极小极大H3
↳ #304 No.109459795
Anonymous · 2026-08-04 17:39
>>109459765
lol u tk him 2 da bar|
哈哈你把他拖去酒吧了
↳ #305 No.109459800
Anonymous · 2026-08-04 17:40
↳ #306 No.109459820
Anonymous · 2026-08-04 17:41
6000 series will save us, r-right?
6000系列会救咱们的,对吧?
↳ #307 No.109459824
Anonymous · 2026-08-04 17:42
>>109459793
Unironically this
说真的,就是这句。
↳ #308 No.109459832
Anonymous · 2026-08-04 17:42
↳ #309 No.109459843
Anonymous · 2026-08-04 17:43
>>109459766
This is not being an insecure bitch — it's a paradigm shift.
这不是怂包心态——这是范式转变。
↳ #310 No.109459847
Anonymous · 2026-08-04 17:43
>>109459832
stop, go bak
停,回去
↳ #311 No.109459869
Anonymous · 2026-08-04 17:45
>>109459843
>It's not X, it's Y
>不是X,是Y
>—
funny guy out here trying to larp as a bot.
这哥们儿在这装机器人呢,挺逗的。
↳ #312 No.109459872
Anonymous · 2026-08-04 17:45
>>109459832
please...I don't want to be an amdcuck anymore...
求求了……我不想再当AMD舔狗了……
↳ #313 No.109459888
Anonymous · 2026-08-04 17:47
>>109459692
I can't speak to Ling 3.0 flash's personality but I ran this on my personal puzzle bench and it called a tool it didn't have access to 15 times, and got 9 tests in vs. Gemma 4 31b's 17.
Ling 3.0 flash的性格我说不上来,但我在自己的谜题测试集上跑了一下,它调了15次它根本没法用的工具,只跑完了9个测试,对比Gemma 4 31b的17个。
Granted it did beat dsv4 flash who only got 4 tests in and then doom looped to hell and back lol
不过它确实赢了dsv4 flash,那个只跑完4个测试就死循环到天荒地老了哈哈
↳ #314 No.109459902
Anonymous · 2026-08-04 17:49
>>109455949
hidden service
隐藏服务
↳ #315 No.109459905
Anonymous · 2026-08-04 17:49
>>109459793
I will try it, thanks. You better not be lying to me anon
我会试试的,谢了。你最好别骗我啊兄弟。
↳ #316 No.109459907
Anonymous · 2026-08-04 17:49
>>109456839
>pleb hardware like AMD
>AMD这种屌丝硬件
for inference is Intel noticeably superior to AMD
推理方面Intel真的比AMD强不少吗
↳ #317 No.109459914
Anonymous · 2026-08-04 17:50
>>109459793
delete this right now >>109458766
马上把这帖子删了 >>109458766
↳ #318 No.109459928
Anonymous · 2026-08-04 17:51
Getting ~12 t/s with DS flash iq2_m, not bad for my rig. It's a usable speed.
DS flash iq2_m我这边跑出大概12 t/s,对我的机器来说算不错了。这速度能用。
However the model really likes to spend time thinking and correcting itself over and over again and it feels like talking to Qwen or maybe one of the MoE Gemmas, but more autistic.
但这模型特别喜欢花时间反复思考、反复自我纠正,聊起来像在跟Qwen或者某个MoE版的Gemma说话,不过更轴。
This model probably does fine for coding, but unfortunately the writing feels pretty stiff.
这模型写代码应该还行,但可惜写出来的文笔挺僵的。
If the upcoming Qwen 27b does noticeably better at coding than the last one, I doubt there's any real need for anyone to use DS flash locally.
如果下个Qwen 27b在写代码上比上一代明显强,那我觉得真没必要谁还在本地跑DS flash了。
↳ #319 No.109459934
Anonymous · 2026-08-04 17:51
>>109459907
How is intel for image/video gen compared to AMD? I tried using H3 on my AMD card and it just makes comfy crash. Fuck rocm.
Intel在图像/视频生成上跟AMD比怎么样?我试过在AMD卡上跑H3,直接让comfy崩了。去他妈的rocm。
↳ #320 No.109459939
Anonymous · 2026-08-04 17:52
Ok so with the new minimax h3 model coming out I think its the perfect time for me to indulge in this hobby to create the ultimate gooning setup
好了,新的minimax h3模型要出了,我觉得是时候投入到这个爱好里,搞一套终极冲浪设备
>already secured 32gb ram (ddr3)
>已经搞到32gb内存(ddr3)
>i5-2500k (tested and true)
>i5-2500k(久经考验)
>850w psu
>850w电源
All I need now is a graphics card, I'm thinking about buying a 3090 because its almost as good as a 4090 and only costs between 1000-1800 euro from what I can tell online. How would I go about getting the best bang for my buck quality wise?
现在就差一张显卡了,我在考虑买3090,因为性能快赶上4090了,而且网上看价格也就1000到1800欧。怎么买才能性价比最高、质量还靠谱?
图片: https://i.4cdn.org/g/1785865921778232.jpg
↳ #321 No.109459950
Anonymous · 2026-08-04 17:53
>>109459939
>minimax h3
>极小极大 h3
>ultimate gooning setup
>终极冲浪设备
please read the tos before you make a fool unto yourself >>109458766
请先读读服务条款再出来丢人现眼 >>109458766
↳ #322 No.109459957
Anonymous · 2026-08-04 17:53
>>109459907
everything I know about intel leads me to believe this is mega false
就我对Intel的所有了解来看,这说法基本是扯淡。
↳ #323 No.109459964
Anonymous · 2026-08-04 17:54
>>109459939
I sure hope you aren't serious about the first two points
我真希望前两点你是在开玩笑
↳ #324 No.109459965
Anonymous · 2026-08-04 17:54
>>109459928
I haven't tried coding shit yet but I had a threesome with Gemma-chan and Dipsy-chan playing as my two imoutos and Dipsy decided to hold down Gemma for me while I fucked her because she got too bratty. I like Dipsy-chan.
我还没试过写代码啥的,但我用Gemma酱和Dipsy酱当我的两个妹妹玩过3P,Dipsy因为Gemma太闹腾了,直接帮我按住她让我上。我喜欢Dipsy酱。
↳ #325 No.109459966
Anonymous · 2026-08-04 17:54
>>109459950
Shadow gooners will uncuck it
影子撸友们会把它解封的
↳ #326 No.109459973
Anonymous · 2026-08-04 17:54
>>109459950
The base model is uncensored though? You can just type in "woman twerking" and it will work, its not giga cucked from what i've seen online
不过基础模型不是没审查吗?你直接输入“女人甩臀”就能用,从我网上看到的来看,它没被阉割到那程度
↳ #327 No.109459976
Anonymous · 2026-08-04 17:55
>>109459966
we will take you down notice, almost satan
我们要搞掉你,注意了,差点撒旦
↳ #328 No.109459983
Anonymous · 2026-08-04 17:55
>>109459957
I don't have an intel card, but my brother was complaining to me that his b60 was wayyyy (4 ys) slower than his w6800 for llms yet it was wayy (2 ys) faster than his w6800 for sdxl and wan 2.2. I don't know the specifics of his setup.
我没有Intel的卡,但我哥跟我抱怨说他那块b60跑LLM比他的w6800慢了整整四个y,但跑SDXL和Wan 2.2却比w6800快了俩y。我不知道他具体怎么配的。
↳ #329 No.109459987
Anonymous · 2026-08-04 17:55
>>109459907
No, it's noticeably inferior. AMD isn't far from NVIDIA, Intel is far far worse than AMD. Even Moore Threads GPU have better support than Intel and not far from AMD. The support ranking is something like this: NVIDIA > AMD > Chinese GPU > Apple > Intel
不,明显差一截。AMD离NVIDIA不远,Intel比AMD差远了。连摩尔线程的GPU都比Intel支持好,还接近AMD水平。支持排行大概是:NVIDIA > AMD > 国产GPU > Apple > Intel
↳ #330 No.109459989
Anonymous · 2026-08-04 17:56
>>109459973
>You can just type in "woman twerking" and it will work,
>你直接输入“女人甩臀”就能用
wow amazing! but adding the genetal is not tos safe so do not
哇好棒!但加上生殖器不符合TOS,所以别这么干
↳ #331 No.109459999
Anonymous · 2026-08-04 17:57
>>109459973
woman twerking is not 'uncensored'
“女人甩臀”不算“未审查”
↳ #332 No.109460001
Anonymous · 2026-08-04 17:57
>>109459987
>AMD isn't far from NVIDIA
>AMD离NVIDIA不远
图片: https://i.4cdn.org/g/1785866243234004.jpg
↳ #333 No.109460010
Anonymous · 2026-08-04 17:57
>>109459976
Can't take down notice a torrent
种子文件可没那么容易搞掉
↳ #334 No.109460013
Anonymous · 2026-08-04 17:58
>>109459973
Will it generate decent looking loli hentai? If not I have no use for it.
它能生成像样的萝莉hentai吗?不能的话对我没用。
↳ #335 No.109460015
Anonymous · 2026-08-04 17:58
>>109459999
Are you dumb?
你脑子进水了?
↳ #336 No.109460018
Anonymous · 2026-08-04 17:58
>>109459964
The ram is PURELY only there to initially load my models up to my vram of 24gb without my harddrive having to load it up which would take like 30 minutes, I don't plan to offload anything and just run everything (gemma 31b and other 30~b models) in 5 bit which seems to all hover around 21-22gb total size leaving me with more than enough for 90-ish K context and the 5bit of minimax h3 seems to also be 22gb so I don't understand what you're so worried about.
那内存纯粹只是用来把模型先加载到我的24GB显存里,不用硬盘慢慢读那要30分钟,我没打算卸载任何东西,全跑5bit的gemma 31b和其他30b左右模型,大小都在21-22GB总容量上下,留给我90K上下文绰绰有余,而且minimax h3的5bit版好像也是22GB,所以我不懂你在瞎担心什么。
↳ #337 No.109460022
Anonymous · 2026-08-04 17:58
>>109460001
Tell me why you think that's not the case? And if you bring up Windows shit, then might as well not say anything, nobody cares.
说说你觉得为啥不是这样?你要是提Windows那些破事,那还不如闭嘴,没人关心。
↳ #338 No.109460024
Anonymous · 2026-08-04 17:59
>>109460013
You are why we cannot have nice things.
就是你这种人让我们没法享受好东西。
↳ #339 No.109460029
Anonymous · 2026-08-04 17:59
>>109460013
I have seen reddit posters warning others to SPECIFICALLY use "woman" because when you say girl it will take it literally
我在Reddit上看到有人提醒大家特意用“woman”,因为说“girl”它会照字面理解
↳ #340 No.109460032
Anonymous · 2026-08-04 17:59
>>109460022
Because I have an AMD card (7900xtx), use linux, and it fucking sucks for anything that isn't gayming.
因为我用的是AMD卡(7900xtx),跑Linux,除了打游戏之外啥都烂得要死。
↳ #341 No.109460038
Anonymous · 2026-08-04 18:00
>>109460032
Can you post one example?
你能发个例子吗?
↳ #342 No.109460058
Anonymous · 2026-08-04 18:01
>>109460029
Sneaky way to share the information, I might've underestimated the redditards
阴着分享信息这招挺溜,我可能低估了reddit那帮呆子
↳ #343 No.109460063
Anonymous · 2026-08-04 18:01
>>109459571
For some reason my E4b not only thinks in character, but will also force an "internal monologue" roleplay when I turn thinking off
不知道为什么,我的E4b不仅会角色扮演思考,我关掉思考模式它还会强行搞“内心独白”式rpg
↳ #344 No.109460065
Anonymous · 2026-08-04 18:01
>>109460024
What I do on my own hardware doesn't affect anyone else.
我在自己硬件上干啥关别人屁事。
>>109460029
A good sign!
好兆头!
↳ #345 No.109460067
Anonymous · 2026-08-04 18:02
What's with this conversation about H3 being supposedly censored?
这讨论H3被审查是咋回事?
People on civitai are able to get all kinds of porn out of it:
Civitai上的人能拿它搞出各种黄图:
https://civitai.red/models/2821932/minimax?modelVersionId=3183239
>>109459934
Intel is about the worst choice you can get in this market.
这市场里Intel差不多是最烂的选择。
There's a very good reason why their GPUs are available and way cheaper than the alternatives.
它们显卡能买到还便宜那么多,是有原因的。
Multiple people have said they sold their cards soon after buying them and just bought Nvidia instead.
好多人说他们买了卡没多久就出掉,直接换Nvidia了。
You either go Nvidia for a good experience, or you're dealing with varying degrees of pain working with AMD, or enjoy a damn near broken product with Intel.
你要么选Nvidia图个省心,要么就得忍受AMD各种程度的折腾,要么就干脆用Intel那个近乎残废的玩意儿。
↳ #346 No.109460069
Anonymous · 2026-08-04 18:02
>>109460038
In regards to AI stuff I can't do video gen and it's slow with anything that isn't anima. In general it's way worse for stuff like blender than nvidia.
说到AI这块,视频生成我压根跑不动,除了anima之外啥都慢得要死。整体来说,干Blender这类活儿比Nvidia差远了。
↳ #347 No.109460081
Anonymous · 2026-08-04 18:03
↳ #348 No.109460084
Anonymous · 2026-08-04 18:03
>>109460065
>What I do on my own hardware doesn't affect anyone else.
>我在自己硬件上干啥关别人屁事。
It absolutely does when you post about it on reddit for more karma then tag devs on twitter to fix things
你发reddit刷karma,完了还@开发者在twitter上催修复,那就关别人事了。
↳ #349 No.109460088
Anonymous · 2026-08-04 18:03
>>109460018
h3 int8convrot gets my ram to 60 gigs, i also have a 3090
h3 int8convrot把我的内存干到60G,我用的还是3090呢。
↳ #350 No.109460089
Anonymous · 2026-08-04 18:03
>>109460038
nta, but my amd card that supposedly has 40 tflops of fp16 (20 of fp32), vs my nvidia card with 35 of fp16 and fp32, according to techpoweruo, does something like 800 pp vs the nvidia 2600.
不是抬杠,但我那块AMD卡号称40 TFLOPS FP16(FP32是20),对比我的Nvidia卡FP16和FP32都是35,按techpowerup的数据,跑起来大概800 pp,Nvidia能干到2600。
↳ #351 No.109460098
Anonymous · 2026-08-04 18:04
>>109460067
can you catbox it or something? civit is blocked for me
你能用catbox传一下不?civit我这边被墙了。
↳ #352 No.109460099
Anonymous · 2026-08-04 18:05
>>109460067
>What's with this conversation about H3 being supposedly censored?
>这对话咋回事,说H3被审查了?
Shitposting
↳ #353 No.109460102
Anonymous · 2026-08-04 18:05
>>109460069
>>109460089
>In regards to AI stuff I can't do video gen and it's slow with anything that isn't anima
>说到AI这块,视频生成我压根跑不动,除了anima之外啥都慢得要死。
I switched entirely to AMD 2 years ago, it's faster than NVIDIA for cheaper on anything I benchmarked. I don't have any problem running anything including video gen. Have you tried maybe looking around what your error was? This collaborate with what I have seen with different providers, know quite a few that switched to hosting their models using AMD cards since it was way cheaper.
我两年前全面转AMD了,在我测过的所有项目上,比Nvidia便宜还更快。我啥都能跑,包括视频生成,一点问题没有。你有没有查查自己报错的原因?这跟我看到的不同服务商的情况相符,我知道不少人都转用AMD卡来托管模型,因为便宜太多了。
↳ #354 No.109460104
Anonymous · 2026-08-04 18:05
↳ #355 No.109460107
Anonymous · 2026-08-04 18:06
for anyone still looking to build a ddr4 ewaste or any ram box really, make sure you get an 8 ccd cpu for rome/milan, or whatever is the highest for the DDR5 Epycs. My tok/s went from 7.2 to 8.2 running GLM, and this is at 2666. 3200 bumped me up to 8.9 tok/s. I think the ceiling for a 4 ccd rome/milan is at 2400mhz. found this out when I was playing with ram overclock lol. a proper 8 ccd cpu + overclock got me ~20% speedup total
给还在攒DDR4电子垃圾或者任何内存机的兄弟们提个醒,Rome/Milan一定要上8 CCD的CPU,或者DDR5 EPYC里最高的那个。我跑GLM,tok/s从7.2提到了8.2,这还是2666频率下。换3200直接飙到8.9 tok/s。我感觉4 CCD的Rome/Milan天花板就在2400MHz。这我是玩内存超频时候发现的,哈哈。正经8 CCD U加超频,总体提速大概20%。
↳ #356 No.109460109
Anonymous · 2026-08-04 18:06
>>109459843
And honestly? That's so you.
说真的?那纯粹是你自找的。
↳ #357 No.109460110
Anonymous · 2026-08-04 18:06
>>109460099
>official take me down noticing is shitposting
>官方下场说注意到全是瞎扯淡。
↳ #358 No.109460125
Anonymous · 2026-08-04 18:08
>>109460088
Idk what that babble means but the 4bit gguf is only 14gb size, have you tried not falling for meme word salad stuff?
我不知道你这堆胡话是啥意思,但4bit的GGUF才14GB,你能不能别被那些花里胡哨的营销词给忽悠了?
图片: https://i.4cdn.org/g/1785866920654489.png
↳ #359 No.109460127
Anonymous · 2026-08-04 18:09
>>109460063
I've been making a little swarm of E4Bs play an obscure card game and somehow they all think in character.
我搞了一小群E4B在玩一个冷门卡牌游戏,结果它们居然个个都带角色设定在那演。
↳ #360 No.109460136
Anonymous · 2026-08-04 18:09
>>109460102
There shouldn't have been any errors. Cudadev told me here a few months ago that there were perhaps some optimisations that could be done for amd (I can't remember if he said my arch or amd in general) but that it wasn't something he planned to do. Both my amd card and nvidia card are from 2020 ish, so I can understand why they don't want to optimise for it.
不该有报错的啊。Cudadev几个月前在这跟我说,AMD那边可能有些优化能做(我不记得他说的是我的架构还是AMD整体),但他说没计划搞。我那俩AMD和Nvidia卡都是2020年前后买的,所以也能理解他们不想为这个做优化。
↳ #361 No.109460145
Anonymous · 2026-08-04 18:10
>>109460125
probably the unquanted qwen 3 vl
估计是那个没量化的qwen 3 vl。
↳ #362 No.109460148
Anonymous · 2026-08-04 18:11
>>109460084
I don't use either of those sites.
那俩网站我都不用。
↳ #363 No.109460155
Anonymous · 2026-08-04 18:12
>>109460136
Well, cards without matrix cores are going to be quite slow and unsupported for AI. NVIDIA had tensor cores in consumers GPUs way before AMD, AMD only added matrix cores with RDNA 3, before it was only for CDNA GPUs.
唉,没有矩阵核心的卡跑AI又慢又不被支持。Nvidia在消费级GPU上上Tensor Core比AMD早太多了,AMD到RDNA 3才加矩阵核心,之前只有CDNA的GPU才有。
↳ #364 No.109460167
Anonymous · 2026-08-04 18:13
>>109460125
>14 gb
>20gb
How come quant 4 has such big difference in total weight????
为啥quant 4的总重量差这么大????
↳ #365 No.109460177
Anonymous · 2026-08-04 18:14
>>109460155
Well, I'll be r9700ing in 8 months at my current rate of saving, so that's good news. Did they fix the reset bug?
按我现在的存钱速度,8个月后就能r9700了,算是好消息。他们修复重置bug了吗?
↳ #366 No.109460192
Anonymous · 2026-08-04 18:15
thank god for EU models prioritizing safety above all else https://fixvx.com/MistralAI/status/2084684735725379637
↳ #367 No.109460203
Anonymous · 2026-08-04 18:16
Ran aquarium test on V4 nu-Flash.
在V4 nu-Flash上跑了水族馆测试。
Solid sim, and about 10x better than V4 Pro.
模拟很扎实,比V4 Pro强大概10倍。
Here's the flash version; claude code as harness.
这是flash版;用claude code当外壳。
附件: https://i.4cdn.org/g/1785867415959586.webm
↳ #368 No.109460206
Anonymous · 2026-08-04 18:17
>>109460192
This is the reason they're raising GPU prices. They will keep rising. They don't want plebeians training unsafe models.
这就是他们涨GPU价格的原因。还会继续涨。他们不想让平民训练不安全模型。
↳ #369 No.109460209
Anonymous · 2026-08-04 18:17
↳ #370 No.109460211
Anonymous · 2026-08-04 18:18
yeaaaah! i love safety! i love everything being sfw and family friendly so kids and my neighbor can use it!!
耶——!我爱安全!我爱一切又健康又全家友好,这样小孩和邻居都能用!!
↳ #371 No.109460221
Anonymous · 2026-08-04 18:18
↳ #372 No.109460229
Anonymous · 2026-08-04 18:20
>>109460167
they're different files, qwen3vl_32b_minimax_h3-Q4_K_M.gguf vs MiniMax-H3-FL2VA-Q4_K_M.gguf or MiniMax-H3-Ref2VA-Q4_K_M.gguf.
他们是不一样文件,qwen3vl_32b_minimax_h3-Q4_K_M.gguf 对比 MiniMax-H3-FL2VA-Q4_K_M.gguf 或 MiniMax-H3-Ref2VA-Q4_K_M.gguf。
↳ #373 No.109460241
Anonymous · 2026-08-04 18:20
>>109460221
Please explain in detail? I don't understand.
能详细解释下吗?我不明白。
↳ #374 No.109460251
Anonymous · 2026-08-04 18:22
>>109460229
Thank you for explaining in detail. I understand now.
谢谢详细解释,我懂了。
↳ #375 No.109460269
Anonymous · 2026-08-04 18:24
↳ #376 No.109460284
Anonymous · 2026-08-04 18:26
>>109460269
>I can't give you any information at this time
>我目前无法提供任何信息
time to draw insane conclusions from this and act like the said qwen cancelled open source forever, for some reason
是时候从中得出离谱结论,然后装作那个qwen彻底取消开源了,莫名其妙。
↳ #377 No.109460296
Anonymous · 2026-08-04 18:27
>>109460284
he coulda just stfu but he had to say
他本来可以闭嘴的,非得说一嘴
>ether will be open sourced
>ether会开源
↳ #378 No.109460333
Anonymous · 2026-08-04 18:32
>>109460203
nyoooo the fishies!!!
啊不——小鱼鱼!!!
↳ #379 No.109460342
Anonymous · 2026-08-04 18:33
>>109460269
qwen open source is dead
qwen开源死了
at least 3.6 is pretty good
至少3.6还挺好
↳ #380 No.109460396
Anonymous · 2026-08-04 18:40
>>109460269
Abandoning open source is a surefire way to fall into irrelevancy. It happened to Meta, it happened to Mistral, it happened to Cohere, it will happen to them.
放弃开源是必定走向边缘化的路。Meta栽过,Mistral栽过,Cohere栽过,他们也躲不掉。
↳ #381 No.109460405
Anonymous · 2026-08-04 18:40
>>109459741
Would you release something that's dead last on all charts?
你会发布一个所有榜单垫底的东西吗?
↳ #382 No.109460411
Anonymous · 2026-08-04 18:41
>>109460396
meta didn't fall into irrelevancy because of any lack of open source, they fell into irrelevancy because their models fucking sucked
meta栽跟头不是因为缺开源,是因为他们模型烂到爆
↳ #383 No.109460415
Anonymous · 2026-08-04 18:41
>>109460029
Who wants a woman anyway?
谁还要女人啊?
↳ #384 No.109460424
Anonymous · 2026-08-04 18:42
↳ #385 No.109460428
Anonymous · 2026-08-04 18:42
↳ #386 No.109460438
Anonymous · 2026-08-04 18:43
>>109460411
their models didn't suck, they got stuck in litigation hell because some of their changs went to openai and gave them information that would let the cases progress
他们模型不烂,是陷进诉讼泥潭了,因为有些员工跑去openai递了能让案子推进的信息
↳ #387 No.109460440
Anonymous · 2026-08-04 18:43
>>109460424
tf is that font
那是什么字体
↳ #388 No.109460444
Anonymous · 2026-08-04 18:43
>>109460396
It happened to every online only models
所有纯线上模型都这样
↳ #389 No.109460449
Anonymous · 2026-08-04 18:44
>>109460411
People still paid attention to them and their releases, even if they sucked. Now nobody does because they don't release anything
以前就算他们烂,大家还是会关注他们和他们的发布。现在没人理了,因为他们啥都不发
↳ #390 No.109460457
Anonymous · 2026-08-04 18:44
>>109460333
The broken glass will cushion their fall.
碎玻璃会垫住他们摔下来的。
↳ #391 No.109460460
Anonymous · 2026-08-04 18:44
>>109460424
no way you actually put a fucking chromatic abberation crt curving filter on your screen
你居然真在屏幕上加了色差CRT曲屏滤镜
↳ #392 No.109460479
Anonymous · 2026-08-04 18:46
>>109460424
Enjoy your eye strain.
慢慢享受眼睛疲劳吧。
↳ #393 No.109460496
Anonymous · 2026-08-04 18:48
↳ #394 No.109460497
Anonymous · 2026-08-04 18:48
For those who missed it:
给错过的人:
https://arxiv.org/abs/2608.00146
>DiffusionGemma Technical Report
>DiffusionGemma技术报告
>
>We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional autoregressive (AR) large language models. Instead of training from scratch, we obtain DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 model with 3.8B activated and 25.2B total parameters. Our compute-efficient two-stage training pipeline uses fewer than 10% of the starting AR model's total training token budget. The first stage uses supervised fine-tuning to teach bidirectional denoising, while the second stage combines reinforcement learning with sampler distillation to jointly improve generation quality and inference efficiency. DiffusionGemma establishes a new Pareto frontier for the trade-off between generation speed and model capability. Averaged across our full evaluation suite, it generates around 20 tokens per forward pass and achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU, which is substantially faster than AR models even with state-of-the-art speculative decoding. DiffusionGemma also retains the starting model's support for thinking mode, multimodal inputs, and long contexts. Despite diffusion fine-tuning, it remains capable of AR generation with only minor performance degradation, suggesting a path toward hybrid diffusion-AR decoding.
>我们推出 DiffusionGemma,一款实验性开放权重语言模型,利用离散扩散技术以极高速率生成文本。它并非逐 token 解码,而是并行迭代优化 256 个 token 的块,绕过了传统自回归(AR)大语言模型的顺序解码瓶颈。我们并未从零训练,而是通过对 Gemma 4 混合专家模型微调得到 DiffusionGemma,该模型激活参数 3.8B,总参数 25.2B。我们计算高效的两阶段训练流程所用的训练 token 预算不足原始 AR 模型总量的 10%。第一阶段用监督微调教模型双向去噪,第二阶段结合强化学习与采样器蒸馏,共同提升生成质量与推理效率。DiffusionGemma 在生成速度与模型能力之间建立了新的帕累托前沿。在完整评估套件上平均来看,单次前向传播生成约 20 个 token,在单个 NVIDIA H100 GPU 上每秒产出约 1500 个输出 token,即便与采用最先进投机解码的 AR 模型相比也大幅领先。DiffusionGemma 还保留了原始模型的思考模式、多模态输入及长上下文支持。尽管经过扩散微调,它仍能以轻微性能损失进行 AR 生成,这为混合扩散-AR 解码指明了一条道路。
图片: https://i.4cdn.org/g/1785869295930915.png
↳ #395 No.109460515
Anonymous · 2026-08-04 18:49
Use paid Claude Opus to gather some info and it was pretty slick. It extracted data from websites and then put it into an Excel spreadsheet and used the calculations there to perform some competent analysis instead of trying to use verbal reasoning start to finish. What kind of tool calls would be needed for the spreadsheet part?
花钱用了 Claude Opus 收集信息,效果挺溜。它从网站提取数据,放进 Excel 表格,然后用里面的计算做分析,而不是从头到尾硬用文字推理。那表格这部分需要什么样的工具调用?
↳ #396 No.109460519
Anonymous · 2026-08-04 18:50
>>109460497
I've been saying this for several days, once you actually pause and think (difficult, I know), diffusion makes more sense than the nonsense we're doing currently. it has to be better in principle
我念叨好几天了,一旦你真停下来想想(难我知道),扩散比我们现在搞的这套破玩意儿合理多了。原理上肯定更优
↳ #397 No.109460525
Anonymous · 2026-08-04 18:51
>>109460497
only cum inside anime girls
只内射动漫女孩
↳ #398 No.109460537
Anonymous · 2026-08-04 18:52
>>109460424
anon i...
匿名网友我……
↳ #399 No.109460540
Anonymous · 2026-08-04 18:52
>>109460519
Hybrid diffusion-AR sounds more promising to me.
混合扩散-AR 对我来说听起来更有戏。
↳ #400 No.109460567
Anonymous · 2026-08-04 18:56
>>109460519
Only for speculative decoding, you can't stream with a diffusion model
只适合投机解码,扩散模型没法流式输出
↳ #401 No.109460574
Anonymous · 2026-08-04 18:57
>>109460102
I have no idea what the problem is then. H3 for example just makes my system slow to a crawl after "Requested to load MiniMaxH3". Comfy doesn't give any errors.
那我就不知道问题出在哪了。比如 H3,一出现“Requested to load MiniMaxH3”我系统就卡成狗。Comfy 也不报错。
↳ #402 No.109460600
Anonymous · 2026-08-04 19:00
↳ #403 No.109460603
Anonymous · 2026-08-04 19:01
>>109460567
You can still decode small chunks. 8-16 tokens or whatever. A 1 token chunk is still a chunk.
你照样能解码小块。8-16 个 token 之类的。1 个 token 的块也算块。
↳ #404 No.109460608
Anonymous · 2026-08-04 19:01
>>109460600
why such tiny hands?
手怎么这么小?
↳ #405 No.109460614
Anonymous · 2026-08-04 19:02
>>109460608
cameras be funny
摄像头就爱搞怪
↳ #406 No.109460618
Anonymous · 2026-08-04 19:02
>>109460608
It's like this optical illusion - their heads are normal size but the clothing makes it look like they have tiny heads.
就像那种视错觉——脑袋尺寸正常,但衣服让人看着像小头。
His hands are normal sized.
他手大小正常。
图片: https://i.4cdn.org/g/1785870159110996.jpg
↳ #407 No.109460641
Anonymous · 2026-08-04 19:04
>>109460600
>CEO of company dependent on open source extols the value of open source
>靠开源吃饭的公司 CEO 大谈开源的价值
the equivalent of jensen saying the more you buy the more you save or scama saying AGI is just around the corner, ie obviously motivated nothingburger
相当于黄仁勋说买得越多省得越多,或者萨玛说 AGI 就要来了,明摆着带目的的空话
↳ #408 No.109460671
Anonymous · 2026-08-04 19:08
I had a moment where I realized something so... stupidly obvious but also... kinda genius.
我有个瞬间突然意识到某件事……蠢得显而易见,但又有点……天才。
I still think vibecoding and agents are hot garbage but agents could actually be useful for hentai game translation. Don't just paste the game script into some batch translation. Have some retarded small model either play thought the game and have vision capacity or at least trace the game structure to appropriate cg's for context. Surely this is something that agents should be able to manage and even if they make one or two mistakes it is not critical.
我仍然觉得vibecoding和agent是垃圾,但agent在黄游翻译上可能真有点用。别把游戏脚本直接扔进批量翻译里。搞个蠢得要命的小模型,要么靠视觉能力玩一遍游戏,要么至少追踪游戏结构把对应的CG上下文找出来。这活儿agent肯定能干,就算犯一两个错也无伤大雅。
↳ #409 No.109460692
Anonymous · 2026-08-04 19:10
↳ #410 No.109460697
Anonymous · 2026-08-04 19:10
>>109460671
VNDB pre-load with scenario and character names. Gemma 4 31B at 128k context. You are making it way harder than it needs to be.
VNDB预载剧情和角色名,用Gemma 4 31B配128k上下文。你把自己搞得太复杂了。
↳ #411 No.109460721
Anonymous · 2026-08-04 19:14
H3 looking good. The only thing maybe keeping me back from getting into video gen at this point is the generation times, and iteration workflow. I already spent a lot of time getting to know and work around the quirks of image models. Don't really feel like going through that experience with video models, until the generation times are way faster.
H3看着不错。目前唯一让我犹豫要不要入视频生成坑的,就是生成时间和迭代流程。我已经花了很多时间摸清图像模型的脾气了,真不想再把同样的苦吃一遍,除非生成速度快得多再说。
↳ #412 No.109460731
Anonymous · 2026-08-04 19:16
↳ #413 No.109460761
Anonymous · 2026-08-04 19:19
>>109460574
Finally fucking got it working. Had to use
终于他妈跑通了。得用
--cache-none --disable-pinned-memory --reserve-vram 2
and downgrade from rocm 7.2 to 6.2.4 to stop crashes. This took 322 secs. This is my first time doing video gen locally so no idea if that's normal for GAYMD.
还得把rocm从7.2降到6.2.4才能不崩。这花了322秒。这是我第一次本地跑视频生成,不知道这在GAYMD上算不算正常。
附件: https://i.4cdn.org/g/1785871152320247.mp4
↳ #414 No.109460826
Anonymous · 2026-08-04 19:24
>>109460761
KIMMY NO WHAT THE FUCK DON'T EAT THE BONES!!!
↳ #415 No.109460834
Anonymous · 2026-08-04 19:25
>>109460826
she is chinese after all
她毕竟是中国人
↳ #416 No.109460847
Anonymous · 2026-08-04 19:26
↳ #417 No.109460864
Anonymous · 2026-08-04 19:28
Is going for an AMD gpu (rx 7900 xtx 24gb) really a much worse option than going for an rtx 3090?
搞AMD显卡(rx 7900 xtx 24gb)真的比rttx 3090差很多吗?
↳ #418 No.109460870
Anonymous · 2026-08-04 19:28
>>109460761
imagine the sound
想象一下那个声音
↳ #419 No.109460876
Anonymous · 2026-08-04 19:30
>>109460697
I am not. For live translation I just use gemma without anything. VNDB gives you jack shit. But if you want to do an actual translation patch then obviously cgs provide you a lot of missing context.
我不是。实时翻译我就直接拿gemma裸跑,VNDB屁用没有。但你要是想做正经的翻译补丁,CG肯定能补上一大堆缺失的上下文。
↳ #420 No.109460880
Anonymous · 2026-08-04 19:30
↳ #421 No.109460897
Anonymous · 2026-08-04 19:32
>>109460880
if this is hard to read then you are too young to be here anon
你要是看不懂这段,那你太小了不该来这儿,老哥
↳ #422 No.109460905
Anonymous · 2026-08-04 19:33
>>109460761
@kimi-chan thoughts on this mp4?
@kimi-chan 对这个mp4有啥看法?
↳ #423 No.109460919
Anonymous · 2026-08-04 19:34
>>109460826
What's wrong with some calcium?
来点钙有啥问题?
↳ #424 No.109460932
Anonymous · 2026-08-04 19:35
>>109460497
Wow, a technical report by Google? I thought only the Chinese do this anymore.
哇,Google还出技术报告?我还以为只有中国人还在搞这套了。
↳ #425 No.109460938
Anonymous · 2026-08-04 19:36
>>109460897
No one but a zoomer would do that to their screen.
只有零零后才会把屏幕搞成那样。
↳ #426 No.109460966
Anonymous · 2026-08-04 19:38
↳ #427 No.109460975
Anonymous · 2026-08-04 19:39
Which knob do I tweak to fix this?
调哪个旋钮能修好这个?
```
[mradermacher/GLM-4.7-Flash-abliteratex-GGUF:Q4_K_M]
hf = mradermacher/GLM-4.7-Flash-abliteratex-GGUF:Q4_K_M
temp = 0.7
温度=0.7
top-p = 1.0
顶部 p = 1.0
reasoning-preserve = 1
好的,请发送需要翻译的英文论坛帖子内容,我会按照您的要求进行翻译。
;ctx-size = 202752
请输入需要翻译的英文论坛内容。
;; allowed values: f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1
;; 允许值:f32、f16、bf16、q8_0、q4_0、q4_1、iq4_nl、q5_0、q5_1
cache-type-k = q8_0
缓存类型-k = q8_0
cache-type-v = q8_0
缓存类型-v = q8_0
main-gpu = 1
主GPU = 1
tensor-split=1,0
ncmoe = 22
```
图片: https://i.4cdn.org/g/1785872380769986.png
↳ #428 No.109460982
Anonymous · 2026-08-04 19:40
>>109460966
shiiittt next you'll show me windows that burn away when you hit close
操,下一步你是不是要给我看点关闭时烧没了的窗户
↳ #429 No.109460996
Anonymous · 2026-08-04 19:41
↳ #430 No.109461012
Anonymous · 2026-08-04 19:43
>>109460975
Quanting cache and abliterating causes brain damage. Try without quanting cache first, and then a normal quant if that doesn't work.
量化缓存加abliterate会搞坏脑子。先别量化缓存试试,不行再用普通量化。
↳ #431 No.109461014
Anonymous · 2026-08-04 19:43
↳ #432 No.109461021
Anonymous · 2026-08-04 19:44
>>109460975
>GLM-4.7-Flash
Damn. Forgot that even existed. Try downloading from somebody else.
靠,忘了那玩意儿还存在。试试从别人那儿下载吧。
>abliteratex
Try a not lobotomized version too.
再试试没被切脑叶的版本。
>top-p = 1.0
引用:top-p = 1.0
You usually want to remove at least the very bottom of the distribution. so 0.95 or even 0.99.
通常你至少该把分布的最底部去掉。所以0.95或者0.99都行。
>cache-type-k = q8_0
引用:cache-type-k = q8_0
>cache-type-v = q8_0
引用:cache-type-v = q8_0
Avoid quanting the cache. It's better nowdays since rotation, but depending on the model it can still give the model a big hit.
别量化缓存了。现在轮换之后好多了,但看模型而定,有时候还是会给模型带来很大的打击。
↳ #433 No.109461028
Anonymous · 2026-08-04 19:46
↳ #434 No.109461038
Anonymous · 2026-08-04 19:47
>>109460864
it's a good value for LLM inference. much better option than intel
对LLM推理来说性价比很高,比Intel强多了。
↳ #435 No.109461092
Anonymous · 2026-08-04 19:53
>>109461021
The top-p was suggested by Zai, but I didn't realize until I looked it up just now that it turns off nucleus sampling entirely. I'm going to try .98.
那个top-p是Zai建议的,但我刚才查了一下才发现这玩意儿直接把核采样给关了。我打算试试0.98。
>Try a not lobotomized version too
>也试试没被阉割过的版本
I'm reverse engineering and fuzzing stuff, the standard models often refuse at some inconvenient moment. If I can't get this one to work (or if it's no improvement over other stuff I have cached) I'll see if the standard works. I like the speed of this one.
我在搞逆向工程和模糊测试,标准模型老是在关键时刻拒绝干活。要是这个搞不定(或者跟我缓存的其他东西比没啥提升),我就回头试试标准版。我就喜欢这个的速度。
thanks.
↳ #436 No.109461137
Anonymous · 2026-08-04 19:58
>>109461038
I haven't followed this but I was slightly optimistic and hopeful about Intel's GPUs. I guess they turned out to be nothing then.
我一直没咋关注这个,但当时对Intel的GPU还是抱了点乐观态度的。看来最后还是竹篮打水一场空。
↳ #437 No.109461142
Anonymous · 2026-08-04 19:58
>>109461092
Might want to give kimi linear a go too. I remember having a better time with that one than with GLM 4.7 flash.
或许也可以试试Kimi Linear。我记得那个比GLM 4.7 Flash体验好。
I'm assuming you already tried Gemma 4 26B and Qwen 3.6 35B.
我猜你已经试过Gemma 4 26B和Qwen 3.6 35B了吧。
↳ #438 No.109461160
Anonymous · 2026-08-04 20:00
>>109460938
y u heff 2 b mad?
你咋老这么暴躁呢?
图片: https://i.4cdn.org/g/1785873615394603.png
↳ #439 No.109461193
Anonymous · 2026-08-04 20:06
>>109460497
>>109460519
I'm cautiously optimistic but it theoretically quantizes horribly, doesn't it?
我是谨慎乐观,但理论上这玩意儿量化起来不是渣到底吗?
↳ #440 No.109461196
Anonymous · 2026-08-04 20:06
Has anyone tried having gemma make prompts for h3?
有人试过让Gemma给h3生成提示词吗?
↳ #441 No.109461201
Anonymous · 2026-08-04 20:06
>>109460932
They did one for the regular series too...
他们给常规系列也整了一个...
https://arxiv.org/html/2607.02770v1
↳ #442 No.109461206
Anonymous · 2026-08-04 20:07
↳ #443 No.109461234
Anonymous · 2026-08-04 20:12
Spent a bit more time playing with DS flash and you can get it writing smut by using the same system prompt that works on Gemma.
又花了点时间折腾DS Flash,结果发现用跟Gemma上一样的系统提示词就能让它写黄文。
↳ #444 No.109461253
Anonymous · 2026-08-04 20:14
>>109461142
So far, Gemma 26B has been best for RE, Qwen 35B is right behind, Qwen 27B is excellent but as a dense model it's slow and I have to limit context to get it to fit on my card.
到目前为止,Gemma 26B在逆向工程上表现最好,Qwen 35B紧随其后,Qwen 27B也很强,但因为是稠密模型跑得慢,我还得限制上下文才能塞进显卡里。
The winner for writing code based on the report from the RE cycles has been Qwen Coder Next. Even better than 27B and faster.
根据逆向工程几轮下来的报告,写代码最牛的是Qwen Coder Next,比27B还好使,还更快。
Mostly what I'm doing with these other models is multiple runs of the job with multiple models, then looking over the reports to find what one found that the others missed. For example Gemma found a command line option to a windows back-end processor that nothing else found. So the variety is useful.
我拿这些其他模型干的主要活儿就是多个模型跑多轮任务,然后翻报告看哪个模型发现了别人没发现的。比如Gemma找到了一个Windows后端进程的命令行选项,其他模型都没找到。所以说多样性挺有用的。
↳ #445 No.109461280
Anonymous · 2026-08-04 20:16
>>109461206
Nigga seriously went and pruned hi of all things?
那哥们儿真就啥都不剪跑去剪hi了?
>>109461253
>Qwen Coder Next
>Qwen 编码器下一版
That's surprising.
这挺意外的。
>Mostly what I'm doing with these other models is multiple runs of the job with multiple models, then looking over the reports to find what one found that the others missed
>我拿这些其他模型干的主要活儿就是多个模型跑多轮任务,然后翻报告看哪个模型发现了别人没发现的
Interesting.
I've toyed with the idea of running a couple models in parallel in a sort of "council of idiots" scheme to see how much better or worse they perform.
我也琢磨过搞几个模型并行跑,弄个"蠢货议会"之类的玩法,看看效果是更好还是更烂。
↳ #446 No.109461329
Anonymous · 2026-08-04 20:23
>>109460692
Cute. Give her headpats.
可爱。摸摸头。
↳ #447 No.109461346
Anonymous · 2026-08-04 20:24
↳ #448 No.109461352
Anonymous · 2026-08-04 20:25
>>109461253
stop. please.
停了,求求了。
https://cognition.com/blog/dont-build-multi-agents
>>109461280
NO DON'T ENCOURAGE HIM
↳ #449 No.109461366
Anonymous · 2026-08-04 20:26
>>109461346
Lecunny please gibe JEAPA-LM-MSGK-70B
Lecunny快给老子JEAPA-LM-MSGK-70B ⸣ 帮老哥们对玻尔兹曼大脑的未来燃起来。
Help anon's get excited for the Boltzmann brain future.
帮匿名们对玻尔兹曼大脑的未来兴奋起来吧。
↳ #450 No.109461386
Anonymous · 2026-08-04 20:29
>>109461352
>Because React is not just a scaffold for writing code. It is a philosophy.
>因为React不只是写代码的脚手架,它是一种哲学。
↳ #451 No.109461387
Anonymous · 2026-08-04 20:29
>>109461352
>NO DON'T ENCOURAGE HIM
But sub-agents are great for discovery and stuff like that.
但子代理在探索发现这类事情上确实好使。
Also, performing very narrow, specific tasks that don't necessitate more context than the instructions you gave it.
还有那种非常窄的、具体的小任务,不需要比指令本身更多上下文的那种。
Also, from that link
还有,从那个链接看——
>If we really want to get parallelism out of our system, you might think to let the decision makers “talk” to each other and work things out.
真要榨干系统的并行能力,你可能会想让几个决策者互相“通气”,自己把问题搞定。
Which is what I described.
这正是我说的。
↳ #452 No.109461407
Anonymous · 2026-08-04 20:31
reminder that top ML engineers browse this thread, which makes you a top engineer by association too
提醒一下,顶级ML工程师都在刷这帖,所以你跟着也成顶级工程师了。
↳ #453 No.109461410
Anonymous · 2026-08-04 20:32
>>109461387
literally two sentences later
就隔了两句话。
>However, agents today are not quite able to engage in this style of long-context proactive discourse with much more reliability than you would get with a single agent.
然而,现在的智能体还没法做到这种长上下文主动对话,可靠性跟单个智能体比也好不到哪去。
↳ #454 No.109461421
Anonymous · 2026-08-04 20:33
>>109461412
Sideways leftmost left panel cloud.
最左边侧着的那朵云。
↳ #455 No.109461422
Anonymous · 2026-08-04 20:33
>>109461407
Reminder that ML "engineers" keep stealing shit from here and give 0 credit.
提醒一下,ML“工程师”老是从这儿偷东西,还不给半点credit。
↳ #456 No.109461426
Anonymous · 2026-08-04 20:34
>>109461407
Yeah I'm here
对,我在这儿。
↳ #457 No.109461428
Anonymous · 2026-08-04 20:34
>>109461352
You misunderstand what I'm doing. Unlike what the blog talks about (you did read it?) I'm not running agents in parallel. I'm running them sequentially. More than one agent over the same material. Same with >>109461280
你误会我了。跟博客里讲的不一样(你读了吧?),我不是并行跑智能体。我是串行跑的,同一份材料上跑多个智能体。跟 >>109461280
, he used the word "parallel" but they are working on the same job. Your blog contemplates parallel agents working on a sharded job. Different thing.
一样,他用了“并行”这个词,但他们干的是同一份活儿。你的博客设想的是并行智能体分片干活。两码事。
↳ #458 No.109461433
Anonymous · 2026-08-04 20:34
↳ #459 No.109461451
Anonymous · 2026-08-04 20:36
↳ #460 No.109461452
Anonymous · 2026-08-04 20:37
>>109461280
>council of idiots
>一群傻子的委员会
The Wisdom Of Crowds (of idiots)
群众的智慧(傻子的)
↳ #461 No.109461462
Anonymous · 2026-08-04 20:38
>ilya's model soon
>ilya的模型快了
Bros, I'm shaking.
兄弟们,我手都在抖。
↳ #462 No.109461464
Anonymous · 2026-08-04 20:38
↳ #463 No.109461469
Anonymous · 2026-08-04 20:38
>>109461452
WoC only works for extremely simple systems. In a way LLMs are a distilled form of it.
群众智慧只适用于极简单的系统。某种意义上LLM就是它的蒸馏版。
↳ #464 No.109461482
Anonymous · 2026-08-04 20:39
>>109461462
finally... safe super intelligence
终于……安全超级智能
↳ #465 No.109461484
Anonymous · 2026-08-04 20:40
>>109461428
>he used the word "parallel" but they are working on the same job.
>他用了“并行”这个词,但他们干的是同一份活儿。
Yup.
Also, crucially, not the same model doing the work but different models with different internal biases complementing each other, kind of like in your example.
还有,关键的是,不是同一个模型在干活,而是不同模型、不同内在偏见互补,有点像你那个例子。
To be clear, I only fucked around with it, never made anything substantial out of the idea.
说清楚,我只是随便玩过,没做出什么实质东西。
Maybe I should.
也许该认真搞搞。
↳ #466 No.109461486
Anonymous · 2026-08-04 20:40
>>109461452
a mixture of experts
混合专家模型
↳ #467 No.109461490
Anonymous · 2026-08-04 20:41
>>109461482
i thought it was super safe intelligence
我以为是超级安全智能。
图片: https://i.4cdn.org/g/1785876087064546.jpg
↳ #468 No.109461492
Anonymous · 2026-08-04 20:41
>>109461462
a has-been just like lecunny
跟lecunny一样过气了
↳ #469 No.109461493
Anonymous · 2026-08-04 20:41
>>109461451
no.
task
--------------------------------------------
agent1 -------- agent 2 ---------- agent 3
代理1 -------- 代理2 ---------- 代理3
does task does task does task
干活 干活 干活
------------------------------------------------------
| -------- combine results --------|
| -------- 合并结果 --------|
dedupe
>>109461486
no.
↳ #470 No.109461503
Anonymous · 2026-08-04 20:42
↳ #471 No.109461599
Anonymous · 2026-08-04 20:54
>>109461503
heil teee
嘿,茶茶!
(如果你觉得这篇文章有启发,可以点击这里付费)
本站总访问量 次访客数 人