EUREKAR
Eurekar·TOP
捕捉真实世界的英语信号
34信息来源
13,006精选文章
168单词卡片
18照片图片
全部8,248口语2,075外贸4免费1,129帖子4,467新闻1,602hackernews1,197techmeme752tmz657slashdot529techcrunch507arstechnica431随笔330外刊291Cards168simonwillison100sethgodin60图片18information3

Sup /g/

4chan /g/ · A hydrus-like image tagger · 2026-08-19 06:21 · 22 帖 · 原文 ↗

← 上一篇返回列表下一篇 →
#1 Sup /g/
A hydrus-like image tagger · 2026-08-19 06:21

Sup /g/

Just released the first version of naiad-net, a hobby project that I ended up pushing into a usable state. The tl;dr of it is you add picture folders to it, it indexes them into a sqlite db, and you can tag shit.

刚刚发布了 naiad-net 的第一个版本,这是一个业余爱好项目,我最终将其推入可用状态。它的要点是你向其中添加图片文件夹,它将它们索引到一个 sqlite 数据库中,然后你可以标记狗屎。

But I also wanted what the Hydrus network has, in a more lightweight package. So I have a dev server that ingests hydrus ptr updates every night. This way you can also pull tags from the net, in a quicker and only slightly less private way. Basically instead of downloading the entire db, you request buckets via hash prefixes. It's pretty fast this way.

但我也想要 Hydrus 网络拥有的东西,并且是一个更轻量级的包。所以我有一个开发服务器,每天晚上都会接收 Hydrus ptr 更新。通过这种方式,您还可以以更快且私密性稍差的方式从网络中提取标签。基本上,您不需要下载整个数据库,而是通过哈希前缀请求存储桶。这种方式速度相当快。

Shit's made in rust, optimized for speed, interface is wip but imo right now the basic functionality of 'just put tags on my images' works. It's completely slopped together, and I'm actually surprised I got this far. If anyone cares enough I'll keep adding planned features, like plugins to scrape/autotag via model, but right now only basic functionality is done.

该死的,是用 Rust 制作的,针对速度进行了优化,界面是 wip,但在我看来,现在“只需在我的图像上添加标签”的基本功能就可以了。它完全倾斜在一起,我真的很惊讶我能走到这一步。如果有人足够关心,我将继续添加计划的功能,例如通过模型进行抓取/自动标记的插件,但现在只完成了基本功能。

图片: https://i.4cdn.org/g/1787120474869864.png

#2 No.109594825
Anonymous · 2026-08-19 10:52

>>109593708

tried hydrus once and uninstalled the turd 10m later. explain how your fecal matter is different and what it does better than windows explorer.

尝试了一次 Hydrus,10m 后卸载了 turd。解释一下您的排泄物有何不同,以及它比 Windows 资源管理器更好的地方。

#3 No.109594835
Anonymous · 2026-08-19 10:55

>>109593708

what gui library is used?

使用什么图形用户界面库?

#4 No.109594947
Anonymous · 2026-08-19 11:14

>>109594835

Claude webshit

克劳德网络狗屎

#5 No.109595002
Anonymous · 2026-08-19 11:24

>>109593708

A thread died for this. Fuck you.

一个线程为此死亡。去你的。

#6 No.109595046
Anonymous · 2026-08-19 11:32

>>109594825

You unzip it, add an image folder, and 5 minutes later you've got hydrus tags for them.

你解压它,添加一个图像文件夹,5 分钟后你就得到了它们的 Hydrus 标签。

>>109594835

Svelte

>>109594947

Yea

#7 No.109595285
Anonymous · 2026-08-19 12:21

Oh right, forgot to put up the github.

对了,忘记放github了。

https://github.com/scoopscoop/naiad-net

#8 No.109595418
Anonymous · 2026-08-19 12:47

>>109594825

>filtered by autismo software

>由自闭症软件过滤

lol

#9 No.109595448
Anonymous · 2026-08-19 12:53

>>109595046

>Shit's made in rust, optimized for speed

>狗屎是用铁锈制成的,优化了速度

>Svelte

lol

#10 No.109595482
Anonymous · 2026-08-19 13:01

>>109595285

gave it a 10 seconds look

看了 10 秒

>uses "jpeg_decoder" for thumbnails when "image" with jpeg decoding is in the dependency tree

>当具有jpeg解码的“图像”位于依赖树中时,使用“jpeg_decoder”作为缩略图

>(picky) uses sync-only ureq in the client when async is already used in the workspace

>(picky) 当工作区中已使用异步时,在客户端中使用仅同步的 ureq

is this the power of trillion dollar intelligence? lol.

这就是万亿美元情报的力量吗?哈哈。

a funnier game for someone with more time, is to figure out what parts were "borrowed" from where.

对于有更多时间的人来说,一个更有趣的游戏是弄清楚哪些部分是从哪里“借来”的。

#11 No.109595582
Anonymous · 2026-08-19 13:18

>>109595448

It can filter through a 90k image library in seconds, and pull just as many hashes from a repo in minutes, that's good enough for me.

它可以在几秒钟内过滤 90k 图像库,并在几分钟内从存储库中提取同样多的哈希值,这对我来说已经足够了。

#12 No.109595815
Anonymous · 2026-08-19 13:55

>>109595582

i made that comment before seeing you posted the repo.

我在看到你发布回购协议之前就发表了这样的评论。

>>109595482 is may last comment.

>>109595482 可能是最后一条评论。

#13 No.109595968
Anonymous · 2026-08-19 14:18

>>109595448

Svelte is not what's doing the actual numbers crunching in this sort of programs, it's just gui library.

Svelte 并不是在此类程序中进行实际数字运算,它只是 gui 库。

>>109595582

>It can filter through a 90k image library in seconds

>它可以在几秒钟内过滤90k图像库

That's really bad to be honest. People use Hydrus for collections with tens of millions files and billions of mappings.

说实话,这真的很糟糕。人们使用 Hydrus 来收集包含数千万个文件和数十亿个映射的集合。

You gotta optimize it to under 100ms and hopefully it scales sub linearly with the database size.

您必须将其优化到 100 毫秒以下,并希望它能够随数据库大小呈亚线性扩展。

#14 No.109596150
Anonymous · 2026-08-19 14:41

>>109595968

I haven't tried going past 90k to benchmark it even further, though it would be worth trying to duplicate my library a few times to see where it might grind. Like I said, it started as a hobby project so currently I'm just doing whatever features I want to personally use. Next up would be a plugin system that lets you write scrapers/importers or accept tags from autotagging models.

我还没有尝试过超过 90k 来进一步对它进行基准测试,尽管值得尝试复制我的库几次,看看它可能会在哪里磨损。就像我说的,它最初是作为一个业余爱好项目,所以目前我只是做我想个人使用的任何功能。接下来是一个插件系统,可让您编写抓取器/导入器或接受来自自动标记模型的标记。

>>109595482

I mean there might be unused/legacy code in there or suboptimal patterns, but as far as I know I'm the only user and it's not bothering me, so it doesn't make sense to look into it atm.

我的意思是那里可能有未使用/遗留代码或次优模式,但据我所知,我是唯一的用户,它并没有打扰我,所以研究它没有意义。

#15 No.109596271
Anonymous · 2026-08-19 14:57

>>109596150

>though it would be worth trying to duplicate my library a few times to see where it might grind

>尽管值得尝试复制我的库几次,看看它可能会在哪里磨损

https://github.com/funmaker/hygen

Here is a simple image + tags(sidecar) generator I use to benchmark my project.

这是一个简单的图像+标签(sidecar)生成器,我用它来对我的项目进行基准测试。

For the performance tips, I personally do not know how hydrus actually does it, but it archives pretty good performance with sqlite alone. I am using postgresql and use intarray extension. Basically, I store sorted list of post ids for each tag and then do set operations on them. Sorted arrays are very fast for this, you can get unions, intersections and differences in O(n) with simple algorithms. If you use Rust, you might do all of that on the Rust side and even add some smart parallelism to really speed things up. But that if you ever feel like optimizing it all. Just be mindful that if your solution is worth shit, people with millions of images will flood in and constantly bitch about performance.

对于性能提示,我个人不知道 Hydrus 实际上是如何做到的,但它单独使用 sqlite 就可以实现相当好的性能。我正在使用 postgresql 并使用 intarray 扩展。基本上,我存储每个标签的帖子 id 的排序列表,然后对它们进行设置操作。排序数组对此非常快,您可以通过简单的算法在 O(n) 内获得并集、交集和差集。如果您使用 Rust,您可能会在 Rust 方面完成所有这些工作,甚至添加一些智能并行性来真正加快速度。但如果你想优化这一切的话。请注意,如果您的解决方案一文不值,那么拥有数百万张图像的人就会涌入并不断抱怨性能。

Good luck!

祝你好运!

#16 No.109596327
Anonymous · 2026-08-19 15:06

>>109596271

Thanks anon

谢谢anon

>Just be mindful that if your solution is worth shit, people with millions of images will flood in and constantly bitch about performance.

>请注意,如果您的解决方案一文不值,那么拥有数百万张图像的人就会涌入并不断抱怨性能。

Yeah, I get that, but I'm not intending to replace hydrus, in fact somewhere down the line I want the server to push tags with good enough sources (direct booru scrapes or manual tags) to the hydrus PTR, but once again that depends on demand

是的,我明白了,但我并不打算取代 Hydrus,事实上,我希望服务器将具有足够好的来源(直接 booru scrapes 或手动标签)的标签推送到 Hydrus PTR,但这又取决于需求

#17 No.109596880
Anonymous · 2026-08-19 16:21

>>109593708

BASED BASED BASED I always wanted something like this. Always felt getting used to hydrus was a waste of time because now everything is fragmented under a trillion tags and easy to lose track of compared to just using subfolders

基于基于基于我一直想要这样的东西。总是觉得习惯 Hydrus 是浪费时间,因为现在所有内容都被一万亿个标签碎片化,与仅使用子文件夹相比很容易丢失

#18 No.109596918
Anonymous · 2026-08-19 16:26

>>109596880

>because now everything is fragmented under a trillion tags and easy to lose track of compared to just using subfolders

>因为现在所有内容都分散在数万亿个标签下,并且与仅使用子文件夹相比很容易丢失

You can use tags like paths. Tags are more general structure.

您可以使用路径等标签。标签是更通用的结构。

#19 No.109597568
Anonymous · 2026-08-19 17:47

>>109596880

Naiad doesn't move anything, it just creates a list of hashes from your files, creates a thumbnail cache, then whatever tags you have are stored in a sqlite db.

Naiad 不会移动任何内容,它只是从您的文件创建一个哈希列表,创建一个缩略图缓存,然后您拥有的任何标签都存储在 sqlite 数据库中。

#20 No.109597673
Anonymous · 2026-08-19 18:00

No, thanks.

不,谢谢。

图片: https://i.4cdn.org/g/1787162437092222.png

#21 No.109597717
Anonymous · 2026-08-19 18:05

>>109593708

What are you using to tag the images?

你用什么来标记图像?

#22 No.109597796
Anonymous · 2026-08-19 18:14

>>109597717

Currently just manually, but you can pull tags from a hydrus-derived database sitting in a public repo. It doesn't have tags for everything, but at least half of everything I throw at it gets tagged pretty easily with just a download.

目前只能手动操作,但您可以从公共存储库中的 Hydrus 派生数据库中提取标签。它没有所有东西的标签,但至少有一半我扔给它的东西只需下载就能很容易地被标记。

← 上一篇返回列表下一篇 →

(如果你觉得这篇文章有启发,可以点击这里付费