EUREKAR
Eurekar·TOP
捕捉真实世界的英语信号
34信息来源
13,130精选文章
168单词卡片
18照片图片
全部8,372口语2,075外贸4免费1,128帖子4,467新闻1,602hackernews1,226techmeme755tmz663slashdot561techcrunch529arstechnica452随笔330外刊291Cards168simonwillison100sethgodin60图片18information14
← 返回

恐怖爬虫

2026-08-29 hackernews

← 上一篇返回列表下一篇 →

# Creepy Crawlies

# 恐怖爬虫

428 points · 202 comments | 作者: zdw | 原文: news.ycombinator.com

时间: 2026-08-29T17:49:01Z


### 热门评论

1. [Demiurge]

> I maintain a formerly popular gaming website, and it used to have hundreds of legitimate requests per second. The load would be especially high during popular event times. So, it’s always been running on a dedicated server.It also has an “online users” counter, which attempted to count real user sessions of unauthenticated user which still maintained a session, which lets them comment, or modify certain filter and display options. It never counted the Google bot.Over the last few years this counter went from 100-200 users online to thousands. I have been very hands-off with it for many years, doing minor upgrades and backups. However, the site also has gotten quite slow, these sessions were obviously impacting it. So, I finally investigated these crawlers, and yes, it turns out it’s an insane amount of traffic that is entirely artificial, the site has just a handful of real users, and thousands of these crawling sessions that actually try to do everything they can, click every button. It doesn’t help that sort and search were implemented using GET links.I fixed the counter to exclude the crawlers, but I have a bit of a dilemma. I don’t want to stop the bots from updating their knowledge based on all the content.The best solution I could find is the new CloudFlare feature where they might charge the crawlers for every request, or otherwise block them. I think that’s a fantastic idea for the internet, at large. I signed up for the beta access, but haven’t heard from them again. I do think it’s unfortunate that this requires CloudFlare and the middleman.Overall, it seems like the LLM are really straining the internet economy, the openness of it. Email spam used to be the worst, but the organized trillionaire labs sucking up the entire internet is going to break something if we don’t preempt them better.It’s too bad the copyright and public internet systems are not acting quick enough. And I think there is no reason to act like this race really has to be at such a break

⋯ 继续阅读请开通会员 ⋯

1测一测:这篇你记住了吗?只用 20 秒,带走一个线索。不确定也没关系。

下面哪项最准确概括这篇文章要带走的关键信号?

不想现在做也没关系,继续往下看即可。

🔒

MEMBERS ONLY

这篇是会员专享内容,你看到的是预览段。

会员每天解锁 6000+ 篇真实英语素材——双语科技、口语、外刊、单卡,不设上限。

年会员 ¥365 —— 一天一块钱,续费一直能用。

了解会员 →

已是会员?点此登录解锁全文。

学习留言 0 条

记录你在这篇文章里学到的表达、疑问或感受。

登录后记录你的学习留言 →

还没有学习留言。写下第一个收获吧。

← 上一篇返回列表下一篇 →