I think people are missing one of the points of the article. It's not just that agents are hammering the site, it's that there might be lurking vulnerabilities that allow malicious usage, which is why it went down, then came back up with a fraction of the data and a reduced design.
The article says (speculates?) that malicious users are trying to get privileged access for an edge in prediction market betting. From the article:
> If you could see The Numbers data before everyone else, every single week, you would have a significant edge over all the other traders - learning the answers slightly ahead of publication would allow you to front-run the trades.
If it's the case that prediction market traders were trying to front-run the market outcomes, the prediction markets should put the fix on their end by locking down trading before a market closes with some time buffer to prevent this.
That's assuming the prediction markets all play fair and by the same rules. If Polymarket were to do that, its competition could choose not to. Since these exchanges are not tied to any one country (or jurisdiction), its users would jump ship. Because the users don't care.
See also stock market trading, where companies would go to great lengths to shave nanoseconds off of getting information and putting in trades. Now apply this min/maxing to worldwide, unregulated and anonymous.
"One Reddit theory even suggested it was a deliberate rug pull designed to cripple the free site to push people towards paid products."
I know this isn't really the point of the article, but I've been thinking about this a lot. I wonder if we're going to see more resources go this way.
I used to publish little doo-dads as opensource software. Not because it was something that was legitimately ground breaking or anything (it absolutely wasn't) but because I had a problem, and I thought "heh wouldn't it be cool if someone else had a similar problem and could use my resource for it."
But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
And I can see why maintainers of sites like this, or other free but incredibly useful resources might start to get irate at that.
This attitude makes zero sense to me. You benefit from the "training data" just like everyone else does. If you don't, that's a problem with you, not a problem with AI models.
AI solves exactly the meta-problem you describe: "I had a problem and needed to write a one-off doo-dad utility program to solve it." Now you can do something with your time besides writing pointless one-off doo-dads.
As for monetizing the training data, (a) it cost hundreds of millions of dollars to generate the weights, so why begrudge the companies that made the investment and did the research necessary to make it happen?; and (b) rest assured, whatever your doo-dad does, an open-weight model like GLM 5.2 can generate it for free using your own hardware.
So you don't have to pay anyone in that case. Well, except nVidia, I guess. Point granted there.
I don't know how else to explain it. It doesn't feel the same to have my contributions smushed into a linear-algebra machine? For me, the incentive was the idea that my contributions might have directly helped someone. That absolutely does not feel the same when I think "My answer has modified the back propagation of a Machine Learning training run."
It doesn't have to make sense to you - I just believe that I'm not exactly alone in this thought.
Now apply this to Art, free stories, writing etc. It feels bad to have your free contributions hoovered up and monetized. It doesn't feel particularly fair or ethical to me. And it'd make me double think before making something free and publicly available.
Understood, there's certainly nothing invalid about your point of view here. It's just not one that I personally can come to terms with.
I've spent a lot of time in your shoes, wasting time on busy-work needed to accomplish a larger goal (and absolutely sharing the results freely, over multiple decades)... and I don't miss that part of it one bit.
The godlike tools are paid subscriptions at best and restricted to a handful of corpos by the government at best. The free scraps thrown to the proles are mostly novelty. The other day I asked 5.6 Sol to identify a push prop airplane and it said “this is a helicopter.”
> Now you can do something with your time besides writing pointless one-off doo-dads
Writing software one-offs to scratch an itch was historically one of my most enjoyable past-times. AI trivializing that has been a very real theft of joy in my life.
Solving the actual problem was never the point, it was just motivation to do geek-out and craft some code.
AI is rapidly diminishing many interesting hobbies (coding, art, music, writing).
Having more free time when there's nothing fun nor exciting to do with it isn't really a benefit.
It does make sense. When podgietaru wrote their doo-dad programs, I guess they made their work available for humans to use and they didn't mind giving their work away for free. Some other person might say "I'm giving my work away free, I don't care if humans use it or AI, for any purpose". Both are legitimate choices, each according to their own.
As for "everyone benefiting from training data" - sure, but the AI companies are not investing Billions of dollars out of the goodness of their hearts, it is towards one and only singular goal of making profits (at some point). People might be sympathetic to these AI companies if they at least behave decently - they take everyone's work (text, software, fiction, music, images, videos...) without paying a penny to anyone. If they take everyone's work for free, they should give away anything that is built on that work also for free. This is before we even get to environment, privacy, hammering sites by not respecting robots.txt etc issues.
If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no? Even if I spent my own money making the meal, it was made from stolen raw material...
If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no?
Yes, and that's exactly my point. We are in violent agreement. It's shared with me, with you, with OP, and with everybody else. We can get a delicious meal for free or we can pay somebody else to serve us a slightly-tastier version.
Here, disregarding copyright law has fulfilled the very purpose of copyright law: to advance the useful arts and sciences. Copyright law was the best tool we had to accomplish that before, and now we have something better.
Sure they are. Surf around on HuggingFace and you'll see dozens of open-weight models up to 120B parameters in size from for-profit US companies, freely downloadable (if not freely runnable, alas). Probably hundreds of them, at this point.
And a 3T parameter model is scheduled to be dropped by the Chinese on Monday.
They should certainly publish more, and if somebody were to argue that model weights trained by scraping copyrighted data should inherently be accessible to everyone, I'd be 100% in favor of that.
But the McDouble isn't an adequate comparison. Yes, the companies who opened Michelin-rated restaurants with the food they stole from your garden are giving you McDonalds'-level food for free and charging for the rest. Meanwhile, some other companies who raided your garden are giving you the plant-based equivalent of Ruth's Chris or Fogo de Chão for free.
The other thing is, after they stole all that stuff from your garden, it was somehow still there. Your neighbors on Hacker News say that some bandits raided your garden, but you can plainly see that no one has picked any fruit or uprooted any plants, and your security cameras reveal nothing more rapacious than a rabbit or two. You begin to suspect that your neighbors are gaslighting you.
The plant metaphor breaks down when you pause, take a breath, put the McDouble down and realize that intellectual property is a completely different ownership concept than physical property.
>As for "everyone benefiting from training data" - sure, but the AI companies are not investing Billions of dollars out of the goodness of their hearts, it is towards one and only singular goal of making profits (at some point).
Its not guaranteed they will ever get there, every day it seems increasingly likely that without massive government intervention we are just waiting for local open weights models to become popular.
Meanwhile my ISP gatekeeps the same free content behind a service payment. Their motive isnt free love and world peace, its also profit.
>If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no?
If you cloned all the veggies in my garden you are free to clone them further and fill your belly you owe me nothing.
>Now you can do something with your time besides writing pointless one-off doo-dads.
This is such a disgusting statement that goes against the very core of what open-source stands for. If even only one person got some use out of what they made, it is not pointless by definition.
I used to rank pretty well trashing crappy credit cards and encouraging people to switch to better options, then Google decided that 10 results for the card issuers website was better.
Same shit when I manually wrote proto-gethuman posts on calling telecoms/banks/etc (and also pushed visitors to try an Indy ISP or credit unions), then Google felt it was better to drive users to the telecom’s website that wants you to do anything but call them.
Do you have any active sites or social media with your recommendations?
(I never publish my recommendations because they seem to complicated for people. Like using 30 GB plan/$10/monthly from T-Mobile for data and then using Tello for $10 plan for voice/text. This would require a 2 esim/sim card phone. I am currently using a Moto G Power 2024 phone from eBay $90/new, which was better than the $200 slightly used Pixel 6a from Swappa. Most people would just save the hassle, get a Galaxy phone with a phone contract.)
Nothing active anymore, all converted to static sites.
If you’re on page 2 of the results, you effectively don’t exist so I stopped bothering/benefiting from display ads.
Coincidentally, I did try to get deeper into the cellular service side (it’s another high margin and high customer value segment), but I did better on the finance side.
> But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.
> And I can see why maintainers of sites like this, or other free but incredibly useful resources might start to get irate at that.
Those who object to the scraping fall into several camps, but the biggest complaint I am hearing is that it increases both maintenance costs and time. In other words: it sucks when people are using your work in a manner that you find offensive, but it goes beyond that by doing actual harm.
> The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.
People doing it with a couple of machines and residential proxies versus Anthropic doing it with 2 data centers worth of machines.
> The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.
> In other words: it sucks when people are using your work in a manner that you find offensive
That to me still reeks of Dog in the Manger mentality. If you publish something for the world to use, you should neither care nor even track, much less discriminate by (or suddenly seek compensation for) who is using it.
Agreed, this is what you agreed to (implicitly or explicitly) when uploading stuff to the internet. In fact, open source embraces this (hence the 'open').
But people are free to not publish things or post things online with a more restrictive license. Not that a license stops things from being indexed.
> But people are free to not publish things or post things online with a more restrictive license. Not that a license stops things from being indexed.
Right. In fact, people are also free to publish things with licenses that condition access on compensating the author/publisher, and they have both social and legal backing to enforce it. This is called "proprietary", and it's not a wrong choice - in fact outside of software, it's the default choice.
The problem is when people publish "free" and "open" as a marketing tactic, where in fact they really want to control and charge for access (whether dollars or karma or credit). That is just plain dishonesty.
> Agreed, this is what you agreed to (implicitly or explicitly) when uploading stuff to the internet.
A more pragmatic approach is to acknowledge that concerning one's self with how something is used once it has been released is an emotional drain. It is a bit much to suggest that someone agreed to something, even if that agreement is implicit, just because they released it.
> But people are free to not publish things or post things online with a more restrictive license.
Licences are meaningless unless you have the ability to enforce them (e.g. to sue). That's why so many companies are willing to ignore the terms of open source licenses. It's also why the attempts of enforcement that we do hear about are usually backed by a third party, rather than being done by the software developer themselves. Simply put, the individual developer (or even small project) trying to make a contribution to the community would be better served by not publishing (instead of using a restrictive license) if they are concerned about how their work is used.
I don't even know if there is a good way to resolve the problem. Consider something like a DMCA Takedown notice. It removes the administrative and legal overhead to copyright infringement, yet it is also easy to abuse. For example: businesses have weaponized it by using it against individuals. Perhaps my cynicism is taking over here, but I suspect any easily accessible mechanism for enforcement would be similarly abused.
This is absolutely fine, until some shitheel decides they're going to continually scrape your website, thousands of requests per second, and knock it offline constantly, so they don't have to cache even a single thing, or write better scrapers. They'll just suck up your bandwidth for it. After all, it's you that's paying, not them.
Search engines were absolutely fine. They were respectful of the sources they scraped. There's now a bunch of money-obsessed shitheels doing whatever they can, as blatantly and as lazily as they can, because they believe they'll become rich. Every one of these fuckers needs to be dead or in jail before the world becomes whole again.
If nobody takes action against these fuckers, everything you care about will just nope out of existence, and these fuckers will be all you have left.
Major AI companies aren't indiscriminately scrapping any more than search engines were (which themselves caused some problems in the past too, until that got sorted out). Your gripe is with SEO people and content marketers.
I'd be more sympathetic if it was really about this, though. However, the majority of the complaints seems to come from people personally offended by the possibility of their content, which they claimed was published for everyone to benefit, ending up as training data, and thus actually benefiting humanity, at scale far beyond the original publication ever would. Which is the Dog in the Manger attitude I point out.
The problem is the rat race for the pot of gold at the end of the rainbow.
A similar thing happened when crypto was in ascendence. Web pages were getting crypto-miners injected into them. Celebrities were shilling NFTs. Everyone and their dog was on the make. The thought of riches broke the minds of millions.
The same is happening with AI. Whether it's the "major AI companies" or millions of self-interested also-rans with fewer moral scruples scraping, the problem still exists. The problem will continue to exist until you can't conceiveably make money by scraping like a bastard. If everybody identified themselves up front and respected robots.txt, there would not be a problem. But they don't, and they don't, and they pummel websites for no fucking reason, and they don't care, and they won't stop.
Websites get hammered by millions of unique IP addresses from residential ISPs which happen to belong to botnets, none of them identifying themselves as a bot user agent, all just pretending to be some slightly out-of-date version of Chrome.
And it's not like the "major AI companies" hands are clean. Who bought up all the RAM and compute power? Who's building datacenters in the third-world parts of the USA and powering 24/7 them with diesel generators, or sucking up most of the power generator, and of course using up all the drinking water to cool their machines?
Everything in humanity and nature is there for the AI bros to use and dispose of, provided they come out on top. This is the attitude that needs to be taken down.
You may be right about the scale of also-ran operations, even though I disagree with the core comparison. Unlike crypto coins, which require superlinearly growing energy waste just to sustain their basic guarantees, and were created to solve "problems" that aren't problems and don't need solving (hint: trust is a feature, not a bug), AI actually works. It delivers real value for cheap. The growth isn't artificial, it's a product that doesn't even need much marketing[0] - it exploded organically since ChatGPT, since it's obviously that immediately useful for approximately everyone in some aspects of their work or life.
One nit though:
> And it's not like the "major AI companies" hands are clean. Who bought up all the RAM and compute power? Who's building datacenters in the third-world parts of the USA and powering 24/7 them with diesel generators, or sucking up most of the power generator, and of course using up all the drinking water to cool their machines?
They paid for it, much like everyone else.
> Everything in humanity and nature is there for the AI bros to use and dispose of, provided they come out on top. This is the attitude that needs to be taken down.
Here you're arguing against the basic market economy. They aren't using and disposing of anything they couldn't buy for that purpose like literally everyone else. There's no theft or trickery going on here. There's a boom, because AI is that useful, but it's still all normal resource allocation.
--
[0] - Of course the competing players invest tons in marketing to gain an edge against the other players.
> But now I'm really reluctant to give more stuff to the free web.
I feel the same way too. But guess what, all the code you did not publish gets into the training corpus anyway (when you gave Codex or whatever full read access to your filesystem).
The vast majority of people are in fact this stupid.
I recently had trouble connecting to a new AP on my laptop. After a quarter hour of frustration, I connected to the AP of my phone, asked Claude Code what the problem is, and she found the issue in seconds. I didn't allow her to make the actual changes, but she did have read access to everything, and helped me considerably.
So network-manager gets a new bug report about too-long non-ASCII AP names, I get online, and I don't know maybe Anthropic sneakily learned something from my local python projects. I am one of those vast-majority stupid people.
> Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
I suspect the next AGPL (if not normal GPL) will explicitly prohibit training of closed models against the source material. Ideally it will prohibit training of open weight models too. Either share the entire process, or go piss up a rope.
>But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
I dont get this.
You provided something for free to help people, but dont want to do that anymore because it might go into training data and help many many more people?
So far LLMs have been loan funded donations of loss leading services. They might never actually make their first dollar of profit.
Meanwhile ISPs the world over have been monetizing access to your content.
I feel like what you mean is that you want control over attribution.
I think the sheer magnitude of the economics have made the scales fall from a lot of people's eyes. For decades people put stuff on the internet for free on the assumption it was "not worth" anything. It turns out that as soon as that commons can be enclosed, we can marshal hundreds of dollars for every single living human, to pay for this commons to be repackaged. The money is there, and we're happy to spend it, we just won't spend it on you.
So basically the “information is free, encyclopedias are expensive” phenomenon from the pre-internet days? Collection, collation, and distribution have more and different qualitative value than the sum of the individual bits.
I also recently made an open-source project with 200 GitHub stars private. I never had a problem with others using it as the basis for their own projects. In fact, that happened, and I received credit for it. But LLMs just hoover up everything, process it, and then spit it back out as if it were their own.
Why would you offer data for free if people are going to pay someone else to access it?
The whole internet is about to experience USA-level greed because people are too dumb not to use AI. Just like Google killed small sites because people were too lazy to look at page 2.
A couple years ago, I was running a website which allowed the public to view all the US government handouts to small businesses during the COVID-19 pandemic. It also tracked all the fraudulent loans being prosecuted by the DOJ, and allowed anyone to run structured queries over the public dataset. There was a "donate" button which took in ~$2k in donations over the lifetime of the site, and you could download the entire underlying dataset (around 10 GB uncompressed) for free directly on the site.
Despite the "download all data" link being prominently placed on the front page, the AI scrapers decided it would be more efficient to download terabytes upon terabytes of raw HTML by paginating through every possible facet on the search endpoint. Even with CloudFront caching results and a fairly efficient backend setup, the monthly bill ended up with around $1k just going toward network ingress/egress, so I shut down the site the following month.
Completely agree, I've worked at several companies where I saw the AWS bill balloon from five to six figures, usually because of pointless over-provisioning. However, it all became worth it when I saw Jeff Bezos go to space. [1]
Because the crawlers would still have hammered their site, though.
(The GP post doesn’t actually meaningfully address the issue being raised. Adding BigQuery or whatever would not change the fact that (a) they already offered a method of getting all of the data in a cost effective way, and (b) the issue was that the crawlers hammered the site hard enough to make it economically unviable.)
The comment I’m speaking of was from a different user, and it’s giving a strong smell that both users are bots unfortunately. Maybe I’m wrong, maybe it’s a coincidence those exact strings of characters were formed in isolation from each other in the same thread. I hope I’m wrong, but it’s a red flag
I'm not a bot, and the other commenter was restating the question as the solution provided did not seem to actually solve the problem, just redesign the entire project (See comment from u/taneq).
We didn't form "those exact strings of characters in isolation from each other in the same thread" they are directly referencing my comment in a reply thread to said comment.
Sure, but now you're moving the site from "Free data presented in a pleasant way to view" to a "pay-as-you-go database". Your audience shifts dramatically, and you lose the ability to share the data you're trying to present.
Yes, if I wanted to put a nice HTML interface over the query results then I would still end up with the same problem where some combinatoric explosion of query parameters to the `/search` endpoint, most of which are cache misses, leads to many many TBs of network egress.
not the OP but I'd say that google crawls you once and AIs scrape your page every time someone asks them a question that they think your page might be relevant to.
There's been an explosion of vibe-coded scrapers that behave poorly and ignore robots.txt. Presumably OP disallowed crawling of the search endpoint for the reason stated (that crawling it would result in endless permutations of search filters.)
They were both more exhaustive and more frequent. I don't remember the exact numbers, but it was definitely over 100x the traffic from search engines. By the way, search engines were allowed under robots.txt - I did want all 11.5 million loans to be individually indexed so they would pop up in Google search results, and I actually did receive/forward multiple tips about fraudulent loans because someone searched a business name on Google. All of this traffic was barely a blip, and my AWS bill for my hobby data science account was only ~$50-$100/month.
When I looked at the logs after getting the billing alert, 99.99% of the requests were to the "/search" endpoint with virtually every permutation of ~10-12 facets in the query parameters. There was only one scraper, but it triggered an enormous amount of network egress since it ended up missing the cache on the majority of queries.
I deployed a Lambda function behind CloudFront which rendered a simple HTML page with the query results from executing some SQL over the dataset. I served millions of page views for next to nothing because most pages were already in the cache.
I don't know where you get this expectation that people should anticipate that a crawler might try every possible combination of query parameters, thereby missing the cache on each one. Most people consider it a bitter and arrogant perspective, which is why this got downvoted.
the AI scrapers decided it would be more efficient to download terabytes upon terabytes of raw HTML by paginating through every possible facet on the search endpoint.
No, they decided that would be a great way to convince you of the narrative and persuade you to pay for "security" services that further the incumbent browser monopoly.
They're not "AI scrapers", they're DDoS'ers manufacturing consent.
One thing I’m wondering is whether we need an open source set of technical patterns and libraries for dealing with this changing traffic mix.
Millions of small sites and creators don’t have the ability to design their own protections against large scale automated access. If useful content now attracts aggressive crawler traffic, many sites will be too expensive or unreliable to run.
Is part of the answer a community response? Perhaps a community-maintained toolkit, based on traffic data, that host sites can apply? It could include standard agent identification, rate limiting, traffic classification, access policies, caching, challenge mechanisms, logging, attribution and usage control etc etc.
In effect, we need much stronger road rules for today’s automated traffic, available as open technical patterns and libraries rather than every site owner having to invent this alone (they won’t).
The community response was respecting robots.txt, but since that was more of a gentleman's agreement than a legal requirement, hungry AI scrapers disregarded it.
Which brings us to the old fashioned flood control mechanisms. That is, the toolkit you propose already exists and has been used for decades in various iterations to protect against various forms of attack (slashdotting, DDOS attacks, overzealous search engines, and now AI scrapers).
Have you looked into those before? Companies like Cloudflare have been at the forefront of this field for a long time now.
Yes, a community effort against the botnets would be great.
The "good bots" which identify themselves and follow the rules are easy enough to block, so not a problem. It is the "bad bots" which pretend to be real users and hide behind residential proxies, and so are almost impossible to block at the moment, which are the problem.
Given that there are big companies openly (i.e. on the clearweb, not even darkweb) selling access to these residential proxy botnets of compromised smart TVs[0] and mobile phones and other devices, can we not simply get a database of the these IPs and block access from them? I would venture that almost every single one of those residential users are unaware that they have devices in their homes which have been compromised and are being abused in this way, so if they were to start seeing messages from more and more sites along the lines of "Access to this site has been blocked because unusual traffic has been detected from your computer network. Please check all devices on your network and remove any malware which may be routing this traffic." then maybe we could start addressing the problem at the source.
> The "good bots" which identify themselves and follow the rules are easy enough to block, so not a problem.
And then they get blocked, which is a problem. In particular, the modern agentic AI tools interacting with web services to fulfill user queries - they are acting as user agents, and they should not be discriminated against.
So I'd say the first pattern that needs to be broadly adopted is non-discrimination of user agents.
But of course we've tried that in the past, the whole problem is that non-browser user agents == end-user automation, which is anathema to pretty much every on-line business out there, as money made online is primarily conditioned on users wasting their own lives on interacting with services directly.
I don't think there is a technocratic solution to this problem. Very rich people are using their money to DDoS the internet, for no good reason. We just need to identify these people and fine or imprison them until they stop doing it. Residential proxy providers would be a good start.
To a point; as the post mentions, they added instructions for LLMs to their website and increased their licensing inquiries tenfold (the article didn't state anything about actual licensing payments being made though).
I don't think there would be any issues per se if the scrapers just paid licensing fees to get the good / complete data. But the issue was that they started to try and find exploits to get to data earlier.
At the risk of oversimplifying things from a distance, this site -- especially the free, public-facing part of it -- seems like it would be an ideal candidate for a rewrite using static site generator/framework. That, coupled with a bot-aware CDN should keep them online at a reasonable cost for many years to come.
Otherwise, I'm very curious to know more about their old and new architecture and what sorts of mitigation/scaling strategies they've started using to keep the site online.
Yeah this seems like one of the easiest access models to create an excellent security model for -- basically just static publishing. Sounds like they were in need of a rewrite anyway!
It probably seems daunting but to be honest this feels like a weekend's work at this point with LLM assistance. Not to be glib!
But it also destroys the business model behind the site.
Sure, you could technically redesign to handle the bot traffic, but if the bot traffic is just taking the data and reducing any need for humans to visit the site, why is he putting in the effort to maintain the site?
Well, why is he? If it's for ad revenue, this problem is surely global to the web. If it's because he likes to do it, it shouldn't matter if one AI or a trillion scrape the site. I feel like we badly need a rethink of the web architecture after DoubleClick anyway. Maybe the name of the game should be to cut down a site's assets very hard and use static hosting for them. This interferes with crummy sites that show a mess of inlined ads every refresh, but now that there are a billion 'poor users' this is no longer feasible.
Maybe more an inherent problem with these chatbots.
LLMs are great, but they aren't producing new information. You still need people for that. But if you cut down any incentive for the people to do that, the LLMs will starve.
Because... ? If I'm operating a site, and I want $X to allow a bot to scrape my site, why shouldn't the bot be allowed to make the purchasing decision to scrape the site? Obviously the bot owner would have their own set of guardrails, but if it allows sites like The numbers to stay up because bots aren't going to look at ads so the old model of displaying ads isn't going to work, I have a hard time seeing that as a bad thing.
According to the post inquiries for purchasing the data increased as a result though so they might not make the decision, but they do make the suggestion.
Ya this problem was solved 20 years ago with Varnish cache and Coral cache/CDN.
I think what's really going on is that bots expose how underpowered web servers has gotten in recent years. In the 2000s, even poorly-architected PHP sites tended to serve about 200 requests per second, with 1000+ being common for static sites. I remember when Node.js came out and claimed that it could serve more like 100,000 RPS due to its cooperative threading model. But today sites have a remarkable slowness to them, running many hundreds or thousands of database queries due to ORMs and N+1 problems, so that response times can be 500 ms or more and even 1000 simultaneous users stresses servers.
What went wrong is that nobody solved stuff like Russian doll caching in a general way separate from the programming language and database. We should have had ways to make dependency graphs using Etag headers as keys with real cache invalidation of dependent data. I went down that rabbit hole 10 years ago using touch events in Laravel with callbacks to handle cache invalidation when class model data was saved to the database. Also a query cache using Redis which I think might have been handled better at the database level anyway. After that experience, I can honestly say that cache invalidation is so difficult to get right that it's effectively an open problem. Meaning that programmings should use a package instead of rolling it by hand, and it should be a major concern from the start (along with sharding by user id or using something like Firebase).
Don't get me started on how the web should have been a P2P content-addressable memory anyway. Nearly everything should be available from a nearby edge peer, similarly to BitTorrent. But nobody bothered to solve how to make that work with HTTPS/SSL. I suspect that has to do with early flaws in the browser security model where the whole page has to be behind HTTPS or warnings appear. So it was never clear what was personally identifiable information (PII) or merely public data being served over HTTPS. To really solve that, we probably need real trust networks and maybe even zero-knowledge proofs.
Since these problems are so challenging to fix, and big companies can't be bothered to do it since they pulled the ladder up behind them, we're probably stuck with banal "are you human" challenge screens for the foreseeable future.
They are indeed quite challenging. We literally wrote a paper about it (https://arxiv.org/abs/2605.09114) ... you almost describe some of our original architecture too! You might find it interesting -- or not. https://getswytch.com
> Ya this problem was solved 20 years ago with Varnish cache and Coral cache/CDN.
It was... if you are paying datacenter rates for the bandwidth
If you're paying cloud provider per GB pricing, nah, even text will add up if you happen to be targeted by a bunch of bots
> What went wrong is that nobody solved stuff like Russian doll caching in a general way separate from the programming language and database. We should have had ways to make dependency graphs using Etag headers as keys with real cache invalidation of dependent data.
we did that with nested ESI includes in Varnish so every "box" of content on the page was cached separately + some piping for invalidation, so if a given piece in the database was changed it sent invalidation to all nodes. There was also some grace so if the thing you wanted got updated RIGHT NOW you might get stale version while the new one is updated in background, and don't pay the latency cost
> I think what's really going on is that bots expose how underpowered web servers has gotten in recent years
More often than not they're running on a VPS, and cloud providers have been pushing the envelope of what a vCPU is. Amazon still bills hyperthreads as a single core!
Add to that their servers are often aging and you get a recipe for slow web services
What a throwback. Back in 2015 I have started Applaudience, which at the time was the only provider of real-time cinema ticket sales data. I still remember comparing our numbers against TheNumbers.com as part of calibration. I have since moved on to other businesses, but this remains one of my favorite pieces of technology that I have developed. Would love to bring it back one day.
90% of bot traffic on my network of websites is through headless Chrome, via residential bots nowadays. Impossible to block. Not even for Google, as they inflate my Adsense numbers as well.
Claude Bot still (at least last month, and it's been doing it for over a year now) seems to have a bug when traversing (at least my sites), wherein it drops the trailing slash of a directory (which is present in the a href tag), then makes the request to the subdirectory without the slash I put in the link, then Caddy automatically responds to that (via the built-in File handler) with a redirect telling it to add the trailing slash, and Claude Bot then makes the request again with the trailing slash that should have been there in the first place.
So most subdirectory URLs get two requests from Claude bot, the first one needless because that wasn't the URL in the tag.
This only works for crawlers that play fair, unfortunately there's many that don't. Including ones that use consumer devices like TVs [0] and mobile phone apps that offer incentives to consumers (if they're even open about it) to use their internet connection to e.g. crawl websites. There's probably an army of hacked toothbrushes and the like doing the same thing.
Thanks God i can serve up to a terabyte per month of traffic easily, otherwise it would be a catastrophe. I am not against bot scraping, but I worry about stability for meat visitors.
So, I think of enabling payments for website visits cloudflare recently developed.
I also have the problem of old technology like the mentioned site. While my personal blog uses static generator, archive website uses ancient Drupal version, which has no security patches for many years already.
Genuine question: do prediction markets open up a new revenue stream for TheNumbers.com?
Specifically: they suspect that the motivation for trying to hack their site (for at least some people) is that they wanted early access to numbers that folks were betting on in prediction markets. Since they've got those numbers they can just bet on them, then benefit from their perfect knowledge.
On the one hand this does seem incredibly unethical (it's clearly insider trading).
On the other hand the CEO of PolyMarket has said that insider trading is part of the point of PolyMarket: https://youtu.be/ZN4njIQcSR4?si=ztyTtgjeHSJbNjSZ&t=1566
There’s an easy solution: Bruce should place large bets on the prediction markets right before publishing the relevant data. It is legal, ethical, and, for him, risk free.
It feels like there is fundamentally missing infrastructure here that is needed to make these problems go away.
Basically bots need to be (somehow) paying for the traffic they create, or prevented from creating it, or told to go away and then fined if they violate the request. No idea how to do these or even at what level in the stack they should happen, but they need to happen eventually somehow.
Every website I visit could get a fraction of a cent in my Cloudflare Wallet. A human browsing incurs a few dollars a month. Plus, any website I access in this way decides not to show me ads either and just charges me the cost of serving the data plus a nominal profit margin.
Disclosure: I am a shareholder and would love for them to solve the AI bot problem and the ad problem like this.
we've been working on basically the same problem in the email space for a few decades (legit email vs. mass spam). its a very hard (i think impossible) problem.
Well, it is technologically not that hard to imagine a solution---the problem is social, getting everyone to agree on how to do it, figure out the policy around it, etc. And the situation is getting so bad that maybe it is time for someone (maybe someone reading this thread) to figure it out.
I set Cloudflare up a couple of months ago specifically to block bot traffic. It didn't do anything for me. Dumb bots were still hitting every special link on my wiki fast enough that the server was continually swamped running Lua scripts. 65% of the traffic for my English-language site was coming from Vietnam. But I didn't want to block Vietnam altogether, because my hobby site has genuine users from there too.
I eventually just shut down my mediawiki instance. I couldn't find a way to keep it online and still run on an affordable VPS.
I think I probably use AI like a lot of consumers out there. Search engines have gotten bad and AI really good at answering fairly specific questions. Often pointing at sites like wikipedia. Pretty clear changing my behavior will have zero impact but it is certainly part of the problem. Feels a lot like my CO2 consumption.
Only solution I see is to ban user that abuse the system. Too much requests, too much bandwidth and you get into the blacklist and are cutoff from the human side of the internet.
With the amount of money AI labs are burning, someone should just set up some infra to host the data they want and charge for access to it, instead of abusing the goodwill of legitimate pages that offer it for free.
The Numbers is such a brilliant site, this explains why it came back in such a stripped down form. AI and prediction markets giving me more reason to hate them.
If the AI companies destroy the open web, eventually they'll need to start curating knowledge sources just like netflix makes movies and amazon has physical stores...
Already happening. Frontier LLM vendors have been hiring human domain experts specifically to create their own proprietary training data in targeted verticals.
The real story here is that prediction markets were banned for a reason and loosening the rules is causing chaos just as was expected. AI plays little role in this story, hacking by humans would also be motivated by financial returns, unless the element is that AI hacking is cheaper and the returns are not so big.
You missed the speculated motivation: unrestricted prediction markets, which are an open and broad incentive to do whatever actions might provide a slight edge in betting.
Webscraping is among the more benign things that gambling-on-anything can drive. And even that has a negative impact, as seen here.
John Brunner and Alvin Toffler both made stark warnings wrapped in futuristic giddiness about things like it. People like Fuller no doubt thought polling on a large scale was terrific. There were old usenet groups and BBS subs (minus the money aspect) experimenting with the model. I do not believe they ever are or were good in a largescale model (money or not).
I mean wisdom of crowds, super-forecasters, calibration and pre-registration are useful tools that can result in better predictions, turning it into online gambling is where it went sideways.
Mini-rant: Promoters claim that the system serves a public good by helping society discover/converge on useful truths sooner.
Even where the question/event is of public interest, the opposite happens instead. People with expertise or non-public information are incentivized to misdirect and delay as long as possible, as that maximizes what they can make from betting.
Who are running these bots? I presume developers at all of the frontier labs know (or at least would know to look for) Wikipedia has bulk APIs for automated access. Unnecessary scraping increases their workload/costs too, so why in 2026 is this still a problem?
Black market and gray market data. All the firms want data. All the other firms want data. The banks want data. The other criminals also want data for their crimes and schemes. Oh insurance companies, and the ATS systems. Everybody wants as much data as they can get and they don't care how they get it.
> Read the Docs, a non-profit that hosts documentation for open-source software, who watched a single crawler download 73 terabytes of zipped HTML in one month, costing it over $5,000 in bandwidth
From the article.
Not the same site, but an example of the same issue.
All new content will be behind paywalls. This will be the new way the Internet works and I've said this for the last several years. There's no value in writing unique content just to have AI steal it and distribute it without you getting any clicks. The only way it works if it there's a licensing agreement and if they don't want to pay, then who cares, you weren't going to make any money from them anyway.
Why would you produce content and have no one read it or visits your site, but OpenAI and Anthropic make millions from it? At some point it becomes stupid to just give free money to these companies when they steal literally all your traffic and content. As per the article, Anthropic sends 1 page view for every 38,000 views they get.
The "old web" hasn't existed in 30 years now. I was a part of the old web and if people in the early 90s knew that someone was profiting off their work it would have killed it immediately.
I always liked this site, but reading this and seeing the anger about expecting the site maintainer to do things for you is repulsive. Frankly, if he wanted to pull his site down with no notice that is perfectly within his right. It was/is his site. He doesn't owe anyone a .tar.gz either. His work.
I agree. I've seen it happen before on a project I use. I decided to take a look at the repo for one of the plugins, and I saw a heinous issue that basically was TELLING (not even asking) the maintainer to fix it.
Like dominoes, as soon as it is accepted in a few places, people think it is acceptable to push to 'share'. It's almost terroristic sometimes, the pressure some maintainers are under.
This is another reason we can't have nice things. The level of entitlement required to send an angry message to someone complaining about their free resource being offline is mind-boggling, though.
Every example of this confirms my view that we should treat the spread of these AI scrapers like we treat the proliferation of drugs. We should be seeking to bust AI rings like we seek to bust drug cartels.
> The world we have built thus far is so incredibly ill-prepared for the power and scale of the AI models we all have access to.
Exactly, so use the AI to secure your servers. Ask the AI to audit your site for any security holes. If unable to rewrite, at least harden the existing code. AI’s are really really cheap (and fast) security consultants now.
This is an odd example to present this argument through.
I'm not defending AI bots overwhelming websites or hackers motivated by Polymarket, etc., but I don't really believe a 30 year old website with "approximately 160,000 source files serving around 2 million pages" is a good litmus test for the state of the online world. What's worse is the absurdity that this basic site offering niche data would be targeted because Polymarket depends on it for some of their bets, something that most websites don't have to deal with. Frankly, that seems like a much more interesting angle to explore than "this old website that should have been re-written multiple times over the last three decades now doesn't have a choice but to be re-written."
Oddly, it seems like the solution, at least in this specific case, is relatively simple (and ironic): pay for an LLM to re-write this site in a modern, scalable, secure way and have that LLM monitor the site to ensure it is behaving correctly. Put this thing in a modern platform behind a proper cache (i.e., throw the whole thing up in Cloudflare) and many of these problems just don't exist anymore. If this site is as basic as it seems, this could be a weekend project.
It's certainly an interesting story, and I don't blame the owner of The Numbers for handling his site the way he has, but framing this as "AI is destroying our beloved internet" just seems obtuse, at least through this lens.
Isn't this the NRA's argument? The only way to stop a bad guy with a gun is a good guy with a gun.
Or at least somewhere between that and full on protection racket.
LLMs come on the scene, hammer the site until it goes offline or racks up bills that threaten bankruptcy, and the solution then is to pay for an LLM to fix it and monitor it.
That's a real nice site ya got there, it would be a shame if massive datacenters going up around the world were to start hammering it from tens of thousands of IP addresses....
I think people are missing one of the points of the article. It's not just that agents are hammering the site, it's that there might be lurking vulnerabilities that allow malicious usage, which is why it went down, then came back up with a fraction of the data and a reduced design.
The article says (speculates?) that malicious users are trying to get privileged access for an edge in prediction market betting. From the article:
> If you could see The Numbers data before everyone else, every single week, you would have a significant edge over all the other traders - learning the answers slightly ahead of publication would allow you to front-run the trades.
If it's the case that prediction market traders were trying to front-run the market outcomes, the prediction markets should put the fix on their end by locking down trading before a market closes with some time buffer to prevent this.
Why should that be necessary?
You can unilaterally stop trading 'before a market closes with some time buffer to prevent this.' No need for centralised action.
Because markets succeed or fail based on perceived fairness.
It is in a prediction market’s best interest to not become the place where you go to get fleeced.
Today they’re the Wild West and run like the early days of darknet markets, but if the
Could you please rephrase your last sentence? I appears to be incomplete.
That's assuming the prediction markets all play fair and by the same rules. If Polymarket were to do that, its competition could choose not to. Since these exchanges are not tied to any one country (or jurisdiction), its users would jump ship. Because the users don't care.
See also stock market trading, where companies would go to great lengths to shave nanoseconds off of getting information and putting in trades. Now apply this min/maxing to worldwide, unregulated and anonymous.
This is what tech libertarians / cryptobros want.
"One Reddit theory even suggested it was a deliberate rug pull designed to cripple the free site to push people towards paid products."
I know this isn't really the point of the article, but I've been thinking about this a lot. I wonder if we're going to see more resources go this way.
I used to publish little doo-dads as opensource software. Not because it was something that was legitimately ground breaking or anything (it absolutely wasn't) but because I had a problem, and I thought "heh wouldn't it be cool if someone else had a similar problem and could use my resource for it."
But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
And I can see why maintainers of sites like this, or other free but incredibly useful resources might start to get irate at that.
This attitude makes zero sense to me. You benefit from the "training data" just like everyone else does. If you don't, that's a problem with you, not a problem with AI models.
AI solves exactly the meta-problem you describe: "I had a problem and needed to write a one-off doo-dad utility program to solve it." Now you can do something with your time besides writing pointless one-off doo-dads.
As for monetizing the training data, (a) it cost hundreds of millions of dollars to generate the weights, so why begrudge the companies that made the investment and did the research necessary to make it happen?; and (b) rest assured, whatever your doo-dad does, an open-weight model like GLM 5.2 can generate it for free using your own hardware.
So you don't have to pay anyone in that case. Well, except nVidia, I guess. Point granted there.
| Now you can do something with your time besides writing pointless one-off doo-dads.
...thus eliminating all the tedious chatting, relaxing and making friends that people were previously forced to do while waiting for elevators.
...thus eliminating all the tedious chatting, relaxing and making friends that people were previously forced to do while waiting for elevators
That also makes no sense, but I don't know what else I should have expected.
I don't know how else to explain it. It doesn't feel the same to have my contributions smushed into a linear-algebra machine? For me, the incentive was the idea that my contributions might have directly helped someone. That absolutely does not feel the same when I think "My answer has modified the back propagation of a Machine Learning training run."
It doesn't have to make sense to you - I just believe that I'm not exactly alone in this thought.
Now apply this to Art, free stories, writing etc. It feels bad to have your free contributions hoovered up and monetized. It doesn't feel particularly fair or ethical to me. And it'd make me double think before making something free and publicly available.
Understood, there's certainly nothing invalid about your point of view here. It's just not one that I personally can come to terms with.
I've spent a lot of time in your shoes, wasting time on busy-work needed to accomplish a larger goal (and absolutely sharing the results freely, over multiple decades)... and I don't miss that part of it one bit.
This is such a sad worldview. Life without any challenges. Every problem immediately solved without learning anything.
What's sad is somebody giving you a godlike tool and watching you mope around muttering about "life without any challenges."
The godlike tools are paid subscriptions at best and restricted to a handful of corpos by the government at best. The free scraps thrown to the proles are mostly novelty. The other day I asked 5.6 Sol to identify a push prop airplane and it said “this is a helicopter.”
It's only considered "godlike" by those satisfied with mediocrity
The person I replied to is apparently very content with mediocrity.
> Now you can do something with your time besides writing pointless one-off doo-dads
Writing software one-offs to scratch an itch was historically one of my most enjoyable past-times. AI trivializing that has been a very real theft of joy in my life.
Solving the actual problem was never the point, it was just motivation to do geek-out and craft some code.
AI is rapidly diminishing many interesting hobbies (coding, art, music, writing).
Having more free time when there's nothing fun nor exciting to do with it isn't really a benefit.
It does make sense. When podgietaru wrote their doo-dad programs, I guess they made their work available for humans to use and they didn't mind giving their work away for free. Some other person might say "I'm giving my work away free, I don't care if humans use it or AI, for any purpose". Both are legitimate choices, each according to their own.
As for "everyone benefiting from training data" - sure, but the AI companies are not investing Billions of dollars out of the goodness of their hearts, it is towards one and only singular goal of making profits (at some point). People might be sympathetic to these AI companies if they at least behave decently - they take everyone's work (text, software, fiction, music, images, videos...) without paying a penny to anyone. If they take everyone's work for free, they should give away anything that is built on that work also for free. This is before we even get to environment, privacy, hammering sites by not respecting robots.txt etc issues.
If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no? Even if I spent my own money making the meal, it was made from stolen raw material...
If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no?
Yes, and that's exactly my point. We are in violent agreement. It's shared with me, with you, with OP, and with everybody else. We can get a delicious meal for free or we can pay somebody else to serve us a slightly-tastier version.
Here, disregarding copyright law has fulfilled the very purpose of copyright law: to advance the useful arts and sciences. Copyright law was the best tool we had to accomplish that before, and now we have something better.
But the AI companies are not publishing the models? Moreover they are charging for access to the models.
Sure they are. Surf around on HuggingFace and you'll see dozens of open-weight models up to 120B parameters in size from for-profit US companies, freely downloadable (if not freely runnable, alas). Probably hundreds of them, at this point.
And a 3T parameter model is scheduled to be dropped by the Chinese on Monday.
They should certainly publish more, and if somebody were to argue that model weights trained by scraping copyrighted data should inherently be accessible to everyone, I'd be 100% in favor of that.
I wouldn’t be happy if those companies used my vegetables to open a Michelin star restaurant, even if I got a free McDouble out of the deal.
Plant based I guess I don’t know metaphors are hard.
But the McDouble isn't an adequate comparison. Yes, the companies who opened Michelin-rated restaurants with the food they stole from your garden are giving you McDonalds'-level food for free and charging for the rest. Meanwhile, some other companies who raided your garden are giving you the plant-based equivalent of Ruth's Chris or Fogo de Chão for free.
The other thing is, after they stole all that stuff from your garden, it was somehow still there. Your neighbors on Hacker News say that some bandits raided your garden, but you can plainly see that no one has picked any fruit or uprooted any plants, and your security cameras reveal nothing more rapacious than a rabbit or two. You begin to suspect that your neighbors are gaslighting you.
The plant metaphor breaks down when you pause, take a breath, put the McDouble down and realize that intellectual property is a completely different ownership concept than physical property.
Exactly. Plants, like flames, reproduce themselves.
>As for "everyone benefiting from training data" - sure, but the AI companies are not investing Billions of dollars out of the goodness of their hearts, it is towards one and only singular goal of making profits (at some point).
Its not guaranteed they will ever get there, every day it seems increasingly likely that without massive government intervention we are just waiting for local open weights models to become popular.
Meanwhile my ISP gatekeeps the same free content behind a service payment. Their motive isnt free love and world peace, its also profit.
>If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no?
If you cloned all the veggies in my garden you are free to clone them further and fill your belly you owe me nothing.
>Now you can do something with your time besides writing pointless one-off doo-dads.
This is such a disgusting statement that goes against the very core of what open-source stands for. If even only one person got some use out of what they made, it is not pointless by definition.
> it cost hundreds of millions of dollars to generate the weights
It cost hundreds of billions of dollars to generate the training data, they just didn't get paid.
I open source as much software as I can because I want the models to train on it and get better at it!
I feel the same about my old blog articles.
I used to rank pretty well trashing crappy credit cards and encouraging people to switch to better options, then Google decided that 10 results for the card issuers website was better.
Same shit when I manually wrote proto-gethuman posts on calling telecoms/banks/etc (and also pushed visitors to try an Indy ISP or credit unions), then Google felt it was better to drive users to the telecom’s website that wants you to do anything but call them.
Please do train on my pre-LLM gold!
Do you have any active sites or social media with your recommendations?
(I never publish my recommendations because they seem to complicated for people. Like using 30 GB plan/$10/monthly from T-Mobile for data and then using Tello for $10 plan for voice/text. This would require a 2 esim/sim card phone. I am currently using a Moto G Power 2024 phone from eBay $90/new, which was better than the $200 slightly used Pixel 6a from Swappa. Most people would just save the hassle, get a Galaxy phone with a phone contract.)
Nothing active anymore, all converted to static sites.
If you’re on page 2 of the results, you effectively don’t exist so I stopped bothering/benefiting from display ads.
Coincidentally, I did try to get deeper into the cellular service side (it’s another high margin and high customer value segment), but I did better on the finance side.
> But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.
> And I can see why maintainers of sites like this, or other free but incredibly useful resources might start to get irate at that.
Those who object to the scraping fall into several camps, but the biggest complaint I am hearing is that it increases both maintenance costs and time. In other words: it sucks when people are using your work in a manner that you find offensive, but it goes beyond that by doing actual harm.
A change in quantity can become a change in quality.
I think the toxicology proverb applies here perfectly: The dose makes the poison.
I believe that it was Stalin who phrased it most eloquently: Quantity has a quality all its own.
Not Stalin: https://klangable.com/blog/quantity-has-a-quality-all-its-ow...
Thank you!
> The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.
People doing it with a couple of machines and residential proxies versus Anthropic doing it with 2 data centers worth of machines.
Scale matters.
> The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.
> In other words: it sucks when people are using your work in a manner that you find offensive
That to me still reeks of Dog in the Manger mentality. If you publish something for the world to use, you should neither care nor even track, much less discriminate by (or suddenly seek compensation for) who is using it.
Agreed, this is what you agreed to (implicitly or explicitly) when uploading stuff to the internet. In fact, open source embraces this (hence the 'open').
But people are free to not publish things or post things online with a more restrictive license. Not that a license stops things from being indexed.
> But people are free to not publish things or post things online with a more restrictive license. Not that a license stops things from being indexed.
Right. In fact, people are also free to publish things with licenses that condition access on compensating the author/publisher, and they have both social and legal backing to enforce it. This is called "proprietary", and it's not a wrong choice - in fact outside of software, it's the default choice.
The problem is when people publish "free" and "open" as a marketing tactic, where in fact they really want to control and charge for access (whether dollars or karma or credit). That is just plain dishonesty.
> Agreed, this is what you agreed to (implicitly or explicitly) when uploading stuff to the internet.
A more pragmatic approach is to acknowledge that concerning one's self with how something is used once it has been released is an emotional drain. It is a bit much to suggest that someone agreed to something, even if that agreement is implicit, just because they released it.
> But people are free to not publish things or post things online with a more restrictive license.
Licences are meaningless unless you have the ability to enforce them (e.g. to sue). That's why so many companies are willing to ignore the terms of open source licenses. It's also why the attempts of enforcement that we do hear about are usually backed by a third party, rather than being done by the software developer themselves. Simply put, the individual developer (or even small project) trying to make a contribution to the community would be better served by not publishing (instead of using a restrictive license) if they are concerned about how their work is used.
I don't even know if there is a good way to resolve the problem. Consider something like a DMCA Takedown notice. It removes the administrative and legal overhead to copyright infringement, yet it is also easy to abuse. For example: businesses have weaponized it by using it against individuals. Perhaps my cynicism is taking over here, but I suspect any easily accessible mechanism for enforcement would be similarly abused.
This is absolutely fine, until some shitheel decides they're going to continually scrape your website, thousands of requests per second, and knock it offline constantly, so they don't have to cache even a single thing, or write better scrapers. They'll just suck up your bandwidth for it. After all, it's you that's paying, not them.
Search engines were absolutely fine. They were respectful of the sources they scraped. There's now a bunch of money-obsessed shitheels doing whatever they can, as blatantly and as lazily as they can, because they believe they'll become rich. Every one of these fuckers needs to be dead or in jail before the world becomes whole again.
If nobody takes action against these fuckers, everything you care about will just nope out of existence, and these fuckers will be all you have left.
Major AI companies aren't indiscriminately scrapping any more than search engines were (which themselves caused some problems in the past too, until that got sorted out). Your gripe is with SEO people and content marketers.
I'd be more sympathetic if it was really about this, though. However, the majority of the complaints seems to come from people personally offended by the possibility of their content, which they claimed was published for everyone to benefit, ending up as training data, and thus actually benefiting humanity, at scale far beyond the original publication ever would. Which is the Dog in the Manger attitude I point out.
The problem is the rat race for the pot of gold at the end of the rainbow.
A similar thing happened when crypto was in ascendence. Web pages were getting crypto-miners injected into them. Celebrities were shilling NFTs. Everyone and their dog was on the make. The thought of riches broke the minds of millions.
The same is happening with AI. Whether it's the "major AI companies" or millions of self-interested also-rans with fewer moral scruples scraping, the problem still exists. The problem will continue to exist until you can't conceiveably make money by scraping like a bastard. If everybody identified themselves up front and respected robots.txt, there would not be a problem. But they don't, and they don't, and they pummel websites for no fucking reason, and they don't care, and they won't stop.
Websites get hammered by millions of unique IP addresses from residential ISPs which happen to belong to botnets, none of them identifying themselves as a bot user agent, all just pretending to be some slightly out-of-date version of Chrome.
And it's not like the "major AI companies" hands are clean. Who bought up all the RAM and compute power? Who's building datacenters in the third-world parts of the USA and powering 24/7 them with diesel generators, or sucking up most of the power generator, and of course using up all the drinking water to cool their machines?
Everything in humanity and nature is there for the AI bros to use and dispose of, provided they come out on top. This is the attitude that needs to be taken down.
You may be right about the scale of also-ran operations, even though I disagree with the core comparison. Unlike crypto coins, which require superlinearly growing energy waste just to sustain their basic guarantees, and were created to solve "problems" that aren't problems and don't need solving (hint: trust is a feature, not a bug), AI actually works. It delivers real value for cheap. The growth isn't artificial, it's a product that doesn't even need much marketing[0] - it exploded organically since ChatGPT, since it's obviously that immediately useful for approximately everyone in some aspects of their work or life.
One nit though:
> And it's not like the "major AI companies" hands are clean. Who bought up all the RAM and compute power? Who's building datacenters in the third-world parts of the USA and powering 24/7 them with diesel generators, or sucking up most of the power generator, and of course using up all the drinking water to cool their machines?
They paid for it, much like everyone else.
> Everything in humanity and nature is there for the AI bros to use and dispose of, provided they come out on top. This is the attitude that needs to be taken down.
Here you're arguing against the basic market economy. They aren't using and disposing of anything they couldn't buy for that purpose like literally everyone else. There's no theft or trickery going on here. There's a boom, because AI is that useful, but it's still all normal resource allocation.
--
[0] - Of course the competing players invest tons in marketing to gain an edge against the other players.
Your rhetoric is harsh but you’re not wrong.
> But now I'm really reluctant to give more stuff to the free web.
I feel the same way too. But guess what, all the code you did not publish gets into the training corpus anyway (when you gave Codex or whatever full read access to your filesystem).
That's only if one is stupid enough to give an LLM access to their computer. Don't do that, it's incredibly irresponsible security wise.
The vast majority of people are in fact this stupid.
I recently had trouble connecting to a new AP on my laptop. After a quarter hour of frustration, I connected to the AP of my phone, asked Claude Code what the problem is, and she found the issue in seconds. I didn't allow her to make the actual changes, but she did have read access to everything, and helped me considerably.
So network-manager gets a new bug report about too-long non-ASCII AP names, I get online, and I don't know maybe Anthropic sneakily learned something from my local python projects. I am one of those vast-majority stupid people.
> Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
I suspect the next AGPL (if not normal GPL) will explicitly prohibit training of closed models against the source material. Ideally it will prohibit training of open weight models too. Either share the entire process, or go piss up a rope.
My understanding of the GPL is that it currently prohibits the training of LLMs, no need to explicitly mention them. MIT does allow it, however.
Exactly. Just like using LibGen is prohibited.
If LLM training is found to be a fair use (looks likely), no license will help.
>But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.
I dont get this.
You provided something for free to help people, but dont want to do that anymore because it might go into training data and help many many more people?
So far LLMs have been loan funded donations of loss leading services. They might never actually make their first dollar of profit.
Meanwhile ISPs the world over have been monetizing access to your content.
I feel like what you mean is that you want control over attribution.
I think the sheer magnitude of the economics have made the scales fall from a lot of people's eyes. For decades people put stuff on the internet for free on the assumption it was "not worth" anything. It turns out that as soon as that commons can be enclosed, we can marshal hundreds of dollars for every single living human, to pay for this commons to be repackaged. The money is there, and we're happy to spend it, we just won't spend it on you.
So basically the “information is free, encyclopedias are expensive” phenomenon from the pre-internet days? Collection, collation, and distribution have more and different qualitative value than the sum of the individual bits.
I think he just means that AI companies are not people.
I also recently made an open-source project with 200 GitHub stars private. I never had a problem with others using it as the basis for their own projects. In fact, that happened, and I received credit for it. But LLMs just hoover up everything, process it, and then spit it back out as if it were their own.
Why would you offer data for free if people are going to pay someone else to access it?
The whole internet is about to experience USA-level greed because people are too dumb not to use AI. Just like Google killed small sites because people were too lazy to look at page 2.
A couple years ago, I was running a website which allowed the public to view all the US government handouts to small businesses during the COVID-19 pandemic. It also tracked all the fraudulent loans being prosecuted by the DOJ, and allowed anyone to run structured queries over the public dataset. There was a "donate" button which took in ~$2k in donations over the lifetime of the site, and you could download the entire underlying dataset (around 10 GB uncompressed) for free directly on the site.
Despite the "download all data" link being prominently placed on the front page, the AI scrapers decided it would be more efficient to download terabytes upon terabytes of raw HTML by paginating through every possible facet on the search endpoint. Even with CloudFront caching results and a fairly efficient backend setup, the monthly bill ended up with around $1k just going toward network ingress/egress, so I shut down the site the following month.
Building on AWS is a financial time bomb.
Are you not able to put limits on how much the site can spend?
Sure, but many a hobbyist has discovered the need for that the hard way.
As of 2026, still not, and probably never.
Completely agree, I've worked at several companies where I saw the AWS bill balloon from five to six figures, usually because of pointless over-provisioning. However, it all became worth it when I saw Jeff Bezos go to space. [1]
1. https://www.youtube.com/watch?v=IOmX793-5t4
BigQuery has "public datasets", so users can even run complex SQL on it, but it's them who pays for it, not you. You only pay for data storage.
The crawlers would have still just hammered their site though, right?
I think the idea is that they could store the data in BigQuery, and point users of the site there.
The crawlers would have still just hammered their site though, right?
Why is your comment exactly word for word of another comment just one level above in the comment chain?
Because the crawlers would still have hammered their site, though.
(The GP post doesn’t actually meaningfully address the issue being raised. Adding BigQuery or whatever would not change the fact that (a) they already offered a method of getting all of the data in a cost effective way, and (b) the issue was that the crawlers hammered the site hard enough to make it economically unviable.)
Maybe because they restated what they said instead of addressing to the previous commenter’s point.
The comment I’m speaking of was from a different user, and it’s giving a strong smell that both users are bots unfortunately. Maybe I’m wrong, maybe it’s a coincidence those exact strings of characters were formed in isolation from each other in the same thread. I hope I’m wrong, but it’s a red flag
I'm not a bot, and the other commenter was restating the question as the solution provided did not seem to actually solve the problem, just redesign the entire project (See comment from u/taneq).
We didn't form "those exact strings of characters in isolation from each other in the same thread" they are directly referencing my comment in a reply thread to said comment.
Sure, but now you're moving the site from "Free data presented in a pleasant way to view" to a "pay-as-you-go database". Your audience shifts dramatically, and you lose the ability to share the data you're trying to present.
True. But it sounds like they already lost that.
No, it's free for everyone for the data sizes they have:
Free tier: 10 GB of active storage and 1 TiB of query data processed per month.
There is also the deep magic...
https://github.com/phiresky/sql.js-httpvfs
I'm curious how this compares to just using DuckDB in the browser?
https://duckdb.org/2021/10/29/duckdb-wasm
Yes, if I wanted to put a nice HTML interface over the query results then I would still end up with the same problem where some combinatoric explosion of query parameters to the `/search` endpoint, most of which are cache misses, leads to many many TBs of network egress.
I think snowflake has similar, you can rent it out or make it free
Thanks for that tip, this is exactly how I would do this if I had to do it from scratch. Just in case it's useful for anyone else: https://docs.cloud.google.com/bigquery/public-data
Do you have any view on why the AI scrapers resulted in a heavier load than existing crawlers from eg search engines?
Where they more exhaustive or more frequent?
not the OP but I'd say that google crawls you once and AIs scrape your page every time someone asks them a question that they think your page might be relevant to.
There's been an explosion of vibe-coded scrapers that behave poorly and ignore robots.txt. Presumably OP disallowed crawling of the search endpoint for the reason stated (that crawling it would result in endless permutations of search filters.)
They were both more exhaustive and more frequent. I don't remember the exact numbers, but it was definitely over 100x the traffic from search engines. By the way, search engines were allowed under robots.txt - I did want all 11.5 million loans to be individually indexed so they would pop up in Google search results, and I actually did receive/forward multiple tips about fraudulent loans because someone searched a business name on Google. All of this traffic was barely a blip, and my AWS bill for my hobby data science account was only ~$50-$100/month.
When I looked at the logs after getting the billing alert, 99.99% of the requests were to the "/search" endpoint with virtually every permutation of ~10-12 facets in the query parameters. There was only one scraper, but it triggered an enormous amount of network egress since it ended up missing the cache on the majority of queries.
That’s a traffic design problem. You should be happy that your work is valuable and also protected it against excessive requests. Simple.
I deployed a Lambda function behind CloudFront which rendered a simple HTML page with the query results from executing some SQL over the dataset. I served millions of page views for next to nothing because most pages were already in the cache.
I don't know where you get this expectation that people should anticipate that a crawler might try every possible combination of query parameters, thereby missing the cache on each one. Most people consider it a bitter and arrogant perspective, which is why this got downvoted.
For what it’s worth, a Hetzner dedicated server has unlimited ingress and egress. It seems like it’s the only provider that does. Egress fees suck.
You can get a beefy one for about $40/mo on their server auction site.
Just... don’t miss payments. Ever. Or they’ll delete your server within a week or so.
the AI scrapers decided it would be more efficient to download terabytes upon terabytes of raw HTML by paginating through every possible facet on the search endpoint.
No, they decided that would be a great way to convince you of the narrative and persuade you to pay for "security" services that further the incumbent browser monopoly.
They're not "AI scrapers", they're DDoS'ers manufacturing consent.
One thing I’m wondering is whether we need an open source set of technical patterns and libraries for dealing with this changing traffic mix.
Millions of small sites and creators don’t have the ability to design their own protections against large scale automated access. If useful content now attracts aggressive crawler traffic, many sites will be too expensive or unreliable to run.
Is part of the answer a community response? Perhaps a community-maintained toolkit, based on traffic data, that host sites can apply? It could include standard agent identification, rate limiting, traffic classification, access policies, caching, challenge mechanisms, logging, attribution and usage control etc etc.
In effect, we need much stronger road rules for today’s automated traffic, available as open technical patterns and libraries rather than every site owner having to invent this alone (they won’t).
The community response was respecting robots.txt, but since that was more of a gentleman's agreement than a legal requirement, hungry AI scrapers disregarded it.
Which brings us to the old fashioned flood control mechanisms. That is, the toolkit you propose already exists and has been used for decades in various iterations to protect against various forms of attack (slashdotting, DDOS attacks, overzealous search engines, and now AI scrapers).
Have you looked into those before? Companies like Cloudflare have been at the forefront of this field for a long time now.
Yes, a community effort against the botnets would be great.
The "good bots" which identify themselves and follow the rules are easy enough to block, so not a problem. It is the "bad bots" which pretend to be real users and hide behind residential proxies, and so are almost impossible to block at the moment, which are the problem.
Given that there are big companies openly (i.e. on the clearweb, not even darkweb) selling access to these residential proxy botnets of compromised smart TVs[0] and mobile phones and other devices, can we not simply get a database of the these IPs and block access from them? I would venture that almost every single one of those residential users are unaware that they have devices in their homes which have been compromised and are being abused in this way, so if they were to start seeing messages from more and more sites along the lines of "Access to this site has been blocked because unusual traffic has been detected from your computer network. Please check all devices on your network and remove any malware which may be routing this traffic." then maybe we could start addressing the problem at the source.
[0] https://news.ycombinator.com/item?id=49000864
> The "good bots" which identify themselves and follow the rules are easy enough to block, so not a problem.
And then they get blocked, which is a problem. In particular, the modern agentic AI tools interacting with web services to fulfill user queries - they are acting as user agents, and they should not be discriminated against.
So I'd say the first pattern that needs to be broadly adopted is non-discrimination of user agents.
But of course we've tried that in the past, the whole problem is that non-browser user agents == end-user automation, which is anathema to pretty much every on-line business out there, as money made online is primarily conditioned on users wasting their own lives on interacting with services directly.
I don't think there is a technocratic solution to this problem. Very rich people are using their money to DDoS the internet, for no good reason. We just need to identify these people and fine or imprison them until they stop doing it. Residential proxy providers would be a good start.
AI companies really do socialize the costs and privatize the profits.
Sites like The Numbers have to take on the cost of surviving the AI onslaught and the AI companies return nothing back to them.
worse; it isn't just no return, it is also cost.
To a point; as the post mentions, they added instructions for LLMs to their website and increased their licensing inquiries tenfold (the article didn't state anything about actual licensing payments being made though).
I don't think there would be any issues per se if the scrapers just paid licensing fees to get the good / complete data. But the issue was that they started to try and find exploits to get to data earlier.
At the risk of oversimplifying things from a distance, this site -- especially the free, public-facing part of it -- seems like it would be an ideal candidate for a rewrite using static site generator/framework. That, coupled with a bot-aware CDN should keep them online at a reasonable cost for many years to come.
Otherwise, I'm very curious to know more about their old and new architecture and what sorts of mitigation/scaling strategies they've started using to keep the site online.
Yeah this seems like one of the easiest access models to create an excellent security model for -- basically just static publishing. Sounds like they were in need of a rewrite anyway!
It probably seems daunting but to be honest this feels like a weekend's work at this point with LLM assistance. Not to be glib!
But it also destroys the business model behind the site.
Sure, you could technically redesign to handle the bot traffic, but if the bot traffic is just taking the data and reducing any need for humans to visit the site, why is he putting in the effort to maintain the site?
Well, why is he? If it's for ad revenue, this problem is surely global to the web. If it's because he likes to do it, it shouldn't matter if one AI or a trillion scrape the site. I feel like we badly need a rethink of the web architecture after DoubleClick anyway. Maybe the name of the game should be to cut down a site's assets very hard and use static hosting for them. This interferes with crummy sites that show a mess of inlined ads every refresh, but now that there are a billion 'poor users' this is no longer feasible.
Maybe more an inherent problem with these chatbots.
LLMs are great, but they aren't producing new information. You still need people for that. But if you cut down any incentive for the people to do that, the LLMs will starve.
The business model of the site is apparently private data sales, not ad revenue.
Bots don't make purchasing decisions.
Some people are already trying to make that happen, but I'm convinced it'll be a net-negative.
Rephrase: Bots SHOULDNT make purchasing decisions.
Because... ? If I'm operating a site, and I want $X to allow a bot to scrape my site, why shouldn't the bot be allowed to make the purchasing decision to scrape the site? Obviously the bot owner would have their own set of guardrails, but if it allows sites like The numbers to stay up because bots aren't going to look at ads so the old model of displaying ads isn't going to work, I have a hard time seeing that as a bad thing.
I mean it from the consumer perspective.
According to the post inquiries for purchasing the data increased as a result though so they might not make the decision, but they do make the suggestion.
Or just use a cache.
Ya this problem was solved 20 years ago with Varnish cache and Coral cache/CDN.
I think what's really going on is that bots expose how underpowered web servers has gotten in recent years. In the 2000s, even poorly-architected PHP sites tended to serve about 200 requests per second, with 1000+ being common for static sites. I remember when Node.js came out and claimed that it could serve more like 100,000 RPS due to its cooperative threading model. But today sites have a remarkable slowness to them, running many hundreds or thousands of database queries due to ORMs and N+1 problems, so that response times can be 500 ms or more and even 1000 simultaneous users stresses servers.
What went wrong is that nobody solved stuff like Russian doll caching in a general way separate from the programming language and database. We should have had ways to make dependency graphs using Etag headers as keys with real cache invalidation of dependent data. I went down that rabbit hole 10 years ago using touch events in Laravel with callbacks to handle cache invalidation when class model data was saved to the database. Also a query cache using Redis which I think might have been handled better at the database level anyway. After that experience, I can honestly say that cache invalidation is so difficult to get right that it's effectively an open problem. Meaning that programmings should use a package instead of rolling it by hand, and it should be a major concern from the start (along with sharding by user id or using something like Firebase).
Don't get me started on how the web should have been a P2P content-addressable memory anyway. Nearly everything should be available from a nearby edge peer, similarly to BitTorrent. But nobody bothered to solve how to make that work with HTTPS/SSL. I suspect that has to do with early flaws in the browser security model where the whole page has to be behind HTTPS or warnings appear. So it was never clear what was personally identifiable information (PII) or merely public data being served over HTTPS. To really solve that, we probably need real trust networks and maybe even zero-knowledge proofs.
Since these problems are so challenging to fix, and big companies can't be bothered to do it since they pulled the ladder up behind them, we're probably stuck with banal "are you human" challenge screens for the foreseeable future.
They are indeed quite challenging. We literally wrote a paper about it (https://arxiv.org/abs/2605.09114) ... you almost describe some of our original architecture too! You might find it interesting -- or not. https://getswytch.com
> Ya this problem was solved 20 years ago with Varnish cache and Coral cache/CDN.
It was... if you are paying datacenter rates for the bandwidth
If you're paying cloud provider per GB pricing, nah, even text will add up if you happen to be targeted by a bunch of bots
> What went wrong is that nobody solved stuff like Russian doll caching in a general way separate from the programming language and database. We should have had ways to make dependency graphs using Etag headers as keys with real cache invalidation of dependent data.
we did that with nested ESI includes in Varnish so every "box" of content on the page was cached separately + some piping for invalidation, so if a given piece in the database was changed it sent invalidation to all nodes. There was also some grace so if the thing you wanted got updated RIGHT NOW you might get stale version while the new one is updated in background, and don't pay the latency cost
> Nearly everything should be available from a nearby edge peer, similarly to BitTorrent.
Back in the days of Gnutella, I remember pushing people to use Magnet links [0] when sharing content.
[0] https://en.wikipedia.org/wiki/Magnet_URI_scheme
> I think what's really going on is that bots expose how underpowered web servers has gotten in recent years
More often than not they're running on a VPS, and cloud providers have been pushing the envelope of what a vCPU is. Amazon still bills hyperthreads as a single core!
Add to that their servers are often aging and you get a recipe for slow web services
What a throwback. Back in 2015 I have started Applaudience, which at the time was the only provider of real-time cinema ticket sales data. I still remember comparing our numbers against TheNumbers.com as part of calibration. I have since moved on to other businesses, but this remains one of my favorite pieces of technology that I have developed. Would love to bring it back one day.
Just FYI, the bigger companies all allow you to block crawlers via robots.txt:
This is good to know, but a bit of all-or-nothing. It's a shame that, for example, Google doesn't support the crawl-delay field so you can tailor their crawling to your setup: https://developers.google.com/crawling/docs/robots-txt/robot...
I presume it would also cut you off even more from referral traffic.
They don't support it yet. It seems as if Cloudflare is trying to use it's power to force Google to change this though.
https://techcrunch.com/2026/07/01/cloudflares-new-policy-pus...
90% of bot traffic on my network of websites is through headless Chrome, via residential bots nowadays. Impossible to block. Not even for Google, as they inflate my Adsense numbers as well.
I once had the dumb idea to make websites in pdf but the idea is growing on me. Perhaps it should even be in animated gif with <img> <map>'s
Claude Bot still (at least last month, and it's been doing it for over a year now) seems to have a bug when traversing (at least my sites), wherein it drops the trailing slash of a directory (which is present in the a href tag), then makes the request to the subdirectory without the slash I put in the link, then Caddy automatically responds to that (via the built-in File handler) with a redirect telling it to add the trailing slash, and Claude Bot then makes the request again with the trailing slash that should have been there in the first place.
So most subdirectory URLs get two requests from Claude bot, the first one needless because that wasn't the URL in the tag.
I won't eat your lunch if you put a sticker on your lunch box telling me not to.
This only works for crawlers that play fair, unfortunately there's many that don't. Including ones that use consumer devices like TVs [0] and mobile phone apps that offer incentives to consumers (if they're even open about it) to use their internet connection to e.g. crawl websites. There's probably an army of hacked toothbrushes and the like doing the same thing.
[0] https://news.ycombinator.com/item?id=49000864
I have small website with archive of older radio broadcasts mostly in Russian. Lately, 90% of traffic comes from USA ;)
Previously it was like 5%.
URL is http://radar.lv btw
Thanks God i can serve up to a terabyte per month of traffic easily, otherwise it would be a catastrophe. I am not against bot scraping, but I worry about stability for meat visitors.
So, I think of enabling payments for website visits cloudflare recently developed.
I also have the problem of old technology like the mentioned site. While my personal blog uses static generator, archive website uses ancient Drupal version, which has no security patches for many years already.
Genuine question: do prediction markets open up a new revenue stream for TheNumbers.com?
Specifically: they suspect that the motivation for trying to hack their site (for at least some people) is that they wanted early access to numbers that folks were betting on in prediction markets. Since they've got those numbers they can just bet on them, then benefit from their perfect knowledge.
On the one hand this does seem incredibly unethical (it's clearly insider trading). On the other hand the CEO of PolyMarket has said that insider trading is part of the point of PolyMarket: https://youtu.be/ZN4njIQcSR4?si=ztyTtgjeHSJbNjSZ&t=1566
There’s an easy solution: Bruce should place large bets on the prediction markets right before publishing the relevant data. It is legal, ethical, and, for him, risk free.
It feels like there is fundamentally missing infrastructure here that is needed to make these problems go away.
Basically bots need to be (somehow) paying for the traffic they create, or prevented from creating it, or told to go away and then fined if they violate the request. No idea how to do these or even at what level in the stack they should happen, but they need to happen eventually somehow.
Every website I visit could get a fraction of a cent in my Cloudflare Wallet. A human browsing incurs a few dollars a month. Plus, any website I access in this way decides not to show me ads either and just charges me the cost of serving the data plus a nominal profit margin.
Disclosure: I am a shareholder and would love for them to solve the AI bot problem and the ad problem like this.
Maybe we'll all eat our words and cryptocurrencies will actually become useful for something.
This does not require cryptocurrency
we've been working on basically the same problem in the email space for a few decades (legit email vs. mass spam). its a very hard (i think impossible) problem.
The same goes for telephony.
I reject all phone calls by default, unless I'm expecting a call.
Well, it is technologically not that hard to imagine a solution---the problem is social, getting everyone to agree on how to do it, figure out the policy around it, etc. And the situation is getting so bad that maybe it is time for someone (maybe someone reading this thread) to figure it out.
I know people have opinions about Cloudflare but why not use it here, at least as a stop gap? Stopping bot traffic is one thing it does very well.
Cost?
It's free for this
It tried it. Hundreds of “genuine” visitors per day on a new website with no search engine presence and no links. That’s a very leaky fence..
Hundreds is better than billions.
I set Cloudflare up a couple of months ago specifically to block bot traffic. It didn't do anything for me. Dumb bots were still hitting every special link on my wiki fast enough that the server was continually swamped running Lua scripts. 65% of the traffic for my English-language site was coming from Vietnam. But I didn't want to block Vietnam altogether, because my hobby site has genuine users from there too.
I eventually just shut down my mediawiki instance. I couldn't find a way to keep it online and still run on an affordable VPS.
I think I probably use AI like a lot of consumers out there. Search engines have gotten bad and AI really good at answering fairly specific questions. Often pointing at sites like wikipedia. Pretty clear changing my behavior will have zero impact but it is certainly part of the problem. Feels a lot like my CO2 consumption.
*Edit - CO2 creation
GPTBot crawled millions of pages on my website. I get a handful of visitors from them. Meanwhile Google sends me 10k visitors each day.
So GPTBot is now blocked.
Only solution I see is to ban user that abuse the system. Too much requests, too much bandwidth and you get into the blacklist and are cutoff from the human side of the internet.
With the amount of money AI labs are burning, someone should just set up some infra to host the data they want and charge for access to it, instead of abusing the goodwill of legitimate pages that offer it for free.
Our internet ecosystem is becoming more and more hostile to open information commons. I like open information commons -- what do we do about this?
The Numbers is such a brilliant site, this explains why it came back in such a stripped down form. AI and prediction markets giving me more reason to hate them.
If the AI companies destroy the open web, eventually they'll need to start curating knowledge sources just like netflix makes movies and amazon has physical stores...
Already happening. Frontier LLM vendors have been hiring human domain experts specifically to create their own proprietary training data in targeted verticals.
Prediction market as well as anything related to gambling must disappear from earth surface. Full stop.
This is such a sad worldview. Life without any challenges and solved using AI. it will vastly decrease our learning
The real story here is that prediction markets were banned for a reason and loosening the rules is causing chaos just as was expected. AI plays little role in this story, hacking by humans would also be motivated by financial returns, unless the element is that AI hacking is cheaper and the returns are not so big.
Do the AI labs in the U.S. publish the AWS IP address sets that their crawlers use? Is that not an effective way to block those crawlers anymore?
They are often from residential IP.there are even businesses that allow you renting such IPs
Nobody serious would ever use an AWS IP lol
Good read, I was so curious how this happened a few months ago
This article, and possibly the owners of the site, seem to muddle together several problems:
1. bot traffic causing infrastructure cost
2. scraping circumventing paying for licenses
3. the risk of hacking
You missed the speculated motivation: unrestricted prediction markets, which are an open and broad incentive to do whatever actions might provide a slight edge in betting.
Webscraping is among the more benign things that gambling-on-anything can drive. And even that has a negative impact, as seen here.
I have no idea why those sites are legal.
Ironic that polymarkets were being advertised as helping society make better predictions.
John Brunner and Alvin Toffler both made stark warnings wrapped in futuristic giddiness about things like it. People like Fuller no doubt thought polling on a large scale was terrific. There were old usenet groups and BBS subs (minus the money aspect) experimenting with the model. I do not believe they ever are or were good in a largescale model (money or not).
I mean wisdom of crowds, super-forecasters, calibration and pre-registration are useful tools that can result in better predictions, turning it into online gambling is where it went sideways.
Mini-rant: Promoters claim that the system serves a public good by helping society discover/converge on useful truths sooner.
Even where the question/event is of public interest, the opposite happens instead. People with expertise or non-public information are incentivized to misdirect and delay as long as possible, as that maximizes what they can make from betting.
If the scrapers are going to get it anyway, put the data up as a zip somewhere.
It would not make any difference.
Indeed, the article mentions Wikipedia experiencing similar scraping pains, even though they already DO have bulk data available.
Who are running these bots? I presume developers at all of the frontier labs know (or at least would know to look for) Wikipedia has bulk APIs for automated access. Unnecessary scraping increases their workload/costs too, so why in 2026 is this still a problem?
Black market and gray market data. All the firms want data. All the other firms want data. The banks want data. The other criminals also want data for their crimes and schemes. Oh insurance companies, and the ATS systems. Everybody wants as much data as they can get and they don't care how they get it.
This data selling also happens with leaks of all kinds like medical data, often to current or future employers, health insurance companies, etc.
> Read the Docs, a non-profit that hosts documentation for open-source software, who watched a single crawler download 73 terabytes of zipped HTML in one month, costing it over $5,000 in bandwidth
From the article.
Not the same site, but an example of the same issue.
All new content will be behind paywalls. This will be the new way the Internet works and I've said this for the last several years. There's no value in writing unique content just to have AI steal it and distribute it without you getting any clicks. The only way it works if it there's a licensing agreement and if they don't want to pay, then who cares, you weren't going to make any money from them anyway.
What if you aren't motivated solely by making money? There will still be new free content created by hobbyists, passionate creators, altruists, etc.
(Agree with you more generally.)
Why would you produce content and have no one read it or visits your site, but OpenAI and Anthropic make millions from it? At some point it becomes stupid to just give free money to these companies when they steal literally all your traffic and content. As per the article, Anthropic sends 1 page view for every 38,000 views they get.
Very few people showed up to read stuff on the old web, and yet: It existed.
The "old web" hasn't existed in 30 years now. I was a part of the old web and if people in the early 90s knew that someone was profiting off their work it would have killed it immediately.
prediction markets were a huge mistake.
I always liked this site, but reading this and seeing the anger about expecting the site maintainer to do things for you is repulsive. Frankly, if he wanted to pull his site down with no notice that is perfectly within his right. It was/is his site. He doesn't owe anyone a .tar.gz either. His work.
I agree. I've seen it happen before on a project I use. I decided to take a look at the repo for one of the plugins, and I saw a heinous issue that basically was TELLING (not even asking) the maintainer to fix it.
Deplorable behavior indeed
Like dominoes, as soon as it is accepted in a few places, people think it is acceptable to push to 'share'. It's almost terroristic sometimes, the pressure some maintainers are under.
Eh, wouldnt just adding some POW challenge, like anubis solve the scraping by making it hurt crawlers wallets?
This is another reason we can't have nice things. The level of entitlement required to send an angry message to someone complaining about their free resource being offline is mind-boggling, though.
What does a cyber attack have to do with AI scraping?
they both have a risk of harm which the site operator was no longer comfortable with.
Every example of this confirms my view that we should treat the spread of these AI scrapers like we treat the proliferation of drugs. We should be seeking to bust AI rings like we seek to bust drug cartels.
Its just too easy for a technically minded bored person to produce slop that hammers websites. There needs to be a penalty for this.
Sorry! Just yesterday many of us decided that AI scraping isn't a real problem, and anytime it is blamed, it's a cover for something else.
https://news.ycombinator.com/item?id=49005747
They should sue Anthropic for this distillation attack
> The world we have built thus far is so incredibly ill-prepared for the power and scale of the AI models we all have access to.
Exactly, so use the AI to secure your servers. Ask the AI to audit your site for any security holes. If unable to rewrite, at least harden the existing code. AI’s are really really cheap (and fast) security consultants now.
This is an odd example to present this argument through.
I'm not defending AI bots overwhelming websites or hackers motivated by Polymarket, etc., but I don't really believe a 30 year old website with "approximately 160,000 source files serving around 2 million pages" is a good litmus test for the state of the online world. What's worse is the absurdity that this basic site offering niche data would be targeted because Polymarket depends on it for some of their bets, something that most websites don't have to deal with. Frankly, that seems like a much more interesting angle to explore than "this old website that should have been re-written multiple times over the last three decades now doesn't have a choice but to be re-written."
Oddly, it seems like the solution, at least in this specific case, is relatively simple (and ironic): pay for an LLM to re-write this site in a modern, scalable, secure way and have that LLM monitor the site to ensure it is behaving correctly. Put this thing in a modern platform behind a proper cache (i.e., throw the whole thing up in Cloudflare) and many of these problems just don't exist anymore. If this site is as basic as it seems, this could be a weekend project.
It's certainly an interesting story, and I don't blame the owner of The Numbers for handling his site the way he has, but framing this as "AI is destroying our beloved internet" just seems obtuse, at least through this lens.
Isn't this the NRA's argument? The only way to stop a bad guy with a gun is a good guy with a gun.
Or at least somewhere between that and full on protection racket.
LLMs come on the scene, hammer the site until it goes offline or racks up bills that threaten bankruptcy, and the solution then is to pay for an LLM to fix it and monitor it.
That's a real nice site ya got there, it would be a shame if massive datacenters going up around the world were to start hammering it from tens of thousands of IP addresses....
The actual solution is damn simple and painfully obvious: make things like so-called "prediction markets" illegal.
The intelligence community came to the conclusion (after research and experiments) that things like "prediction markets" were a bad idea in 1996.
Far worse now than then, with the net and AI of 2026 and the insane number of people now online that want to make an easy buck.
Don't get me wrong; I am guessing some of us on here would do well on those places. But they should not exist.