I really have trouble getting worked up about this. "Rare books" is thrown around regularly but my gut feeling is that's not the case. These are used (often? always?) books and while I'm sure there is waste, in general they just want 1 of every book.
While I wish there was a repository of every book that was already digitized (it pains me this is the best solution), there isn't one and so I think this is not a real problem.
It'd be a different story if they had furnaces that ran only on rare books that they had to continually feed books to but that's not what's happening here. And that 1 destroyed copy will "live on" in a way that it otherwise might not.
It'll live on if they publish those scans or contribute them to a national archives or something. Proprietary data has a habit of being lost over time though.
It seems a little short-sighted to destroy an artifact to get the text.
Perhaps the genetic material that remains in books from the people who handled them, or the pollen from plants in the environment that the book existed in will have value in the future, but we won't know what we lost because some people foolishly destroyed it in a bizarre quest to make AGI that the creators argue could potentially destroy humanity.
The more and more I read about these kinds of people the more I'm starting to realize that they're in the "here for a good time not a long time" group of people and those are the last people you want making long-term decisions.
Millions of books go in the trash every day. Should all private property come with such an asterisk that it ought to be preserved for whatever unlikely contingency you can imagine? I don't know, AI actually seems like a much more worthwhile pursuit.
In one of the first articles that I saw about this topic they were describing a scenario where a botany manuscript from the 19th century that only had three remaining copies was being digitized and then pulped.
They don't care about the R.L. Stein Goosebumps series or Louis L'Amour pulp fiction. They want the stuff that isn't digitized and they're willing to destroy original artifact to do so with no guarantee that even the digitized text will ever be made publicly available.
> In one of the first articles that I saw about this topic they were describing a scenario where a botany manuscript from the 19th century that only had three remaining copies was being digitized and then pulped.
That was something that someone on Twitter made up and then got quoted in an article. All of the books being scanned have ISBNs, meaning they're from 1970 or later. The books being sold off are rare in the sense that there aren't many copies in circulation, but they aren't priceless artifacts. If they were, they'd be being sold individually at high prices by boutique booksellers rather than being sold by the pallet by wholesalers. It's going to mostly be stuff like "The Best Places to Eat in Milwaukee, 1992 edition" or "The 2006 Guide to the Stock Market" or "Teach Yourself Visual Basic 3.0 in 28 Days".
On its own that seems like a pretty weak justification, at that point we would need to preserve every copy of every book forever because you never know what what useful physical material it has on it even beyond its text. That extends to non-book items as well, and while I understand the idea it's just not realistic in any sense.
I think the more important part is to define what "rare" actually means and save/digitize set copies of those books that are actually at risk of being lost completely (rather than letting them rot away somewhere or get bought up to be privately destroyed). I suspect many of them are just not at all interesting enough to justify the expensive though.
My comments on this subject aren't some sort of justification for any particular course of action; they're more a disapproval of the course of action being made by people people who have more dollars than sense.
As a general rule preserving historical artifacts for continued future analysis and appreciation by people who have yet to be born is a noble cause that's considered worthy in and of itself without the need to justify it. We're talking about the richest group of people who have ever existed with the means to preserve them but decline to do so and instead act like the Taliban blowing up statues of Buddha.
I think this is intuitive to most people. If a prerequisite for the birth of AGI was that a humanoid robot had to sit down in the Louvre and eat every single painting there with a knife and fork while occasionally stopping to wipe away historical detritus with the Mona Lisa that it's wearing as a bib people would by and large express a visceral and justifiable outrage.
The people who are doing this know how this looks so they're trying to do it all behind closed doors, through cut-outs and intermediaries.
They are literally not allowed to publish those scans.
Half the reason the books get trashed in this process is because the first sale doctrine keeps copyright from strangling all the freedom in this narrow area.
If they are in the public domain, they could publish them. How many books are in the public domain but have never been digitized and shared in a public archive?
And if the books are not in the public domain, then they should not be allowed to train their AI models with the material without some kind of license or agreement with the owner of the copyright.
> If they are in the public domain, they could publish them.
That'd be an easy fix with new law -- if you are an AI company with book data, you have a burden to make openly available (or require your suppliers to) all public domain book scans.
> They are literally not allowed to publish those scans.
Emphasis mine. Just provide a digital copy, no questions asked. They will distributed to various archives globally. I'll pay for the drives and shipping. I understand and can appreciate the potential liability, and am willing to launder it to preserve the subject collection(s) and dataset(s).
We run a little library in front of our house. It's amazing how many people dump boxes of old books off in front of it hoping that they will find their way back into someone's collection.
The sad truth is the vast, vast majority of printed literature is neither interesting nor useful. People are not dumping off stacks of Umberto Eco. We frankly have to toss a lot of awful cookbooks, self-help books, trashy mass-market "novels", and sketchy religious works. As it is, even the stuff that makes it to the library is not very impressive.
Imagine if we had bad cookbooks, cheap popular novels and stories, tracts on diet and self help from old Rome, or old eras in China, or the equivalent from ages before that.
It's all interesting for something, even if it's just a meta analysis of culture during a certain period or what kind of trashy romance novels were popular in 198X. At least in my view.
This is most of what we have for the great philosophers - half "why is the world like this" and half "how do you live a good life", right? some original plato lectures would be interesting to have.
> Imagine if we had bad cookbooks, cheap popular novels and stories, tracts on diet and self help from old Rome, or old eras in China, or the equivalent from ages before that.
It would be great, but it's a bit off topic here. The point the parent is making is that most of these works are going to be trashed regardless of Amazon's behavior. If Amazon (or any other company) is digitizing it, they're at least preserving it in some form. The alternative may well be that many of these books never get preserved.
Now would I prefer the government or an entity like The Internet Archive do this? Sure.
They were all lost when the library at Alexandria burned, so clearly we should ban libraries. Too much risk of losing unimaginably priceless works all gathered in one place like that
Books are not original manuscripts. Even in low volume cases, they are usually printed hundreds of times. (And usually low volume works aren't all that great...hence the low demand.)
That is a fraction of a percent for books that at some level weren't all that wanted.
Do we know if they actually buy one copy of each book here? Or are they indiscriminately buying books in bulk and just processing all of them?
The latter seems inefficient, so my first assumption would be that they would avoid that. But while I suspect they'd check if they know the book before scanning, I could imagine them not caring that much before buying them and just focus on volume.
The Embassy of the Free Mind (https://www.embassyofthefreemind.com) is a rare book library in Amsterdam that is scanning books the old fashioned way… leading to https://SourceLibrary.org — a collection of over 5,000 books from the renaissance that have never been translated before. Consider donating, if this is a topic you care about!
The length of this article and the unnecessary visualizations felt like a waste after getting to the end and discovering they won’t reveal anything about the books that were scanned.
The closest they got was admitting that the rare books weren’t anything that someone might care about in the sense that people assume when we hear “rare books”
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
Okay? But then why exactly where they considered valuable enough to warrant an entire article about them going to a book scanning facility? Without revealing anything about these books I have no idea if they were classic literary works that were underappreciated, or if this was some old guide about How to Use Microsoft Office 97.
There was a more balanced take on Twitter (which I’m unable to find again, because Twitter) from a book seller who said it was more of the latter type: Books that were rare because they were no longer in demand and most everyone had thrown their copies away. Some parts of the media are doing backflips to try to imply that these are cherished literary classics being fed into the shredder to deprive humanity of something valuable, but the book seller seemed happy to be making sales for useless old books that no human was interested in buying.
Anyone who is into books eventually finds out that we're permanently losing them all the time. Like they're thrown out, and lost forever. University libraries throwing out huge collections of out-of-print material to make room for new books, study spaces, even cafés. Municipal libraries turning over their collections. Books that never made it to libraries going out of print, tossed in the trash after yard sales.
The AI companies digesting this stuff is a net win for humanity. And I'm not a fanboy! Ideally they'd upload them to Anna's archive too, but even if they keep it private forever, at least these books live on in some way in the model weights. Thats better than a landfill.
If or when they go bankrupt or reach agi, they will just delete them. I hope Anna's archive already has them anyways. Apparently openai's newest unreleased model, gpt 6, is capable of continuous training at Inference time, like a person is. That might be enough to delete them
And that’s why imo the outrage should be about copyright law not the businesses. I think it’s a hard problem to solve but right now with copyright on new books being the life of the author + 70 years it’s a bit silly.
I like this and think it could make a lot of sense but you would need to thoughtfully tie it to volume or something similar. I could see publishers gaming the system. I would also add that once it drops out of print that it should belike generic drugs. Anyone can use it. IMO making it quicker free use stops most of what this article is describing.
I would feel much better about this process if they were uploaded and if it were framed as a knowledge preservation project. This would only slightly increase the cost of the project, but have a huge impact on its perception and its net positive impact.
Of course, actually benefiting humanity is only a minor, indirect concern for investors.
I thought this was the original purpose of Google Books. It's actually a little surprising to me that there's anything left out there worth buying in physical form and scanning that hasn't already been digitized in some way.
They could scan and store the data until the copyright expires without issue probably. It's a long time line to work with but you could probably get away with it that way.
It would be really great if the companies who are doing this would commit to placing the scanned files into a public trust that would coordinate with organizations like the Gutenberg project to ensure that the scanned materials enter the public domain on schedule. Publishing encrypted archives with the keys in escrow would be a good first step.
IMO that would go a long way to resolve any concerns about losing books. I still don't like the idea of extremely hard to find or last prints being actually destroyed for this, but it certainly makes it more palatable.
They cannot do those things. The reason they’re scanning the books is because it’s not possible to legally obtain or transfer their digital copies. They have to do their own scans and keep them in house.
First, to be clear, I work for Amazon, and I don't work on this or related efforts, I am a security engineer. I also won't talk about or answer questions at work, and my comments are more generally about the practice (all of the major organizations building AI are doing destructive book scanning). These are my opinions, and do not reflect my employers (past or present).
You might be right that they can't do that now. The simple path forward is to have these companies simply make a public commitment to publish the data when the copyright expires. I also think there is a space to be carved out, probably through regulation, to ensure that there is a clear path for these scans to enter the public domain, at the very least.
This is not just important for these specific books, I have written on other platforms and in other spaces about the importance of media companies and those who benefit from strong copyright laws to protect and generate profits and revenues to repay the public for the cost of that enforcement over time by ensuring that at the appropriate time, those works fully enter the public domain. That could mean a restructuring of the Library of Congress in the United States to become a modern Library of Alexandria to host data, and shifting to a registered copyright model where to gain the protections of the court, you need to upload/submit your copyrighted works for storage and eventual release. I doubt it would ever happen there because of the amount of money invested in tying up IP in the United States, but perhaps a more amenable location like the EU could help with that.
However it might work, part of the promise of the Internet was that information would be liberated, but we see every day how much information gets sent down the memory hole when businesses, sites or services shut down, or how regularly companies abuse IP related regulations to attempt to strangle competition. It would be expensive now, but it would create an incredibly valuable legacy of information for the future, and it will only get more expensive to build such a thing as time goes on.
Great comment. Those who try to excuse book burning cannot be doing it in good faith, this is the acid test.
If AMZN, et al. were burning books in good faith, they would be open about it, point to the party responsible for it, and provide a list of the burned books to allow the community to organize, scavenge and scan the endangered publications. Keeping everything secret is a proof of maliciousness.
Preserving books and providing easy access to the information in them is of utmost importance, this is certainly understood by the public and private entities who could do something positive about it - their actions in the opposite direction should be a wake up call, they're on the wrong side of this issue.
Libraries I volunteered for would coordinate throwing away books late night under dark right before the dumpster was picked up and emptied.
Because people are irrational when it comes to this topic. They say stuff like “book burning” when they see books being destroyed. Doing it in relative secret kept the crazies away.
If you are passionate about this topic, lobby for more funding to store archives in the public good. Expecting private parties to do it for free because it gives you the ick to see books destroyed is not useful.
More books get destroyed each year during estate cleanouts than any AI companies could ever hope to accomplish. Most books donated to goodwill or other thrift shops go straight to the dumpster and might not even have a set of human eyes put on them at all. These places act as sin eaters for folks to leave their trash with.
It would make more sense for government to accept digital copies for any book, and share the out-of-copyright ones. In the United States, we have a Library of Congress who could do it if a law were passed with funding.
What they could do, if they were truly focused on not taking the most destructive path forward, would be to cooperate on a single scanning company that could pool the resulting text for the others on attractive terms, to ensure that once a book was scanned all the others wouldn't need to scan and destroy the same title. That would substantially limit the damage.
Of course, because each company is happy to burn it all down to beat their competitors and being first to any content is most important, this will never happen.
Somehow people are still failing to understand that transferring the scan from a trust, scanning company, or other external party would constitute a transfer of a copy between distinct legal entities that breaches copyright law in an actionable way.
Doing the scan within their own organization does not because it's within a single legal entity and there is no copyright issue involved.
I'm not failing to understand that, I just don't think it's an insurmountable barrier.
You could overcome it through various means, including helping to set up a clearinghouse for bulk rights, implementing a data clean room approach that models could train on in situ, lobbying Congress for copyright law changes especially around orphan works, using Section 108 of the Copyright Act to set up a specific preservation vehicle like the HathiTrust, and other options.
I find it ridiculous that so many in this thread are acting as though AI companies are simply powerless to do anything but buy up and destroy these rare books.
Make a case that it's more profitable to take any of the courses you suggest and you might be on to something but they aren't. All of those things take time and time is money when you're running a business.
I would rather make the case that with the increasing consolidation of media under the big tech companies, regulation is more effective than a business case.
Of course they’re not more profitable, they’re negative externalities that aren’t being priced into a finite commodity.
That’s the whole point: Companies are doing this because it’s easier and cheaper for them, and those doing it now are first-movers who don’t care about how catastrophic the end result will be for the rest of society as they’ll have their training data moat.
It should be obvious, but the things that are most profitable are not always the things that are best for society as a whole. That’s why we have regulation.
That is the problem: they are digesting it. They are not creating a new kind of library, where you could say, show me the text of "How to Fix Your Ice Cream Problems". (An actual book I own) It is not being done for our future reference.
They are digesting it an incorporating into the weights. You won’t be able to get the exact page but you will (if the model is good) be able to get the knowledge out of it in a likely far more concise, relevant, and certainly more widely accessible than the book sitting on your shelf.
> Ideally they'd upload them to Anna's archive too
Not only can they not do that, they must scan physical copies because they are forbidden from using digital pirated copies from sources like this.
Anthropic had a big settlement because they were caught using downloaded digital copies. As a response they’ve ramped up their book scanning and others have followed.
Yeah, the very disjointed way that copyright is applied (I would argue, rights that should also transfer across to digital media but generally have not) is contributing to this situation in the first place.
This is the right answer but the reason they can’t upload the scans is copyright law as demonstrated by Google having to settle with the publishers and allow them to remove their books and limit free access to 20% of text. The AI companies are essentially compressing the information in a huge swath of books that would otherwise be headed to landfill and making them 1000x more accessible. This is unquestionably one of those instances where capitalism is taking money from rich investors and benefiting the 99%.
As if the AI industry is only shredding "crappy" books.
I recently bought a rare book. It was a boat design book written by a famous yacht designer in the 1940s, but it's been out of print for decades and I had to pay $150 for a "fair" copy with missing dust jacket.
I have an interest in older technology and the old ways of doing things (for instance, how do you lubricate the mast of a gaff rigged boat so the gaff jaws don't jam?) I've often found myself reading very old books that have been out of print for a century.
Sometimes I read those books at libraries and I've been the only person to check them out in years (I started doing this back when they still stamped the return date on a card so you could see when it was checked out). Now most of those books have been disposed of by libraries due to yield management software and I've ended up with some of them in my personal collection, but people like me can only save a tiny sliver when most of them are being bought in mass by Sam Altman. When they're gone the knowledge in them is also gone.
We are burning the library of Alexandria and the HN consensus is "those books probably weren't saving anyway."
Once you start to realize just many books got thrown out annually before “AI” you have less sympathy. Libraries are one of the biggest contributors to this because there is simply too many books that nobody cares about.
I suspect folks are over weighting how much knowledge is being destroyed in these books. If someone actually quantified it that would be amazing but as someone who started going to used book sales at a very young age I just have no sympathy. Most books are worthless. I don’t mean that from a text perspective either.
They're not reproducing the information verbatim and with the ability to definitively reference its source. Moreover, not all AI companies are going to get the same books and therefore the same training set, and as companies go bankrupt, training sets are changed, and models abandoned, that information might disappear entirely.
It's not at all the same thing as still having the books themselves available somewhere.
On the other hand, the book can only be in one place and read by one person at a time.
People may not even know the book exists or that it contains the information they want. Obtaining access to a copy may be very difficult, even if it is not particularly rare. The AI is much more accessible.
Then find a way to make them digitally available, verbatim, like Google Books attempted to do. I'm sure with the power that AI companies have they can have legislation amended to make it happen.
That copyright law is a mess doesn't mean that companies that buy up rare books, destroy them, and keep the contents locked up are beyond criticism for their actions.
They don't have to scan and destroy these books, it's an entirely voluntary choice.
> As if the AI industry is only shredding "crappy" books.
Do you, or anyone else, have any source suggesting that they’re buying highly valuable rare books and shredding them?
The kind of $150 rare book that you had to buy from a specialty collector who graded it is in a completely different category. You’re thinking of “rare books” in the historically rare, valuable, and collectible category.
The book sellers shipping off orders of 1000s of books at a time to these facilities are calling the books “rare” because they may only have 1 copy, not because it’s a collectible with a high price tag.
There are a ton of niche books like the one the parent is speaking of which might not be considered hugely valuable by the general public or less specialised second-hand stores, but are both rare and valuable in their niche.
For instance there were only limited production runs for a set of books on South Africa's participation in the Second World War, and getting hold of them is increasingly difficult. But they have ISBN numbers, so they fall within the group of books being collected and destroyed by AI companies, and as a result might disappear altogether, along with the knowledge inside them.
Have you seen AI companies making any effort, with their newfound immense power to get the attention of lawmakers, to lobby for amendments to copyright law to make this exercise a less destructive exercise?
Why would they? They're not the ones upset that they're cutting up books. They fought a court case to allow them to do it a different way that didn't require cutting up any books. They were shut down and had to pay billions, and so they're doing it the only legal way that is left to them.
This is literally what the publishers asked for. If they wanted a different way, they're the copyright lobby. They are making all the rules. They can change it.
Those poor AI companies, they’re the real victims here! /s
There are several things to criticise here, including the US’s terrible copyright laws and the publishers involved in that case who care more about in-print than out-of-print books.
But it’s the AI companies that are actively buying up and destroying books at an unprecedented scale, which is an action they’re choosing do to and certainly don’t have to do. Therefore that’s also where the immediate focus and attention needs to be
> There are a ton of niche books like the one the parent is speaking of which might not be considered hugely valuable by the general public or less specialised second-hand stores, but are both rare and valuable in their niche.
I'll ask again: Does anyone have a single example of this?
Every conversation about this topic has been full of this same assertion, but nobody can ever name even a single example. We're supposed to believe it's factual.
You’re supposed to believe it’s factual because you have investigative reporting by multiple newspapers with good track records, on the record statements by bookstore owners, and evidence presented in court cases like the one regarding Anthropic’s Project Panama.
Notably, one of those piece of evidence from that Anthropic case was an internal presentation clearly stating the intention of Project Panama was to “destructively scan all the books in the world.”
Beyond that, some book sellers have named specific titles, such as “an Estonian translation of John le Carré's The Mission Song, a specific edition of Anne Brontë's Agnes Grey, and the October 1983 issue of Warship monthly magazine” and “everything from 18th-century African agricultural implements to biographies of 1950s racing drivers.”[0]
But getting a full list of all those sold globally is difficult. I imagine we’ll start to see some efforts to try to encourage booksellers to contribute to public lists though.
Weeding (the library term of art for selecting works for deacquisition) is a very import part of collection management. It's actually regionally and nationally coordinated so the interlibary loan network does not throw away the last copies.
It would be nice if your local library hadn't sold their copy but I think your beef is with the libraries, or perhaps with the politicians who failed to fund them in that case.
Actually I expect that a copy of most books probably is still available in a copyright library but it might not be easy to access.
> or if this was some old guide about How to Use Microsoft Office 97.
Materials like that may be very valuable to software / tech / HCI archeologists soon.
I can't get the tools or local know-how to straighten my scythe blade in a country where every cottage had a scythe less than 100 years ago, with the last scythe-native generation rapidly dying out.
Fast forward a civilizational collapse and that M$ Office 97 for Dummies might be as groundbreaking as a Guide to Using Roman Concrete
This feels like a manufactured controversy. What difference does it make to me what someone does with a book after they buy it? It's effectively unavailable to me regardless of what they do. If people are really concerned about these "rare" books, they should lobby the copyright owners to release them online or print more copies.
Do you understand that many of these publishers do not have the original copy anymore, and may have gone out of business entirely? 50 years ago text was not backed up to the cloud.
If there are very few copies and no one wants to even spend $100 to scan it, likely the book wont be bought by anyone in any case and it would rot/thrown away in another 50 years if AI companies had not digitized it.
Which is what has happened over time now that humanity has the ability to digitize them. The problem arises when we are purposefully picking books where we know those scans aren't yet available and purposefully destroying them. There'd be less outrage if any of these companies made statements about guaranteeing the survival of the scans they are taking and the release or planned guaridanship in the case they can no longer maintain the data.
Who is digitizing the books which are rare? Isn't it even more valuable to digitize for them now that they have new buyer which can pay 100s of times more.
Anna's Archive is a controversial example with that as a stated goal. But more to the point most of the people doing it aren't seeking to profit from digitizing these books so the 'monetary value' of them is not the yard stick or motivator. A higher monetary value could actually make it harder for these projects as they can rely on donations of books or even waiting for someone who wants a book they bought so they can spend that on the next one (which is less useful if the next one on your list was bought to destroy by the same person who bought the book you finished scanning).
They're often books tied to a region, i.e. stories about a local fisherman's life and how they overcame adapting a boat to local feedstock, or the engineering that went into building the local train station. That's to say they can have little value to those outside the area but are still a rich read for locals, tourists, and those who move to the region or have a love for a place. I expect we'll see a concerted effort to organise to get these kind of book sent to regional mueseums for scanning now that people are aware they are actively being destroyed and they realise the timeline isn't how long the book can last.
How? Like, do you think a book published 100 years ago has a digital file sitting around a publishing house that might even exist anymore? Re-publishig of old books are often based on scans (the originals were literally pressed into paper by metal, not printed off hard drives), and if the books don't exist, they can't re-scan.
I own some antique Japanese books that have very few copies in existence. You couldn't get them re-printed if you want to if the physical book didn't exist anymore. For a few books, I've actually purchased incredibly high-quality scans that cost me hundreds of dollars because to produce the re-issues, people had to go to museums with high-powered cameras to photograph the pages under supervision of a curator. There's no digital file to print from the publisher. If those books were gone, there's no bringing them back.
I think the term "rare" might be a bit loaded. DOES it mean "only a handful in existence" or more like "1000 copies?" And at any rate, at least by scanning the book they're theoretically making its
contents available to the public. What are the other people who own these books doing besides having them sit on a shelf?
"Only a handful" is a precious or extremely rare book. 1000 copies of a book total would certainly qualify it as rare. That's not a lot of copies.
> What are the other people who own these books doing besides having them sit on a shelf?
Not destroying them. By continuing to exist, it keeps the possibility that the books will eventually go to someone else. Maybe even a library that specializes in rare books.
> By continuing to exist, it keeps the possibility that the books will eventually go to someone else.
Good news, there's hundreds or thousands of copies that continue to exist. Anthropic, Google et al only need a single copy.
And shockingly, for all this talk of rare books, we don't see the other owners of said books stepping up and offering to have these books scanned non-destructively. I won't hold my breath for them to do so in the short or long term.
FURTHERMORE, the books aren't being scanned for no reason, they become part of new model training data. Meaning that whatever insights these books contain might be available to users in the future. Again, more valuable than all the copies rotting away on someone's bookshelf that might be scanned and made available some day.
It's hard to describe what happens to the book's contents, or determine whether it would be more valuable intact. Insights are the result of a person interacting with a book, they aren't like juice that can be squeezed out mechanically.
I care a lot about the high-illustration picture books of the 80s. Since they have ISBNs, they're presumably on the list. Does the LLM ingest the artwork? Even when it comes to text-only novels, it isn't making the experience of reading the book available.
This is a problem whether it's AI companies buying them or three random dudes. The real solution is to get the copyright owners to keep distributing copies, or to change copyright law.
Plus, the physical copy of the book has additional value in terms of forensic verifiability. You can prove the originality of the text and the absence of alterations, and can easily track changes between editions.
Digital copies can be altered at whim, as the only means to provide any sort of tracking or verifiability is through another external software system that itself has to be trusted.
Replace bison with some random animal no one's ever heard about, and photos with literal clones of the animal, and you've got a much better metaphor. Not to forget that in this world, brand new animals are created every day.
As long as digital or physical copies exist, you can. Regardless, let's be real, nobody is destroying 17th-century books. The "rare" books we're talking about here still have copyright owners who just don't bother printing new copies due to a lack of demand.
> What difference does it make to me what someone does with a book after they buy it?
I think this is a bit of a myopic take. It's like saying it's my property I can do what I want, yet there exists designated historic homes or neighborhoods that are deemed to have cultural value where modifications do in fact need to be approved.
I'm with you on the "rare" part. If people are thinking about 70+ year old documents or ancient manuscripts, I doubt that's what AI is being trained on and is being destroyed, but it's reasonable that people find _that_ idea distasteful.
You can say it's manufactured but if these companies ignore this criticism, it's just another way AI companies are committed to losing the public.
> It's like saying it's my property I can do what I want, yet there exists designated historic homes or neighborhoods that are deemed to have cultural value where modifications do in fact need to be approved.
Yes, and when you buy them with that designation, you know what you're getting into. The problem occurs when you already own it, and some group is trying to get it labeled as historic, which will add to your burden and limit what you can do with it.
When I was in a small town, this was actually weaponized. A hotel owner was trying to get another hotel categorized as "historic" and had rallied a lot of people behind his cause. He had a case - the hotel did have some claim to being the "first" in some category or other. But really, he was doing it because it was a competitor. The "historic" hotel owner had to spend a lot of money to fight the cause, because being labeled historic would prevent him from performing various upgrades, making the hotel less attractive to customers (he was already not getting many customers).
> We’re not revealing the titles of the books included in the shipment we tracked, but they are rare, meaning there are not many copies of them in circulation. Sometimes that’s because not many copies of them were ever printed, and sometimes because they are in a foreign language not many people speak.
> The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order.
It would seem the redaction of the rare titles is a way to avoid de-anonymization and subsequent harm to the business of the seller who agreed to place a tracker in one of the books. That being said, maybe they could have chosen a better methodology which would have allowed the disclosure of the title, although ultimately I’m not sure the title matters too much outside of their claim they were “rare”.
The issue is if rare is simply "not many in circulation", that's very very different from rare being "it is a unique artifact". The former kind of rarity doesn't really present an existential end, whereas the latter does.
My neighbor self-published a book, printed I think 100 copies at his own expense. It's literally a rare book. I can't imagine he nor anyone would care if an AI company bought a copy, no matter what they did with it.
Maybe these "rare books" are first edition Mark Twains, and. maybe they're unwanted books that would otherwise have gone to be pulped. The distinction is important and by not giving any evidence or even a qualitative claim about the types of books, 404 media is being pretty weak here.
I don't think HN commenters need to be the arbiters of which books are beneficial to society.
And beyond just benefit to society, don't forget how subjective the decision about what constitutes a "unique artifact" is. That random product manual from 1966 becomes a near-priceless relic to me if it lets me repair and use the sewing machine handed down from my grandmother. Not necessarily the ink printed on paper, but the information contained within.
The above is based on a true story involving archive.org's Manual Library. The idea that even such esoteric and forgotten information would get hoovered up into the walled garden of Amazon's AI and then destroyed in the outside world, such that I have to go pay them to access even a facsimile of it, is frankly disgusting.
>That random product manual from 1966 becomes a near-priceless relic to me if it lets me repair and use the sewing machine handed down from my grandmother. Not necessarily the ink printed on paper, but the information contained within.
On the other hand, now that the manual is in the AI, you can just ask the AI how to repair the sewing machine.
The information contained within has not been lost, and in fact has become much more accessible.
> you can just ask the AI how to repair the sewing machine
Like so often in its usage, the word "just" is carrying immense weight in this sentence. Please see "hoovered up into the walled garden of Amazon's AI [...] such that I have to go pay them to access even a facsimile of it"
> has become much more accessible
This accessibility rests on a number of assumptions, not the least of which is Amazon's (of all companies) charitable good graces in offering access to their AI at an affordable price. It also assumes that the model can accurately regurgitate the text without hallucinating about other, similar machines, and that it can faithfully recreate any diagrams.
> That random product manual from 1966 becomes a near-priceless relic to me if it lets me repair and use the sewing machine handed down from my grandmother.
That ain’t a rare book.
I would also be shocked if you couldn’t find a way to repair and use this 1966 sewing machine… so it seems like a red herring.
Thanks for your opinion. I see you disagree with my first point.
> I would also be shocked if you couldn’t find a way to repair and use this 1966 sewing machine… so it seems like a red herring.
"Repair of the sewing machine is obviously simple and thus left as an exercise to the reader."
Seriously, though, I'm trying to say that you can't make that call for me. You can think that, but I know from experience repairing old machines (with more or less sentimental value) that having the original manual can be the difference between fixing it and irreparably damaging it. I could also, in turn, suppose that you probably do not actually collect and preserve rare "unique artifact" books, and therefore your argument about them is a red herring. But then we would both just be casting aspersions.
And even then, the "rare" qualifier is not needed here. What Amazon/LLM companies are doing is amoral. This is intellectual piracy (not in the "copyright infringement" sense) at the highest level. Stealing and centralizing the accumulation of human knowledge to eventually rob us all and put all power into the hands in the hands of a handful of people, who are not benevolent.
It doesn't really matter. The article suggests that they are selecting and tracking books by ISBN, which means: books that have been published or at least reprinted in the last 50 years or so. And they are trying to get as much of that set as they can, regardless of quantity, quality, or any other consideration. Which means that older books printed before ISBNs became common may be relatively safe, at least from Amazon. And some books just don't come up on the used market very often.
(small historical irony: when Amazon first started selling books, they used the Books in Print database, which included a lot of books not actually in print.)
Rare as in your grandfather's John Deere manual from 1982, not rare as in a test print run of The Great Gatsby. The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
> The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
Citations? Also what exactly are these 'normal means'? As a bibliophile who loves scouring used book stores for out-of-print titles this is a topic I'm very interested in.
Normal means is throwing it in the recycle bin. Especially for stuff like a 1982 John Deere manual. I've never donated an old appliance's manual to the library. Have you?
Amongst my friends, I'm one of the rare folks who donates books to the library. Most people just trash them. And I know the library only wants them to try to sell them in their book sales (or online) so they can get money. Almost nothing one donates to a library actually ends up on the library shelves.
EPA estimates that hundreds of thousands of tons of books are landfilled or recycled every year. That's just the United States. It seems likely that the worldwide figure is in the millions of tons.
The scariest thing for me about companies not caring even minimally about conservation is that when AI gets more powerful than humanity (which is clearly a when not an if, even if there's a lot of uncertainty and differences in opinion about the time horizon here), I want to hope that AI will care more about conservation of human people.
So far from how I see how powerful organizations work, I'm not as certain as I would like to be.
> AI gets more powerful than humanity (which is clearly a when not an if, even if there's a lot of uncertainty and differences in opinion about the time horizon here)
People say this all the time, but so far nothing has convinced me it's true.
LLM development has more or less plateaued, and the current boundaries are very real - energy, resources, capital.
At this point we're talking about marginal improvements against the same asymptotes of all technological innovations.
So, if a corporation can suck up entire books to teach their machines how to think using the information from those books, can we (all humans) join a single corporation that provides all books to its employees? Just need one copy of each and we'll make that copy digitally available for our employees so they can learn from and utilize the knowledge from the books. It's not copyright infringement, they're employees.
That's a bit of a wild interpretation of copyright law.
I mean, Anthropic isn't going to fight it because it lets them do the thing they want to do, so I can see how this never gets beyond the court that allows them to do the thing they want to do.
But would this argument would have flown in the past?
It wasn't even attempted in Sony v Universal. Or any copyright suit up until this point. That doesn't smell funny to you?
Sure, it's a bit wild, and no, the same outcome may not have been reached if the case came up earlier. And if it makes it to the Supreme Court, who knows what they will turn it into.
The flip side is that Alsup (the judge who wrote the opinion) is probably the smartest district court judge we have when it comes to technology, and one of the people I'd trust most to come up good decision.
i would assume that would come into play if they were uploading scans of the books? there must be some gray area where training like this doesnt apply to that. or ya know, just do it and face the consequences later because you already have the data and know that ai obsessed government will just shrug their shoulders.
Not to my understanding. To begin with, it's far from given that "rare" books are all covered by copyright. But if they are, it's at best murky: whether you destroy the original doesn't really have anything to do with what you're doing with scanned contents. The scanned contents themselves may be inherently a copyright issue, regardless of destroying the original. The actual trained model has separate arguments more in its favor, so if no scanned contents exist - IE the data is read once for training and not stored or saved, they have a better argument. But in that case the destruction is totally disconnected from copyright, as they'd be totally okay to rescan the material.
Bias disclaimer: Amazon is my current employer, but I don't work on AI or anything else mentioned in the article.
Yes, this is a result of copyright laws. The other commenters are wrong/uninformed.
If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc. There are essentially zero advantages, other than it is what is required under US copyright law (or at least, it is what their highly paid lawyers believe is required under US copyright law).
A judge a while ago decided that as long as the physical copy is destroyed, and "transformed" into an electronic copy, you can do the upload. But if you preserve the physical copy after scanning it, you are in violation of copyright because you "copied" the book.
That's literally the only reason they are trashing them. It's a legal requirement.
Ah, so I can buy and scan a dvd, destroy it and then legally distribute the legal copy via torrents. Good to know, because that is what the LLM thieves are doing.
Distributing an exact copy of the text would be illegal; if you could get an LLM trained on a book to output the exact copy of the text from the book then that would be illegal as well.
They still made a bunch of copies and reused it to train multiple models after making the first copy. Unless they copy and destroy a book each time they use it as as training sample
Can you find a citation for this? I have heard this claimed rule recently from other people, and I haven't seen this decision (nor do I know what level of court or jurisdiction it might be). This is not a rule that I heard many years ago when working on and adjacent to copyright issues (including book scanning!), although of course the issue has been newly litigated again recently, so there may be new interpretations coming out.
Edit: Someone else linked to an order in Bartz v. Anthropic which appears to emphasize that destroying the original copies improved the defendant's position with respect to the fair use analysis. Is that the decision you're thinking of?
Perhaps partially? I assume that cutting the pages out of the spin makes them easier to scan at least partially. That said I have no insider knowledge of this type of operation so I don't know how much easier that actually makes it.
But ultimately, it's a moot point, because the legal requirement means the books must end up destroyed. Even if the people at Amazon wanted to scan the books in a way that required no destruction at all, it's not currently (legally) possible for them to do so, so they might as well take the easy way out today.
Cutting the pages out makes them machinable. Non-destructive scans involves gently turning pages, and paying a lot of attention to the state of the spine. Destructive scans involve guillotine cutting the spine off, scanning the covers by hand, putting the pages into a hopper, clamping them in and hitting a button. While that book is scanning, you're already cutting the spine off the next book. If the machine jams, try to work the jam out gently, scan the pieces, and let the computer stitch it together.
Those laws are lobbied for by large corporations, these are not just laws that exist outside of that context. They can also be changed, or Amazon could just incur the fines.
Large corporations will move fast and break things when it’s convenient; they don’t care much about the law - just about profit.
I support physical media and doing whatever you want with said physical media. If you want to overpay for a copy of Sharepoint 2007 For Dummies and destroy it, knock yourself out. Just because media has been printed and is "rare" doesn't mean it has any practical value.
You could easily write an article about how Goodwill and the Salvation Army dump millions of "rare" books that no one would ever buy for $0.50. And thrift stores will dump multiple copies of said books. AI companies would only ever need one!
The only thing I disagree with is AI companies feeling entitled to freely use every piece of copywritten work without restriction, but that applies just as much to web scraping, pirated books, etc.
Why is the textual data bound up in these paper books worth the trouble?
These LLM training runs have already ingested essentially the whole public internet. What marginal value is to be gained from scanning and destroying obscure books?
For me the biggest functional issues with LLMs don't seem to have any connection with "I wish they had read this obscure community cookbook from 1946". Is that going to get Claude to stop saying "honestly" to me? Is it going to get LLMs to stop making up sources that don't exist? What is in it for Amazon or any LLM company to chase more obscure data like this.
I'm not an expert in LLM training, but I think we can all agree that the writing on the Internet is generally very low-quality and surface level compared to the depth of books. Most books in the past were even edited by a separate person from the writer!
Granted that that is indeed the case, surely the quantity of high-quality writing available online (including things like Project Gutenberg) still dwarfs the quantity remaining in obscure unscanned books.
Humans manage to be pretty intelligent with only reading perhaps a few thousand books in their lifetime. It seems unlikely that AGI will appear but only once it has read that last out of print 1983 book on knitting patterns.
Frontier researchers have found that dumping more and more data into training is effective at improving LLM capabilities. Nobody has a gears-level understanding of how training on some particular kind of data leads to some particular behaviors, so they generally take the attitude that more is better.
I have to agree; even if they are getting regular 20th century out-of-print books, is that going to add a significant percentage to their training data?
I can only think of it being a 'low-background steel' situation where they want to locate original, non-digitized text for validation or knowledge bases.
We really need to be thinking about the long-term implications of what Amazon, Anthropic, OpenAI, and others are doing here.
Because it's not just one company doing this, it's dozens if not hundreds of them, all with huge budgets and all competing over the same dwindling supply of older books. What those talking about library disposals and similar measures are missing is the sheer scale that destructive book scanning for AI ingestion is operating with. It's unprecedented.
It's not impossible that in a few years the number of older physical books for sale anywhere will plummet to almost nothing and that in many cases that'll include the last copies anywhere of particular titles. Second-hand book stores will close, and a ton of niche knowledge might be lost forever. It'll also all have been done in relative secret, without any transparency or public record kept of what was lost.
Worse, it's something that can only be done once. If our societies don't do something about this now there isn't going to be a chance for a do-over. Once a physical book is destroyed it's gone forever, and we'll be lucky if there's a digital copy left. But even then digital copies don't provide the same level of forensic verifiability that physical copies do. Future researchers looking for material that might've been contained in books like these will be out of luck.
If our societies do nothing to stop this, we'll all be poorer for it.
I don't understand why Amazon still bothers to blow money on their in-house AI effort (outside AWS Bedrock). What have all the scraping, scanning and training led to? Nobody is using their models.
A lot of the rare books seem to be books that are most relevant to specific professions within specific countries. I would expect university or national libraries in those countries to have copies of those books, but they often do not digitize them because controlled digital lending is seen as a legal grey area at best. So these ai companies are doing digitization for them and not publishing it, duplicating work for these libraries once they do decide to digitize.
Many other rare books are just rephrasings of other books on a given topic and this can be done with synthetic data generation now. Same goes for websites
Honestly this feels like a US exclusive problem. Like shipping from overseas would be way too expensive for the cheap books they want. The Chinese just use Anna's archive instead
* They're only buying a single copy of that book as they only need one to scan
* If the book was public domain (or should be), then there should be an effort to "democratize" that data into a public commons of intellectual property?
Would that an enhancement to the Library of Congress or such?
Why do you think for even a moment that they would buy only one copy, rather than every available copy of something rare and difficult to get? Anyone doing this stands to gain from destroying the only copies of something that has scarcity, to stop competitors from ever getting access to it. That's a serious concern.
Why would they do that? What is this 'should' to a company that would do this in the first place? Are you not suggesting the opposite of what they aim to produce?
That seems really speculative. I think you're assuming the maximum possible evil intent for no reason, mostly because you hate AI and anyone involved in it.
Years ago I saw an article on this. For Google Books, they had two processes.
The first was destructive. This was for mainstream books currently being published so they had no value. It's (I believe) where you cut off the spine and scan the pages.
For rarer books, there was a non-destructive process. Basically the book was opened to each page and scanned. This was slower but didn't destroy the book.
I don't understand why these companies haven't just licensed the scans Google has already done. Why is each company doing this rather than just scanning the books once and sharing the scans?
> I don't understand why these companies haven't just licensed the scans Google has already done.
Because for the vast majority of the books, Google doesn't have the legal right to license those scans. It would be legal if the books were out of copyright, but despite the connotation of "rare books" those generally aren't the books we're talking about. Further, in the cases where Google didn't destroy the original of the book they scanned, their scanned copy may be considered infringing under the new standard, so Google doesn't want call undue attention to what they have.
yeah, I think most of this really stems from people streatching the use of "rare books". The cool rare books google scanned one page at a time were borrowed from a library because they are rare and or valuable and usually old (out of copyright by a long time).
If I remember correctly, it wasn't a distinction between common and rare books.
It was a distinction between library books and others. They partnered with libraries, and obviously, libraries didn't want their books destroyed, so they devised a non-destructive book scanner. Some of the books they scanned were indeed rare, but rare or not, you don't destroy books you borrowed from a library!
6 Feb 2021
At the Internet Archive, this is how we digitize a book.
We never destroy a book by cutting off its binding. Instead, we digitize it the hard way--one page at a time.
Wonder the additional cost to add automated page turning. More than Bezos can afford, impoverished chap he is.
While we can quibble (and are) over the details of this particular occurrence, I'm beginning to take the opinion that used/rare booksellers need to implement some form of "KYC" to help guard against the epistemic threats posed by this practice.
The booksellers are probably so ecstatic to get some value out of books otherwise destined for the landfill that they wouldn't lift a finger for "epistemic threats".
According to the article, the scanning facility has been there for a while, so it may well be that they have been destructively scanning books for a while, even before AI. It makes sense for the major book retailer to have digital copies of every book so that it can begin selling those as soon as it can negotiate the rights (if there are any).
In any case, all the indignation about destructive digitization misses the point that rare books takings space in a warehouse for years without being bought will eventually be destroyed anyway.
Something I've been thinking about is that a physical book is a physical artifact, and it's age and authenticity can be verified by treating it as such.
In contrast a digital book has no such marks. It's impossible to tell if it was redacted to hide inconvenient passages, or even completely rewritten or fabricated (which can now easily be done at scale with AI). If the physical book is a primary source then the digital version must be considered a secondary source: A potentially biased retelling.
For this reason, a digital book can never be a perfect substitute for the physical book it was created from.
I'm less concerned with how these rare books didn't rot in large lots of unused books and more concerned with whether or not this helps preserve the content longer.
I don't expect an amazon.com/checkoutrarebooks page but I also don't really think the legality of sharing the content has been much of a concern for these companies either. Ironically, one of the few things Meta got in trouble about with building their AI was helping share the training data.
I seriously doubt any of the companies doing this will preserve the scans for very long. It costs money to store stuff, and none of these companies are have any concern for anyone who isn't them, so they're not going to spend the money or lift a finger unless they work out a way to make it profitable.
archive.org estimates it costs $2/GB to store data in perpetuity (well, at least until decades of storage scaling trends stop of course). That's probably less than it costs to acquire and scan the data in, I'd be surprised if they just tossed it out at the end when it's so cheap to store. The same thing happened at the healthcare data lakes I worked on where the trend switched from the usual "how long are we legally required to store this information" to "how long can we legally hold this information".
So... do the materials that it is printed on make it rare? Or is it the combination of ideas, paragraphs, and words that comprise the text make it rare and therefore valuable? If they're releasing digital versions of this text, i feel that is better off anyway.
I look forward to the day that AI companies pay large advances for new books.
It'll take a while before publishing collapses due to the availability of the same information via an LLM, but once this happens new books will get a lot more expensive for AI companies and a lot cheaper for everyone else.
I've been lead author and co-author on several books on programming. My first book sold tens of thousands of copies even though the writing was not just trash but a crime against humanity against anyone who had to read it. This was back in the day when the source code for the book was provided on request on a floppy, which was actually quite rigid because it was a Mac floppy.
My writing got better and better, and I got better publishers who actually hired editors, but the books sold fewer and fewer copies. Even crappy self-serving poorly written stackoverflow posts are often good enough. Then LLMs killed stackoverflow. Is that real ironic or Alanis ironic?
But I'm not holding my breath waiting for a book deal from OpenAI.
> My first book sold tens of thousands of copies even though the writing was not just trash but a crime against humanity against anyone who had to read it.
I’m sure I never read your book, but there is something cathartic about reading an admission like this. I have no doubt people got value out of your books, but I distinctly remember as a kid saving up to buy a couple expensive programming books from the bookstore and then being sorely disappointed in the way they were written. At the time I thought I was too young to understand adult writing, but when I went back to the books later as an adult it was obvious they were just written by amateurs. I still cherished them and learned a lot, but I will always remember the struggle of trying to follow along with what was probably some first-time writer’s attempt to learn how to write as they went along.
> But I'm not holding my breath waiting for a book deal from OpenAI.
Was this title eligible for class action status in the recent Anthropic scanning case? I know people who popped up on the list decades after writing obscure or forgotten technical titles. The estimated payout is $3k per title, usually split between publisher and author.
There was some specific thing about which ID number, not ISBN as I recall, your book had to have to be part of the settlement. Some of my books qualify some don't.
Copyright office registration. If you or your publisher didn't do it, you don't qualify. Some publishers, even big ones, were exposed as neglecting this basic step when the Anthropic case was settled.
They are only destroying the books because they are required to by copyright law. They obviously wouldn't do so if they were allowed to merely copy the book and preserve the original.
If it was $1 cheaper to destroy the books even if they didn't have to, they probably would. Storing books is expensive and selling it on means it could be picked up and scanned by a competitor.
“And, the digitization of the books purchased in print form by Anthropic was
also a fair use but not for the same reason as applies to the training copies. Instead, it was a
fair use because all Anthropic did was replace the print copies it had purchased for its central
library with more convenient space-saving and searchable digital copies for its central
library — without adding new copies, creating new works, or redistributing existing copies.”
I’m not a lawyer, so it’s possible I’m missing something. But it seems to me like the ruling here implies fair use only holds so long as no net new copies, digital or otherwise, are created. If that’s the case, then it necessitates the destruction of the original.
Not true. To the extent that it is fair use to digitize a work for various purposes, it is also fair use to keep the original. It is only if they wanted to resell the original that they would have to delete their digitized copy. The reason they are destroying them because removing the binding is the most efficient way to scan them, they have no use for the originals after they have been scanned, and don't want to spend money storing them.
No, they aren't, for training AI, at least not based on anything but pure speculation.
The recent trial court decision that keeps being pointed to to support that:
(1) Found that for training AI, digitizing and copying works was fair use, period, with no requirement to destroy.
(2) For creating a centralized digital library for general use, digitizing works while destroying the hardcopy was fair use even with no intent to use them for training AI.
It seems like most bookshop owners today are hoarders-turned-entrepreneur. They rent a storefront, adopt 2-5 cats, and they set up shop with their hoard of books so that they can hoard on a commercial level that would never be possible in a private home.
It also seems that American society has been swept away by the worship of books, rather than literary appreciation or promotion of education. "Banned Books Week" as the primary exhibit here. A tug-of-war over books that are supposedly "banned" when the verb itself has been twisted beyond recognition. I see posters and television shows and public service announcements that promote "Books!" and "Reading!" for no other purpose. Many books are trash and they will fill your head with garbage as sure as social media can, so why the indiscriminate worship of books?
It is conspicuous that, aside from established booksellers, and perhaps the LFL owners, many of the people crying out and moaning over book-destruction seem unwilling or unable to actually take those books in and store them. That is the point, that these books are unworthy of taking up storage space, which is an ongoing cost, and maintenance, and therefore, it is fiscally responsible to sell them, scan them, and destroy them. If you want to dedicate a room full of shelves to books that nobody will ever read, then go ahead and buy out your local bookseller! They will not stop you or cancel your order! Tell them you are saving those innocent books from the Big AI Bugaboo! They'll grovel at your feet!
Also now, I'm seeing small booksellers who feel "suspicious" or "skeptical" about large book orders. There was one Down Under who said "oh we got an order for over 70 and we can't physically handle that!" I feel like it's becoming a "shut up and take their money" situation. These booksellers, as professional business owners, should know that they may never liquidate stock, and this is their golden opportunity to simply get rid of some of that excess hoarded inventory.
But if you're merely cosplaying as a bookseller, and deep down you're a hoarder and a worshipper of books, you may really be reluctant to do transactions or earn money that sells away your books.
In more liberal sense, because you might be destroying something unique that later generations might actually like to see, the life rule that states "Don't be a dick" probably applies trumps even first sale doctrine.
If there was a public interest in these books then it was already attached, and it was an outrage for the sellers to exclusively possess these books, and it was an outrage to sell them to any other private party.
Does anyone else get pretty tired with 404s style of writing? I hate to make this comparison but it feels like a Fox News for the left.
I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it. It is an absolute shame for books to be destroyed but what does it really mean to be rare here? I know they kind of tried to differentiate but it sounds like this could be John Doe’s self help book that never sold well. If you ever are connected to a library you will start to realize how many books simply get thrown out or sold for nothing because nobody wants them.
To me the problem is partly the copyright law. I think it’s. Hard problem but I always lean more towards books having short copyright shelf lives and making it legal for digital copies to be shared after which I think would eliminate I good part of this problem. Not to mention 99% of the books published are probably garbage but that is highly subjective.
So it is bit of a meta rant but I think there a couple holes to go down that could be extremely interesting but they always write these informationally light articles. Like scrolling through a NYT visualization for just some shipping datapoints. Don’t really dig deep on anything and then end with a trust me bro these are rare books that Amazon is destroying for AI when I cannot be that upset with the amazons of the world. There are a lot of reasons for a business to digitize books, most books are worthless and it makes sense to cut the bindings for scanning. I would rather talk about how could you fix copyright to make this less an issue but is it even an issue with how many books get thrown out?
Not really. It's not to my taste, but I've found their journalism to generally be lefty but focused on important events.
Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
The difference is huge between 404 and Fox. Fox is out here trying to tell people there's a trans agenda, and that Biden was a lunatic leftist. They are just making up stories and publishing them because they know their audience engages. 404 definitely make editorial choices about which stories to pursue but I've largely found them to be grounded in real depictions of stuff that is happening.
Again, what rare books? Like they pointed to a book with low published volume or foreign language is “rare”. If you have ever been to annual library book sales you start to realize just how many books get thrown out. I can go buy a stack of books from the 1880s and 1890s on eBay right now for $25. $4 each. Most people don’t realize how worthless books are, even “rare” ones.
I even imagine that the market price for these “rare” books is helping filter out anything truly valuable and rare. It just reads as a rage bait tmz article. The quantity of used books including “rare” books that get thrown into the dump is astronomical.
I worked in a bookshop. I know all about this. I've seen pallets of books with covers torn off so they can be reported destroyed.
Here's a bookseller describing some of the books, which he believes are going to middlemen who are playing a sort of arbitrage with the AI companies by buying obscure titles, reselling them, and tossing away anything that doesn't sell. https://charliebecker.substack.com/p/is-an-ai-company-buying...
The titles mentioned there are not rage-inducing: "The Insider’s Guide to Metro Denver from 1995" and "How to Use Corel WordPerfect 1991". 404 writes and article about book destruction, and then explicitly neglects to mention a single title.
I suspect that's because folks wouldn't react the same way if they realized it was computer manuals for Windows 3.1. Even more obviously, Amazon doesn't want to be paying a lot for these books, so it seems outlandish that they'd be buying valuable rarities, since the booksellers would know the worth of those copies.
Good reporting on this would have included sales prices, volumes, and titles.
They destroy regular books. This was an intentional order from a rare book dealer. There is no reason to believe they use the same destructive process for books they should know to have inherent value and aren't purchasing in significant volume.
> Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
No, destroying collectible books would be a shame, not just any rare worthless books. But these are not collectible. The article tried to dance around it by saying maybe some books have a sentimental value to someone somewhere. But that doesn’t mean any library or collector wants it. Don’t fall for manufactured outrage!
Edit: Here’s an example of an extremely rare book. My great great grandfather published a book of sermons around 1920. That book has zero value to anyone other than my dad. Would I be outraged if it ended up at someone’s estate sale, then a used bookstore, and then an LLM consumed it to learn to read? No; I would have expected it to have been discarded by humans before the LLM even got to it. Most of what we leave behind is discarded.
This was my issue as well. It seems like they tip toed around it in one paragraph saying usually these are books with isbn numbers so not truly collectible or very rare and then in another paragraph plainly stated that it could be foreign language or low volume books that are rare. Rare expresses a different meaning to me and I think it amounts to exactly what you said, manufactured outrage.
We recently had a post highly upvoted here that collected bus tickets from an earlier era. It showed the unique moment in time where such fares were highly detailed, unique pieces of art. Ultimately destined to be single use and largely extremely pedestrian.
We've also seen a website that collects old paper restaurant placemats from across the ages. Literally disposable, zero value items that were made to be discarded by the thousands.
And yet, when someone has an interesting idea that ties them together, makes us recollect or think about our past in some novel way, these useless uncommon things suddenly become quite interesting.
A book on sermons, o its own, from the 1920s is maybe not that interesting. A book of sermons selected from each decade? A collection that compares regional books of sermons? A compare and contrast of the 2020s and 1920s? I can imagine many interesting thesis where a book like that becomes interesting because of the context of other books that are juxtaposed to it.
A book on Detroit motorways and bus schedules from the 30s isn't interesting or valuable on its own. But when you contextualize it, suddenly it might be a way to understand our history, our path through development and redevelopment. Connecting our present moment to the past.
Yes, I collected a couple decades of Muni fast passes too and finally gave them to an artist to destructively turn into some project. I was sad that my “rare” tickets probably won’t be kept forever, but I’m glad that somebody got some use out of it. I feel the same way about LLMs. I’m glad that they are learning from old published materials. In my opinion, 404media is cynically inciting anti-tech anger; they don’t otherwise have any interest in the preservation ecosystem.
You could donate that to the church archives of his denomination or the religious studies department of a university or at least some sort of local historical society.
I was in charge of a "lending library" ministry at my church for a few years. Let me tell ya.
We got started when another ministry moved out, and sort of from zero. My pastor's clear instructions to me were: make sure everything we carry is doctrinally sound.
So we inherited several full collections of books in rapid succession. Some had even belonged to priests and religious. Those gave me a fascinating time, because I could basically rubber-stamp every title that a priest had in his personal collection. But slowly the balance began to tip into rather esoteric volumes that normal laypeople couldn't really use. Literally books full of sermons and other arcane subjects!
I was tasked with discarding/recycling all the rejects. There were tons of rejects, believe me! So with every session when we had boxes full of donation, it was imperative to cull the bad stuff very fast, shelve the rest, and then find somewhere to dump the trash. The manager was encouraging me to recycle, or at least not tip them all into the Dumpster, but it turned out to be a logistical nightmare to find anyplace that would recycle books like that. Having no vehicle, I had to continually figure out ways to cart around heavy loads, just to get them out of church and into the trash somewhere. That was the worst part of my job.
Now it was clear that books were not a very hip or current medium, but there were plenty of elderly parishioners who did appreciate the resource and did compliment my work, but our church was not free of prejudice or judgementalism, and let me just say, there was an angel or entity whose purpose was only to jumble all the books while I wasn't watching, and leave a deliberately unorganized mess for me to confront every week. This made the task distinctly Sisyphean, in addition to the need to constantly discard rejected books.
I finally threw in the towel when large boxes of Spanish-language books were donated; there was no way at all for me to vouch or determine their orthodoxy, and the shelves were full anyways, and I was just tired of propping up a legacy ministry anyway. But it really drove home my opinions about books, hoarders, and that is why I have no troubles with the way books are currently being treated.
So it turns out your organization didn’t need so many donated books. The ideal outcome for your discards should still be sent somewhere to someone who wants to read or possess them, or barring that, to be scanned so their contents are at least preserved somewhere. It belongs in a library, even if incorporeal.
Something can be "rare" while also having no value. One only needs to take a look at their local Facebook Marketplace listings to see this in action.
Rare invokes images of limited edition runs of well loved books, when in reality it's probably extremely outdated software guides, how-tos, technical manuals, etc.
> I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it.
I was going to say a very similar thing, and this is something I strongly disagree with when it comes to HN's moderation, and it goes like this:
The product is the outrage.
And HN should know better and mods should actively discourage, warn and prevent accounts (who karma farm, among other things) from even being able to post rage-bait articles. These aren't "hacker curiosities", they're just insipid bullshit. We wasted time and learned nothing.
I agree with you. And it sucks, because 404 writes about interesting topics! But they're so negatively polarized it often obscures what's actually interesting.
Any bookseller worth their salt will ensure that a truly valuable book will not languish on their shelf for 1 minute longer than it takes to find a buyer willing to pay a fair price.
If AI-scanners are somehow bid-sniping bona fide collectors and wealthy aficionados, there may be cause for concern. But that is most certainly not happening here.
How do you have confidence in that? These buyers are reportedly exceptionally non price conscious.
Personally when seeking out rare books with few copies in existance and extremely scarce availability, I have often found listings that have lasted for quite some time. Not all rare books go to auction. You might be thinking only of some extreme of notable works and not a wider spectrum of desired but scarce publications that does indeed exist contrary to your confident assertion.
Why would that be a reasonable assumption? This is so blown out of proportion. The imagery invoked by the narrative is one of huge corporations destroying the final copies of literary treasures. That's just not happening.
I believe that is the rational default assumption, yes. Presumably there is some known list of books that they purchased and ran through these scanning machines. If there are true literary treasures that are genuinely hard to get access to on that list, then I would join the ideological crusade against.
Where is that list so we can be sure? On some internal systems we can't access. So I'd prefer rare books to not be destroyed if there's a risk that it could be the last remaining copy.
follow the discussion around this on X, i think just a few minutes of research on this topic / reading past the headline you'll find out that this is pretty much a nothingburger
It is interesting to know! My assumption would be it's technical or scientific publications, maybe things like the Springer back catalog (assuming they've not already licensed these). Some of these are "rare" as in they are essentially published PHD theses with very little in the way of sales / print runs.
Of course I could be wrong and they are destroying 14th century monastic scrolls or out of print Mills & Boon editions.
To me, the hate against "destroying books" is kinda silly. The reason we're against destroying books is because the Nazis did it to destroy information.
This situation has absolutely no relation to that. This is the exact opposite, and the only reason this information isn't available publicly and is in risk of getting lost is copyright law.
TLDR: Amazon isn't the nazis in this story, copyright law is.
Making the verbatim content of the book lost forever is not equivalent to burning them? I'd even wager that the CO2 emissions of the scan,shred, and train process is even higher than burning.
> Making the verbatim content of the book lost forever is not equivalent to burning them?
The verbatim content are the words, not the paper.
Books are lost all the time because the last book ended up in a landfill. But if an AI lab digitizes it, now it's stored in an extremely redundant storage lake in a datacenter and the company has huge incentives to make sure they don't ever lose that data.
So is it moral of me to destroy a copy of "Windows 95 for Dummies"?
What if a person with small children and an elderly, incontinent pet with a penchant for peeing on books wants to buy it - can I sell this book to such a dangerous purchaser who might destroy it?
I love book as much as the next person but the hyperbole about "rare books" is absurd. Nobody is buying the Gutenburg Bible and destroying it for AI. The books in question are certainly not rare enough to be in museums - without titles there's no proof these are anything of real value.
But rich people want to violate copyright now! So judges ruled that it's now legal. But, because "the law is the same for everyone" they needed some excuse. NOT because, you know, obviously this demonstrates that very rich companies get to violate the law and you get hit by $30000 per infraction when you do it, when Anthropic ... doesn't even have to stop violating copyright when they do it.
Anyway ... this case is:
Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA, U.S. District Court for the Northern District of California, decided by Judge William Alsup
“the purchased print copy was destroyed and its digital replacement not redistributed, this was a fair use.”
Obviously this violates precedent, I had an internal LLM (probably a frontend for Claude or ChatGPT) find them (just like the convictions for file sharing in the 2000s required counter-to-the-law reasoning by judges, fair use was almost never accepted as a valid excuse, even when it obviously was, but of course Sony was a billion dollar company and needed to be in the right. In fact that this had to happen was explicitly given as a reason to create the DMCA)
Anyway, some precedents:
Hotaling v. Church of Jesus Christ of Latter-Day Saints, 118 F.3d 199 (4th Cir. 1997)
“Although the Church acknowledges that its sole remaining copy is not the one it originally acquired … it maintains that the remaining copy does not infringe Hotaling's copyright because it is a replacement copy…”
Atari, Inc. v. JS & A Group, Inc., 597 F. Supp. 5 (N.D. Ill. 1983)
... defendant sold a device for making backup copies of copyrighted Atari cartridges and argued that §117 permitted replacement/archival copying. The court rejected the broad replacement theory.
This very court has clearly declared that making a copy of a copyrighted work for replacement purposes is illegal, on multiple occasions.
I would like to point out that this isn't Anthropic's only extreme-WTF law violation. When the original judgement against them was made against them, they were forced to admit that using books to train models was illegal if acquired illegally AND THEN WERE ALLOWED TO KEEP DOING IT (thankfully the court never mentioned that part in the judgement so at least they can claim that was never decided when it becomes a huge problem in future cases as it obviously will). But that's not how this works. In my opinion Anthropic and OpenAI and everyone else need to at minimum take training material that was acquired in violation of copyright out of their training data unless and until they have a separate licensing agreement with the copyright holders. As long as Claude knows more about Harry Potter than is said in the promotional summaries it is obviously in violation.
Because of the copyright-filesharing court wars of the 2000s, which were also handled dishonestly by courts (whether we're talking US or EU courts), and the absurd copyright extensions, they had to now make some new excuse, and settled on this very sad, very destructive option. It's not even defensible legally, imho, but of course the biggest wallet must win. I don't understand. It's such a sad joke at this point, and it's not like the courts even still had credibility after the file sharing cases.
I wonder which sad excuse will be forthcoming from the courts when we have someone release a movie made by an AI model that is obviously a direct ripoff from some high-budget studio movie, and Disney needs to be protected from ... say ... "Scorched: Brothers of Aridelle" — In a vast desert kingdom, the royal brothers Elias and Anders grow up together, but Elias secretly possesses dangerous fire magic and isolates himself after accidentally hurting Anders as a child; years later, at Elias’s coronation, Anders announces his engagement to the seemingly charming Princess Hanna, provoking an argument that exposes Elias’s powers and sends him fleeing into the dunes, where he accidentally unleashes an endless heatwave that dries the kingdom’s wells and turns the capital into a furnace ...
I can think of a couple of solutions: 1. Glue a US $1 bill to the spine and take them to court when they destroy it. 2. Send them a license to use the book, with the condition that if they fail to return it within 30 days, they owe $1M. (Hey, if e-books can be licensed, why not physical books?)
In the end its not who can train the better model its who owns the better stack of training data. Acquiring old texts then destroying them makes their information proprietary.
How many AI companies are doing this? Say there are only 5 copies of a book left out there but 100 different companies want to shred it. Rare items could effectively vanish from the market. "Rare" meaning inclusive of high-quality items, disregard the bot opinion that mass book destruction is fine because most books are worthless anyway. And are any safeguards in place to make sure scanned books get saved in their original form for posterity before being recycled into trainingslop? Aside from Google's quasi-legal/ethical mass book scanning operation of course.
This is the citation needed that is missing from every report so far, including this one which deliberately refuses to reveal anything about these books.
You’d think if there were examples of actually valuable, rare books being shredded that the journalists would at least be able to name one such example. Instead it’s always vague posting about the destruction without ever naming any examples.
I think it’s because if they named some example titles, everyone would see that they don’t care about these books being shredded.
On the contrary these companies should be publicly listing the name/info of every book they destroy and use as training data.
If the books are really worthless as you say then their case would be proved transparently.
This article presumably had to maintain confidentiality to protect the seller who agreed to place a tracking device in the shipment.
I would say the burden of proof is on the companies destroying human cultural heritage en masse, not the handful of journalists calling for attention to the matter.
> I would say the burden of proof is on the companies destroying human cultural heritage en masse, not the handful of journalists calling for attention to the matter.
You're literally assuming the conclusion. The exact topic under contention is whether they are "destroying human cultural heritage".
I do work adjacent to the AI book scan-shred pipeline. There are definitely significant books that aren't "Windows 95 for Dummies" which are getting down to single-digit remaining copies.
I just looked for one novel, which wasn't a fantastic book, but it is the first use of a pithy and fun phrase that is so ubiquitous that you'll probably read it a couple of times today. I argued with Claude, GPT and Gemini for ten minutes just now, even knowing the title of the book, to even prove the book exists. It took me years to find a copy originally and then I lost it in a move. I found one more copy today from a rare book seller, but it just sold (to Amazon?).
Is the book valuable? Not particularly, but I feel it's noteworthy and important. I don't want to name it either, because now I have some searches out and the next copy that pops up I'll scan and put on IA. There can only have been a few thousand copies originally published in 1947, it's only in hardcover. I know of a couple of other copies in private hands, so it's not zero copies, but it has to be single-digits.
I have one periodical issue that I know of only one other existing copy (Worthpoint only shows one copy ever sold in their database) and if you look on collector sites there is a blank because nobody even knows what the cover looks like. I can't explain it, since the publication routinely printed hundreds of thousands of copies of each issue, but here we are. Perhaps all the copies were withdrawn and pulped immediately after publication for some reason? It's in my scan pile, so I'll have it uploaded soon. Is it significant? Not hugely, but every other issue of this title has been scanned already, so it's scratching an itch to get this one done.
There's definitely rare stuff getting scanned and shredded. Someone in a comment above said it's not like Nazi book-burning since they were trying to destroy information. But it is like that if you consider there remain no other physical copies and all the electronic copies are locked up in a way that nobody can access except to trick an LLM to spit out a paraphrased copy from its training data.
TechCrunch is trying too hard to make a connection in the title there. Amazon was trying to make money then as it is now. Their selling of books at the beginning was no more principled than the destruction of books now.
This is the legal loophole that allows them to do what they need to do, and the benefit of doing it outweighs the modicum of outrage this title will generate.
Edit. Because I see my statement confused all the HN experts:
Anthropic's version of this was Project Panama [1]. The destruction of the original allows them to keep a digital copy in the AI training dataset. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
all due respect but if they buy them they can do whatever they want with them. not like they are people and not like you have a right to them either. not even like you would have been able to acquire them yourself.
Is it a “loophole” that you can buy something and burn it? That’s how property ownership works. If the books were really so rare and valuable then the original owner should have given them to a museum rather than sell them as scrap. Yet no one who is complaining right now cared about them before Amazon got involved.
No, that's not the loophole. The loophole is that a US federal judge recently ruled that scanning a book for the purpose of training an AI is "fair use" and hence legal _if and only if_ the book is destroyed in the process. If you retain the original, you would be creating an unauthorized copy, but if the original is destroyed, then it is permissable as format shift: https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/...
Sometimes original owner does not know that it is a first edition single copy. For them it is an old book. There have been many cases where owners did not know true value of their antiques.
Well in this case the original owner doesn’t know, the purchaser doesn’t know, the reporter doesn’t know, you and I don’t know. So what actually makes these books rare?
> Is it a “loophole” that you can buy something and burn it? That’s how property ownership works.
It sounds shocking if you're missing the context and only rely on the blogspam which is this TechCrunch article.
They aren't buying the books only to destroy them but to build an AI training dataset, which means they're making a digital copy. This becomes a copyright issue. The destruction is the legal loophole that allows them to keep the digital copy as the only copy in circulation and be considered fair use as decided by a judge (see below).
The concept of ownership means you can do anything you want with the object, the book in this case. Not with the content, like make or distribute copies. The scanning machines are cutting the spine of the book, feeding the pages to scanners, and then destroying the physical copy.
Anthropic's version of this was Project Panama [1]. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
So books that would probably have ended up as trash. These AI training facilities are actually doing these book a service. Not only they are probably going to keep the scans safe (for future training), but having the book end up in an AI model may be the only way it is going to have any use at all.
Let's say for instance that the book in question is about woodworking, and it is not great, lots of mistakes and inaccuracies, unoriginal content, etc... except for a single thing, maybe a trick for making a certain measurement or something like that. Who would read such a crappy book for this single good trick he doesn't know is there, well, an computer will, computers process terabytes of crap without tiring and complaining, that's what they are for, and with a well designed LLM, that one trick may resurface, waiting for someone to ask about that specific measurement.
The problem here is not that rare books end up in AI training facilities, it is that these AI facilities are owned by for-profit companies keeping the data to themselves. These books should go to public libraries instead, for everyone to access, it would be better if these books weren't destroyed in to process too. But the question becomes: why didn't public libraries didn't do that in the fist place? And maybe in a more respectful way. The AI companies would just have had to license the database to libraries, probably simpler and cheaper than having their own scanning facilities.
To me, this mess is a failure of the copyright system. One one hand, large scale book digitization projects intended to preserve and make the original text accessible get lawsuits by publishers, while AI training is "fair use". It means we have built a system that encourages destroying rather than preserving these books!
Yes, we've inadvertently set up a system of incentives that were designed to preserve information and monetize it. And we've instead set up a set of incentives to make it scarce and destroy it.
I really have trouble getting worked up about this. "Rare books" is thrown around regularly but my gut feeling is that's not the case. These are used (often? always?) books and while I'm sure there is waste, in general they just want 1 of every book.
While I wish there was a repository of every book that was already digitized (it pains me this is the best solution), there isn't one and so I think this is not a real problem.
It'd be a different story if they had furnaces that ran only on rare books that they had to continually feed books to but that's not what's happening here. And that 1 destroyed copy will "live on" in a way that it otherwise might not.
It'll live on if they publish those scans or contribute them to a national archives or something. Proprietary data has a habit of being lost over time though.
The problem is that we're ignorant of the true value of objects and we don't know what will be valuable in the future: https://en.wikipedia.org/wiki/Palimpsest
It seems a little short-sighted to destroy an artifact to get the text.
Perhaps the genetic material that remains in books from the people who handled them, or the pollen from plants in the environment that the book existed in will have value in the future, but we won't know what we lost because some people foolishly destroyed it in a bizarre quest to make AGI that the creators argue could potentially destroy humanity.
The more and more I read about these kinds of people the more I'm starting to realize that they're in the "here for a good time not a long time" group of people and those are the last people you want making long-term decisions.
Millions of books go in the trash every day. Should all private property come with such an asterisk that it ought to be preserved for whatever unlikely contingency you can imagine? I don't know, AI actually seems like a much more worthwhile pursuit.
In one of the first articles that I saw about this topic they were describing a scenario where a botany manuscript from the 19th century that only had three remaining copies was being digitized and then pulped.
They don't care about the R.L. Stein Goosebumps series or Louis L'Amour pulp fiction. They want the stuff that isn't digitized and they're willing to destroy original artifact to do so with no guarantee that even the digitized text will ever be made publicly available.
It's foolish and selfish.
> In one of the first articles that I saw about this topic they were describing a scenario where a botany manuscript from the 19th century that only had three remaining copies was being digitized and then pulped.
That was something that someone on Twitter made up and then got quoted in an article. All of the books being scanned have ISBNs, meaning they're from 1970 or later. The books being sold off are rare in the sense that there aren't many copies in circulation, but they aren't priceless artifacts. If they were, they'd be being sold individually at high prices by boutique booksellers rather than being sold by the pallet by wholesalers. It's going to mostly be stuff like "The Best Places to Eat in Milwaukee, 1992 edition" or "The 2006 Guide to the Stock Market" or "Teach Yourself Visual Basic 3.0 in 28 Days".
On its own that seems like a pretty weak justification, at that point we would need to preserve every copy of every book forever because you never know what what useful physical material it has on it even beyond its text. That extends to non-book items as well, and while I understand the idea it's just not realistic in any sense.
I think the more important part is to define what "rare" actually means and save/digitize set copies of those books that are actually at risk of being lost completely (rather than letting them rot away somewhere or get bought up to be privately destroyed). I suspect many of them are just not at all interesting enough to justify the expensive though.
My comments on this subject aren't some sort of justification for any particular course of action; they're more a disapproval of the course of action being made by people people who have more dollars than sense.
As a general rule preserving historical artifacts for continued future analysis and appreciation by people who have yet to be born is a noble cause that's considered worthy in and of itself without the need to justify it. We're talking about the richest group of people who have ever existed with the means to preserve them but decline to do so and instead act like the Taliban blowing up statues of Buddha.
I think this is intuitive to most people. If a prerequisite for the birth of AGI was that a humanoid robot had to sit down in the Louvre and eat every single painting there with a knife and fork while occasionally stopping to wipe away historical detritus with the Mona Lisa that it's wearing as a bib people would by and large express a visceral and justifiable outrage.
The people who are doing this know how this looks so they're trying to do it all behind closed doors, through cut-outs and intermediaries.
They are literally not allowed to publish those scans.
Half the reason the books get trashed in this process is because the first sale doctrine keeps copyright from strangling all the freedom in this narrow area.
The books get trashed because loose pages are much easier to scan.
Depends on the book. I'd wager the majority of the rare books are in the public domain.
But they won't share them because they don't want their competitors to have the data.
If they are in the public domain, they could publish them. How many books are in the public domain but have never been digitized and shared in a public archive?
And if the books are not in the public domain, then they should not be allowed to train their AI models with the material without some kind of license or agreement with the owner of the copyright.
> If they are in the public domain, they could publish them.
That'd be an easy fix with new law -- if you are an AI company with book data, you have a burden to make openly available (or require your suppliers to) all public domain book scans.
> They are literally not allowed to publish those scans.
Emphasis mine. Just provide a digital copy, no questions asked. They will distributed to various archives globally. I'll pay for the drives and shipping. I understand and can appreciate the potential liability, and am willing to launder it to preserve the subject collection(s) and dataset(s).
Rare books tend to be out of copyright, true.
We run a little library in front of our house. It's amazing how many people dump boxes of old books off in front of it hoping that they will find their way back into someone's collection.
The sad truth is the vast, vast majority of printed literature is neither interesting nor useful. People are not dumping off stacks of Umberto Eco. We frankly have to toss a lot of awful cookbooks, self-help books, trashy mass-market "novels", and sketchy religious works. As it is, even the stuff that makes it to the library is not very impressive.
Imagine if we had bad cookbooks, cheap popular novels and stories, tracts on diet and self help from old Rome, or old eras in China, or the equivalent from ages before that.
It's all interesting for something, even if it's just a meta analysis of culture during a certain period or what kind of trashy romance novels were popular in 198X. At least in my view.
If only the Roman's had digitized their self help literature!
This is most of what we have for the great philosophers - half "why is the world like this" and half "how do you live a good life", right? some original plato lectures would be interesting to have.
> Imagine if we had bad cookbooks, cheap popular novels and stories, tracts on diet and self help from old Rome, or old eras in China, or the equivalent from ages before that.
It would be great, but it's a bit off topic here. The point the parent is making is that most of these works are going to be trashed regardless of Amazon's behavior. If Amazon (or any other company) is digitizing it, they're at least preserving it in some form. The alternative may well be that many of these books never get preserved.
Now would I prefer the government or an entity like The Internet Archive do this? Sure.
They were all lost when the library at Alexandria burned, so clearly we should ban libraries. Too much risk of losing unimaginably priceless works all gathered in one place like that
> they just want 1 of every book.
That's the point I come to.
Books are not original manuscripts. Even in low volume cases, they are usually printed hundreds of times. (And usually low volume works aren't all that great...hence the low demand.)
That is a fraction of a percent for books that at some level weren't all that wanted.
Do we know if they actually buy one copy of each book here? Or are they indiscriminately buying books in bulk and just processing all of them?
The latter seems inefficient, so my first assumption would be that they would avoid that. But while I suspect they'd check if they know the book before scanning, I could imagine them not caring that much before buying them and just focus on volume.
Agree. Couldn't care less. There are some neat aspects to "rare books" but overall quite insignificant.
The Embassy of the Free Mind (https://www.embassyofthefreemind.com) is a rare book library in Amsterdam that is scanning books the old fashioned way… leading to https://SourceLibrary.org — a collection of over 5,000 books from the renaissance that have never been translated before. Consider donating, if this is a topic you care about!
The length of this article and the unnecessary visualizations felt like a waste after getting to the end and discovering they won’t reveal anything about the books that were scanned.
The closest they got was admitting that the rare books weren’t anything that someone might care about in the sense that people assume when we hear “rare books”
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
Okay? But then why exactly where they considered valuable enough to warrant an entire article about them going to a book scanning facility? Without revealing anything about these books I have no idea if they were classic literary works that were underappreciated, or if this was some old guide about How to Use Microsoft Office 97.
There was a more balanced take on Twitter (which I’m unable to find again, because Twitter) from a book seller who said it was more of the latter type: Books that were rare because they were no longer in demand and most everyone had thrown their copies away. Some parts of the media are doing backflips to try to imply that these are cherished literary classics being fed into the shredder to deprive humanity of something valuable, but the book seller seemed happy to be making sales for useless old books that no human was interested in buying.
Anyone who is into books eventually finds out that we're permanently losing them all the time. Like they're thrown out, and lost forever. University libraries throwing out huge collections of out-of-print material to make room for new books, study spaces, even cafés. Municipal libraries turning over their collections. Books that never made it to libraries going out of print, tossed in the trash after yard sales.
The AI companies digesting this stuff is a net win for humanity. And I'm not a fanboy! Ideally they'd upload them to Anna's archive too, but even if they keep it private forever, at least these books live on in some way in the model weights. Thats better than a landfill.
If or when they go bankrupt or reach agi, they will just delete them. I hope Anna's archive already has them anyways. Apparently openai's newest unreleased model, gpt 6, is capable of continuous training at Inference time, like a person is. That might be enough to delete them
And that’s why imo the outrage should be about copyright law not the businesses. I think it’s a hard problem to solve but right now with copyright on new books being the life of the author + 70 years it’s a bit silly.
The only way to gatekeep a book behind payment should be, if it is in print.
I like this and think it could make a lot of sense but you would need to thoughtfully tie it to volume or something similar. I could see publishers gaming the system. I would also add that once it drops out of print that it should belike generic drugs. Anyone can use it. IMO making it quicker free use stops most of what this article is describing.
If 23andMe's genetic records were acquired (along with the company itself) by a different entity, why wouldn't this data as well?
Those records are far, far smaller.
I would feel much better about this process if they were uploaded and if it were framed as a knowledge preservation project. This would only slightly increase the cost of the project, but have a huge impact on its perception and its net positive impact.
Of course, actually benefiting humanity is only a minor, indirect concern for investors.
I thought this was the original purpose of Google Books. It's actually a little surprising to me that there's anything left out there worth buying in physical form and scanning that hasn't already been digitized in some way.
Google faced a decade of litigation for making books searchable. It's all downside and little upside for a business.
They could scan and store the data until the copyright expires without issue probably. It's a long time line to work with but you could probably get away with it that way.
It would be really great if the companies who are doing this would commit to placing the scanned files into a public trust that would coordinate with organizations like the Gutenberg project to ensure that the scanned materials enter the public domain on schedule. Publishing encrypted archives with the keys in escrow would be a good first step.
IMO that would go a long way to resolve any concerns about losing books. I still don't like the idea of extremely hard to find or last prints being actually destroyed for this, but it certainly makes it more palatable.
I wonder if there's any way this can be construed to get favorable tax treatment. That'd actually get them doing it.
The real solution right here.
If it's going to cost us tax dollars I'd rather see our own government doing the work of preservation and making those works accessible to the public.
They cannot do those things. The reason they’re scanning the books is because it’s not possible to legally obtain or transfer their digital copies. They have to do their own scans and keep them in house.
First, to be clear, I work for Amazon, and I don't work on this or related efforts, I am a security engineer. I also won't talk about or answer questions at work, and my comments are more generally about the practice (all of the major organizations building AI are doing destructive book scanning). These are my opinions, and do not reflect my employers (past or present).
You might be right that they can't do that now. The simple path forward is to have these companies simply make a public commitment to publish the data when the copyright expires. I also think there is a space to be carved out, probably through regulation, to ensure that there is a clear path for these scans to enter the public domain, at the very least.
This is not just important for these specific books, I have written on other platforms and in other spaces about the importance of media companies and those who benefit from strong copyright laws to protect and generate profits and revenues to repay the public for the cost of that enforcement over time by ensuring that at the appropriate time, those works fully enter the public domain. That could mean a restructuring of the Library of Congress in the United States to become a modern Library of Alexandria to host data, and shifting to a registered copyright model where to gain the protections of the court, you need to upload/submit your copyrighted works for storage and eventual release. I doubt it would ever happen there because of the amount of money invested in tying up IP in the United States, but perhaps a more amenable location like the EU could help with that.
However it might work, part of the promise of the Internet was that information would be liberated, but we see every day how much information gets sent down the memory hole when businesses, sites or services shut down, or how regularly companies abuse IP related regulations to attempt to strangle competition. It would be expensive now, but it would create an incredibly valuable legacy of information for the future, and it will only get more expensive to build such a thing as time goes on.
Great comment. Those who try to excuse book burning cannot be doing it in good faith, this is the acid test.
If AMZN, et al. were burning books in good faith, they would be open about it, point to the party responsible for it, and provide a list of the burned books to allow the community to organize, scavenge and scan the endangered publications. Keeping everything secret is a proof of maliciousness.
Preserving books and providing easy access to the information in them is of utmost importance, this is certainly understood by the public and private entities who could do something positive about it - their actions in the opposite direction should be a wake up call, they're on the wrong side of this issue.
Libraries I volunteered for would coordinate throwing away books late night under dark right before the dumpster was picked up and emptied.
Because people are irrational when it comes to this topic. They say stuff like “book burning” when they see books being destroyed. Doing it in relative secret kept the crazies away.
If you are passionate about this topic, lobby for more funding to store archives in the public good. Expecting private parties to do it for free because it gives you the ick to see books destroyed is not useful.
More books get destroyed each year during estate cleanouts than any AI companies could ever hope to accomplish. Most books donated to goodwill or other thrift shops go straight to the dumpster and might not even have a set of human eyes put on them at all. These places act as sin eaters for folks to leave their trash with.
Copyright does expire, for older books.
It would make more sense for government to accept digital copies for any book, and share the out-of-copyright ones. In the United States, we have a Library of Congress who could do it if a law were passed with funding.
What they could do, if they were truly focused on not taking the most destructive path forward, would be to cooperate on a single scanning company that could pool the resulting text for the others on attractive terms, to ensure that once a book was scanned all the others wouldn't need to scan and destroy the same title. That would substantially limit the damage.
Of course, because each company is happy to burn it all down to beat their competitors and being first to any content is most important, this will never happen.
Somehow people are still failing to understand that transferring the scan from a trust, scanning company, or other external party would constitute a transfer of a copy between distinct legal entities that breaches copyright law in an actionable way.
Doing the scan within their own organization does not because it's within a single legal entity and there is no copyright issue involved.
I'm not failing to understand that, I just don't think it's an insurmountable barrier.
You could overcome it through various means, including helping to set up a clearinghouse for bulk rights, implementing a data clean room approach that models could train on in situ, lobbying Congress for copyright law changes especially around orphan works, using Section 108 of the Copyright Act to set up a specific preservation vehicle like the HathiTrust, and other options.
I find it ridiculous that so many in this thread are acting as though AI companies are simply powerless to do anything but buy up and destroy these rare books.
Make a case that it's more profitable to take any of the courses you suggest and you might be on to something but they aren't. All of those things take time and time is money when you're running a business.
I would rather make the case that with the increasing consolidation of media under the big tech companies, regulation is more effective than a business case.
Of course they’re not more profitable, they’re negative externalities that aren’t being priced into a finite commodity.
That’s the whole point: Companies are doing this because it’s easier and cheaper for them, and those doing it now are first-movers who don’t care about how catastrophic the end result will be for the rest of society as they’ll have their training data moat.
It should be obvious, but the things that are most profitable are not always the things that are best for society as a whole. That’s why we have regulation.
I was thinking that maybe this is not to their benefit.
Just imagine 20 years from now its hard to get books in print. AI companies can just change the history by altering their model's content.
I'm not to keen on corporations holding the world's entire print history in AI models.
That is the problem: they are digesting it. They are not creating a new kind of library, where you could say, show me the text of "How to Fix Your Ice Cream Problems". (An actual book I own) It is not being done for our future reference.
They are digesting it an incorporating into the weights. You won’t be able to get the exact page but you will (if the model is good) be able to get the knowledge out of it in a likely far more concise, relevant, and certainly more widely accessible than the book sitting on your shelf.
That's a fairly limited view of what books can be for.
> You won’t be able to get the exact page
This is a grave problem.
> you will (if the model is good) be able to get the knowledge out of it
That doesn't help with books that aren't factual in nature. "Getting the knowledge" out of a novel makes no sense.
> Ideally they'd upload them to Anna's archive too
Not only can they not do that, they must scan physical copies because they are forbidden from using digital pirated copies from sources like this.
Anthropic had a big settlement because they were caught using downloaded digital copies. As a response they’ve ramped up their book scanning and others have followed.
Yeah, the very disjointed way that copyright is applied (I would argue, rights that should also transfer across to digital media but generally have not) is contributing to this situation in the first place.
This is the right answer but the reason they can’t upload the scans is copyright law as demonstrated by Google having to settle with the publishers and allow them to remove their books and limit free access to 20% of text. The AI companies are essentially compressing the information in a huge swath of books that would otherwise be headed to landfill and making them 1000x more accessible. This is unquestionably one of those instances where capitalism is taking money from rich investors and benefiting the 99%.
As if the AI industry is only shredding "crappy" books.
I recently bought a rare book. It was a boat design book written by a famous yacht designer in the 1940s, but it's been out of print for decades and I had to pay $150 for a "fair" copy with missing dust jacket.
I have an interest in older technology and the old ways of doing things (for instance, how do you lubricate the mast of a gaff rigged boat so the gaff jaws don't jam?) I've often found myself reading very old books that have been out of print for a century.
Sometimes I read those books at libraries and I've been the only person to check them out in years (I started doing this back when they still stamped the return date on a card so you could see when it was checked out). Now most of those books have been disposed of by libraries due to yield management software and I've ended up with some of them in my personal collection, but people like me can only save a tiny sliver when most of them are being bought in mass by Sam Altman. When they're gone the knowledge in them is also gone.
We are burning the library of Alexandria and the HN consensus is "those books probably weren't saving anyway."
Once you start to realize just many books got thrown out annually before “AI” you have less sympathy. Libraries are one of the biggest contributors to this because there is simply too many books that nobody cares about.
I suspect folks are over weighting how much knowledge is being destroyed in these books. If someone actually quantified it that would be amazing but as someone who started going to used book sales at a very young age I just have no sympathy. Most books are worthless. I don’t mean that from a text perspective either.
Libraries usually offer those books for sale to the general public. I just picked up two from my local library that way.
AI companies are not doing the same.
AI companies are building the knowledge from the books into their models, which they offer for sale to the general public.
Far more people will use the model than would ever have read the book.
They're not reproducing the information verbatim and with the ability to definitively reference its source. Moreover, not all AI companies are going to get the same books and therefore the same training set, and as companies go bankrupt, training sets are changed, and models abandoned, that information might disappear entirely.
It's not at all the same thing as still having the books themselves available somewhere.
On the other hand, the book can only be in one place and read by one person at a time.
People may not even know the book exists or that it contains the information they want. Obtaining access to a copy may be very difficult, even if it is not particularly rare. The AI is much more accessible.
Then find a way to make them digitally available, verbatim, like Google Books attempted to do. I'm sure with the power that AI companies have they can have legislation amended to make it happen.
You're defending the indefensible here.
Lol, as if they can just get copyright reform passed because they asked nicely for it.
They can't legally publish the digital copies, and you know it. Training is fair use but direct copies are not.
You are defending something that makes no sense. If you want to get mad at something get mad at copyright law.
That copyright law is a mess doesn't mean that companies that buy up rare books, destroy them, and keep the contents locked up are beyond criticism for their actions.
They don't have to scan and destroy these books, it's an entirely voluntary choice.
Name the rare book.
And libraries throw away a lot of books that go to the dump. So what?
You should be pushing for better copyright law not Amazon for buying up worthless books.
Please scan this book and put it into Anna's library, or keep it for the future.
I would be also interested in participating in your costs.
> As if the AI industry is only shredding "crappy" books.
Do you, or anyone else, have any source suggesting that they’re buying highly valuable rare books and shredding them?
The kind of $150 rare book that you had to buy from a specialty collector who graded it is in a completely different category. You’re thinking of “rare books” in the historically rare, valuable, and collectible category.
The book sellers shipping off orders of 1000s of books at a time to these facilities are calling the books “rare” because they may only have 1 copy, not because it’s a collectible with a high price tag.
There are a ton of niche books like the one the parent is speaking of which might not be considered hugely valuable by the general public or less specialised second-hand stores, but are both rare and valuable in their niche.
For instance there were only limited production runs for a set of books on South Africa's participation in the Second World War, and getting hold of them is increasingly difficult. But they have ISBN numbers, so they fall within the group of books being collected and destroyed by AI companies, and as a result might disappear altogether, along with the knowledge inside them.
Seems like getting a copy and scanning it and releasing it as part of a digital library would be a good way to preserve the knowledge.
Making the contents of the books available verbatim afterward would be a less bad option, yes.
That's not happening though.
Right, but that's because of copyright law.
Have you seen AI companies making any effort, with their newfound immense power to get the attention of lawmakers, to lobby for amendments to copyright law to make this exercise a less destructive exercise?
Because none are.
Why would they? They're not the ones upset that they're cutting up books. They fought a court case to allow them to do it a different way that didn't require cutting up any books. They were shut down and had to pay billions, and so they're doing it the only legal way that is left to them.
This is literally what the publishers asked for. If they wanted a different way, they're the copyright lobby. They are making all the rules. They can change it.
Those poor AI companies, they’re the real victims here! /s
There are several things to criticise here, including the US’s terrible copyright laws and the publishers involved in that case who care more about in-print than out-of-print books.
But it’s the AI companies that are actively buying up and destroying books at an unprecedented scale, which is an action they’re choosing do to and certainly don’t have to do. Therefore that’s also where the immediate focus and attention needs to be
> There are a ton of niche books like the one the parent is speaking of which might not be considered hugely valuable by the general public or less specialised second-hand stores, but are both rare and valuable in their niche.
I'll ask again: Does anyone have a single example of this?
Every conversation about this topic has been full of this same assertion, but nobody can ever name even a single example. We're supposed to believe it's factual.
You’re supposed to believe it’s factual because you have investigative reporting by multiple newspapers with good track records, on the record statements by bookstore owners, and evidence presented in court cases like the one regarding Anthropic’s Project Panama.
Notably, one of those piece of evidence from that Anthropic case was an internal presentation clearly stating the intention of Project Panama was to “destructively scan all the books in the world.”
Beyond that, some book sellers have named specific titles, such as “an Estonian translation of John le Carré's The Mission Song, a specific edition of Anne Brontë's Agnes Grey, and the October 1983 issue of Warship monthly magazine” and “everything from 18th-century African agricultural implements to biographies of 1950s racing drivers.”[0]
But getting a full list of all those sold globally is difficult. I imagine we’ll start to see some efforts to try to encourage booksellers to contribute to public lists though.
[0] https://www.theguardian.com/technology/2026/aug/15/uk-irelan...
Im with you, my take-away from the article was of sadness for all the books that are going extinct because of this.
Real humans sharing their unique knowledge, packaged in a book.
The thing is, I doubt this is even in the top 10 causes of books going extinct. I feel like people often over-romanticise the medium, especially.
Weeding (the library term of art for selecting works for deacquisition) is a very import part of collection management. It's actually regionally and nationally coordinated so the interlibary loan network does not throw away the last copies.
https://cdlib.org/west/ https://papr.crl.edu https://eastlibraries.org
You're fortunate if that's the case for where you live.
In the UK weeding is at the discretion of the Branch Librarian. Books I checked out 15 years ago have now vanished from the system.
Can we copy the books? It seems like in an effort to monetize writing, we've created a system that incentivizes destroying it.
It would be nice if your local library hadn't sold their copy but I think your beef is with the libraries, or perhaps with the politicians who failed to fund them in that case.
Actually I expect that a copy of most books probably is still available in a copyright library but it might not be easy to access.
https://runtimewire.com/article/anthropic-settles-book-pirac...
https://www.gadgetreview.com/we-dont-want-it-to-be-known-ins...
https://www.irishtimes.com/world/europe/2026/08/10/a-mysteri...
https://www.bbc.co.uk/news/articles/cp3rprx2wl4o
https://dallasexpress.com/national/the-vanishing-page-ai-fir...
https://fortune.com/2026/07/31/dutch-bookseller-ai-spam-phis...
This is why we read the comments first :)
Thank you.
> or if this was some old guide about How to Use Microsoft Office 97.
Materials like that may be very valuable to software / tech / HCI archeologists soon.
I can't get the tools or local know-how to straighten my scythe blade in a country where every cottage had a scythe less than 100 years ago, with the last scythe-native generation rapidly dying out.
Fast forward a civilizational collapse and that M$ Office 97 for Dummies might be as groundbreaking as a Guide to Using Roman Concrete
This feels like a manufactured controversy. What difference does it make to me what someone does with a book after they buy it? It's effectively unavailable to me regardless of what they do. If people are really concerned about these "rare" books, they should lobby the copyright owners to release them online or print more copies.
It's better value to me that someone is digitizing and training with books then they just languish somewhere in perpetuity.
If there’s only 3 copies in existence, and everyone on the frontier wants it in their corpus, what do you think will happen?
> what do you think will happen?
They raise the price and print more copies?
What if the book has been out of print for 50 years? This isn't a contrived scenario, most of the rare books are like this.
> What if the book has been out of print for 50 years?
Same answer? They raise the price and print more copies.
You can just print even 1 copy, just the cost per print will be higher.
In fact publishers getting some profit on the tail end of books is the best thing that could happen to authors and publishing industry.
Do you understand that many of these publishers do not have the original copy anymore, and may have gone out of business entirely? 50 years ago text was not backed up to the cloud.
If there are very few copies and no one wants to even spend $100 to scan it, likely the book wont be bought by anyone in any case and it would rot/thrown away in another 50 years if AI companies had not digitized it.
Sorry, that "just-so" explanation is not how the world works.
Which is what has happened over time now that humanity has the ability to digitize them. The problem arises when we are purposefully picking books where we know those scans aren't yet available and purposefully destroying them. There'd be less outrage if any of these companies made statements about guaranteeing the survival of the scans they are taking and the release or planned guaridanship in the case they can no longer maintain the data.
Who is digitizing the books which are rare? Isn't it even more valuable to digitize for them now that they have new buyer which can pay 100s of times more.
Anna's Archive is a controversial example with that as a stated goal. But more to the point most of the people doing it aren't seeking to profit from digitizing these books so the 'monetary value' of them is not the yard stick or motivator. A higher monetary value could actually make it harder for these projects as they can rely on donations of books or even waiting for someone who wants a book they bought so they can spend that on the next one (which is less useful if the next one on your list was bought to destroy by the same person who bought the book you finished scanning).
They're often books tied to a region, i.e. stories about a local fisherman's life and how they overcame adapting a boat to local feedstock, or the engineering that went into building the local train station. That's to say they can have little value to those outside the area but are still a rich read for locals, tourists, and those who move to the region or have a love for a place. I expect we'll see a concerted effort to organise to get these kind of book sent to regional mueseums for scanning now that people are aware they are actively being destroyed and they realise the timeline isn't how long the book can last.
You are not understanding how the market works. Rarity and value do not always correlate with demand. Otherwise I have some beanie babies to sell.
I doubt that your beanie babies are very rare.
> They raise the price and print more copies?
How? Like, do you think a book published 100 years ago has a digital file sitting around a publishing house that might even exist anymore? Re-publishig of old books are often based on scans (the originals were literally pressed into paper by metal, not printed off hard drives), and if the books don't exist, they can't re-scan.
I own some antique Japanese books that have very few copies in existence. You couldn't get them re-printed if you want to if the physical book didn't exist anymore. For a few books, I've actually purchased incredibly high-quality scans that cost me hundreds of dollars because to produce the re-issues, people had to go to museums with high-powered cameras to photograph the pages under supervision of a curator. There's no digital file to print from the publisher. If those books were gone, there's no bringing them back.
I think the term "rare" might be a bit loaded. DOES it mean "only a handful in existence" or more like "1000 copies?" And at any rate, at least by scanning the book they're theoretically making its contents available to the public. What are the other people who own these books doing besides having them sit on a shelf?
"Only a handful" is a precious or extremely rare book. 1000 copies of a book total would certainly qualify it as rare. That's not a lot of copies.
> What are the other people who own these books doing besides having them sit on a shelf?
Not destroying them. By continuing to exist, it keeps the possibility that the books will eventually go to someone else. Maybe even a library that specializes in rare books.
> By continuing to exist, it keeps the possibility that the books will eventually go to someone else.
Good news, there's hundreds or thousands of copies that continue to exist. Anthropic, Google et al only need a single copy.
And shockingly, for all this talk of rare books, we don't see the other owners of said books stepping up and offering to have these books scanned non-destructively. I won't hold my breath for them to do so in the short or long term.
FURTHERMORE, the books aren't being scanned for no reason, they become part of new model training data. Meaning that whatever insights these books contain might be available to users in the future. Again, more valuable than all the copies rotting away on someone's bookshelf that might be scanned and made available some day.
It's hard to describe what happens to the book's contents, or determine whether it would be more valuable intact. Insights are the result of a person interacting with a book, they aren't like juice that can be squeezed out mechanically.
I care a lot about the high-illustration picture books of the 80s. Since they have ISBNs, they're presumably on the list. Does the LLM ingest the artwork? Even when it comes to text-only novels, it isn't making the experience of reading the book available.
This is a problem whether it's AI companies buying them or three random dudes. The real solution is to get the copyright owners to keep distributing copies, or to change copyright law.
The three random dudes will probably not destroy them, quite a significant difference.
Not only that, but some random dudes scan the books and put them in shadow libraries.
If that would be the case, how come only 3 copies are left?
If there are only 3 copies, then each copy is going to cost thousands of dollars.
Sample:
> Thou shalt commit adultery.
https://en.wikipedia.org/wiki/Wicked_Bible
> What difference does it make to me
It's the scale that matters as first, and secondly, most people don't shred their books after reading them once or twice. This is just beyond words.
What difference does it make to me if someone shoots the last bison? I wasn't getting to eat it either way.
If people are really concerned about bison, they should lobby gamekeepers to release photos of them.
https://en.wikipedia.org/wiki/American_bison#/media/File:Bis...
A digital copy of a book is identical in value to a printed copy.
This is entirely untrue.
Simple example:
Actual book from 1732 (rare, original, older than your country): https://blackwells.co.uk/bookshop/product/The-Compleat-City-... -- yours for £1,258.00
A scan of the same edition of that book: https://archive.org/details/bim_eighteenth-century_the-compl... -- free. Nobody's paying any money for the digital copy.
I don't think you understand the book itself is a collectible object with rarity and value, regardless of the information it contains.
Plus, the physical copy of the book has additional value in terms of forensic verifiability. You can prove the originality of the text and the absence of alterations, and can easily track changes between editions.
Digital copies can be altered at whim, as the only means to provide any sort of tracking or verifiability is through another external software system that itself has to be trusted.
NFT's
Replace bison with some random animal no one's ever heard about, and photos with literal clones of the animal, and you've got a much better metaphor. Not to forget that in this world, brand new animals are created every day.
If you destroy a 17th century book, you can't make another one. Eventually you destroy them all.
As long as digital or physical copies exist, you can. Regardless, let's be real, nobody is destroying 17th-century books. The "rare" books we're talking about here still have copyright owners who just don't bother printing new copies due to a lack of demand.
> What difference does it make to me what someone does with a book after they buy it?
I think this is a bit of a myopic take. It's like saying it's my property I can do what I want, yet there exists designated historic homes or neighborhoods that are deemed to have cultural value where modifications do in fact need to be approved.
I'm with you on the "rare" part. If people are thinking about 70+ year old documents or ancient manuscripts, I doubt that's what AI is being trained on and is being destroyed, but it's reasonable that people find _that_ idea distasteful.
You can say it's manufactured but if these companies ignore this criticism, it's just another way AI companies are committed to losing the public.
> It's like saying it's my property I can do what I want, yet there exists designated historic homes or neighborhoods that are deemed to have cultural value where modifications do in fact need to be approved.
Yes, and when you buy them with that designation, you know what you're getting into. The problem occurs when you already own it, and some group is trying to get it labeled as historic, which will add to your burden and limit what you can do with it.
When I was in a small town, this was actually weaponized. A hotel owner was trying to get another hotel categorized as "historic" and had rallied a lot of people behind his cause. He had a case - the hotel did have some claim to being the "first" in some category or other. But really, he was doing it because it was a competitor. The "historic" hotel owner had to spend a lot of money to fight the cause, because being labeled historic would prevent him from performing various upgrades, making the hotel less attractive to customers (he was already not getting many customers).
> We’re not revealing the titles of the books included in the shipment we tracked, but they are rare, meaning there are not many copies of them in circulation. Sometimes that’s because not many copies of them were ever printed, and sometimes because they are in a foreign language not many people speak.
Not even the title of one of those rare books?
> The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order.
It would seem the redaction of the rare titles is a way to avoid de-anonymization and subsequent harm to the business of the seller who agreed to place a tracker in one of the books. That being said, maybe they could have chosen a better methodology which would have allowed the disclosure of the title, although ultimately I’m not sure the title matters too much outside of their claim they were “rare”.
The issue is if rare is simply "not many in circulation", that's very very different from rare being "it is a unique artifact". The former kind of rarity doesn't really present an existential end, whereas the latter does.
Yes this is disingenuous to the extreme.
My neighbor self-published a book, printed I think 100 copies at his own expense. It's literally a rare book. I can't imagine he nor anyone would care if an AI company bought a copy, no matter what they did with it.
Maybe these "rare books" are first edition Mark Twains, and. maybe they're unwanted books that would otherwise have gone to be pulped. The distinction is important and by not giving any evidence or even a qualitative claim about the types of books, 404 media is being pretty weak here.
I don't think HN commenters need to be the arbiters of which books are beneficial to society.
And beyond just benefit to society, don't forget how subjective the decision about what constitutes a "unique artifact" is. That random product manual from 1966 becomes a near-priceless relic to me if it lets me repair and use the sewing machine handed down from my grandmother. Not necessarily the ink printed on paper, but the information contained within.
The above is based on a true story involving archive.org's Manual Library. The idea that even such esoteric and forgotten information would get hoovered up into the walled garden of Amazon's AI and then destroyed in the outside world, such that I have to go pay them to access even a facsimile of it, is frankly disgusting.
>That random product manual from 1966 becomes a near-priceless relic to me if it lets me repair and use the sewing machine handed down from my grandmother. Not necessarily the ink printed on paper, but the information contained within.
On the other hand, now that the manual is in the AI, you can just ask the AI how to repair the sewing machine.
The information contained within has not been lost, and in fact has become much more accessible.
> you can just ask the AI how to repair the sewing machine
Like so often in its usage, the word "just" is carrying immense weight in this sentence. Please see "hoovered up into the walled garden of Amazon's AI [...] such that I have to go pay them to access even a facsimile of it"
> has become much more accessible
This accessibility rests on a number of assumptions, not the least of which is Amazon's (of all companies) charitable good graces in offering access to their AI at an affordable price. It also assumes that the model can accurately regurgitate the text without hallucinating about other, similar machines, and that it can faithfully recreate any diagrams.
> That random product manual from 1966 becomes a near-priceless relic to me if it lets me repair and use the sewing machine handed down from my grandmother.
That ain’t a rare book.
I would also be shocked if you couldn’t find a way to repair and use this 1966 sewing machine… so it seems like a red herring.
> That ain’t a rare book.
Thanks for your opinion. I see you disagree with my first point.
> I would also be shocked if you couldn’t find a way to repair and use this 1966 sewing machine… so it seems like a red herring.
"Repair of the sewing machine is obviously simple and thus left as an exercise to the reader."
Seriously, though, I'm trying to say that you can't make that call for me. You can think that, but I know from experience repairing old machines (with more or less sentimental value) that having the original manual can be the difference between fixing it and irreparably damaging it. I could also, in turn, suppose that you probably do not actually collect and preserve rare "unique artifact" books, and therefore your argument about them is a red herring. But then we would both just be casting aspersions.
And even then, the "rare" qualifier is not needed here. What Amazon/LLM companies are doing is amoral. This is intellectual piracy (not in the "copyright infringement" sense) at the highest level. Stealing and centralizing the accumulation of human knowledge to eventually rob us all and put all power into the hands in the hands of a handful of people, who are not benevolent.
In what way is this stealing?
The rare qualifier is used precisely because no reasonable person thinks this is stealing.
Hoarding.
Nothing good comes from hoarding whether it being toilet paper, money or knowledge.
> Hoarding.
No reasonable person thinks buying 1 of something is "hoarding".
What about 1 of everything?
Isn't that usually called "collecting"?
Collectors usually do the opposite of destroying the thing they are collecting.
https://jskfellows.stanford.edu/theft-is-not-fair-use-474e11...
https://styleblueprint.com/everyday/the-quiet-theft-ai-steal...
The rare qualifier is used to highlight the destruction of rare items.
It is still stealing even if the book is common.
How is it stealing if they paid for it?
Reusing the content without consent, attribution or compensation
Makes you wonder how many packets with rare books.containing airtags said recipient receives. My guess is: one so far.
The warehouse worker getting paid near minimum wage to unpack 100 boxes per hour is probably not taking a ton of time to look for and report AirTags.
However, a title can be cross-referenced against purchase histories in probably under a minute.
As I understand it, "rare" in this context could mean anything—even a washing machine manual from the 1980s...
It doesn't really matter. The article suggests that they are selecting and tracking books by ISBN, which means: books that have been published or at least reprinted in the last 50 years or so. And they are trying to get as much of that set as they can, regardless of quantity, quality, or any other consideration. Which means that older books printed before ISBNs became common may be relatively safe, at least from Amazon. And some books just don't come up on the used market very often.
(small historical irony: when Amazon first started selling books, they used the Books in Print database, which included a lot of books not actually in print.)
Rare as in your grandfather's John Deere manual from 1982, not rare as in a test print run of The Great Gatsby. The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
> The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
Citations? Also what exactly are these 'normal means'? As a bibliophile who loves scouring used book stores for out-of-print titles this is a topic I'm very interested in.
> Also what exactly are these 'normal means'?
Normal means is throwing it in the recycle bin. Especially for stuff like a 1982 John Deere manual. I've never donated an old appliance's manual to the library. Have you?
Amongst my friends, I'm one of the rare folks who donates books to the library. Most people just trash them. And I know the library only wants them to try to sell them in their book sales (or online) so they can get money. Almost nothing one donates to a library actually ends up on the library shelves.
Here's EPA data from as recent as 2018: https://www.epa.gov/facts-and-figures-about-materials-waste-...
EPA estimates that hundreds of thousands of tons of books are landfilled or recycled every year. That's just the United States. It seems likely that the worldwide figure is in the millions of tons.
"rare" is used in these headlines/articles to incite and generate clicks
The scariest thing for me about companies not caring even minimally about conservation is that when AI gets more powerful than humanity (which is clearly a when not an if, even if there's a lot of uncertainty and differences in opinion about the time horizon here), I want to hope that AI will care more about conservation of human people.
So far from how I see how powerful organizations work, I'm not as certain as I would like to be.
> AI gets more powerful than humanity (which is clearly a when not an if, even if there's a lot of uncertainty and differences in opinion about the time horizon here)
People say this all the time, but so far nothing has convinced me it's true.
LLM development has more or less plateaued, and the current boundaries are very real - energy, resources, capital.
At this point we're talking about marginal improvements against the same asymptotes of all technological innovations.
So, if a corporation can suck up entire books to teach their machines how to think using the information from those books, can we (all humans) join a single corporation that provides all books to its employees? Just need one copy of each and we'll make that copy digitally available for our employees so they can learn from and utilize the knowledge from the books. It's not copyright infringement, they're employees.
Almost like… a library?
This is what the copyright laws dictate no ?
No. A copy is still a copy even if you destroy the original.
Until about a year ago this would have been a reasonable and respectable argument, but at least in California you are arguing against current legal precedent: https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/...
That's a bit of a wild interpretation of copyright law.
I mean, Anthropic isn't going to fight it because it lets them do the thing they want to do, so I can see how this never gets beyond the court that allows them to do the thing they want to do.
But would this argument would have flown in the past?
It wasn't even attempted in Sony v Universal. Or any copyright suit up until this point. That doesn't smell funny to you?
Sure, it's a bit wild, and no, the same outcome may not have been reached if the case came up earlier. And if it makes it to the Supreme Court, who knows what they will turn it into.
The flip side is that Alsup (the judge who wrote the opinion) is probably the smartest district court judge we have when it comes to technology, and one of the people I'd trust most to come up good decision.
He's a treasure, and my instinct is that he got it right: https://en.wikipedia.org/wiki/William_Alsup
i would assume that would come into play if they were uploading scans of the books? there must be some gray area where training like this doesnt apply to that. or ya know, just do it and face the consequences later because you already have the data and know that ai obsessed government will just shrug their shoulders.
Not to my understanding. To begin with, it's far from given that "rare" books are all covered by copyright. But if they are, it's at best murky: whether you destroy the original doesn't really have anything to do with what you're doing with scanned contents. The scanned contents themselves may be inherently a copyright issue, regardless of destroying the original. The actual trained model has separate arguments more in its favor, so if no scanned contents exist - IE the data is read once for training and not stored or saved, they have a better argument. But in that case the destruction is totally disconnected from copyright, as they'd be totally okay to rescan the material.
Bias disclaimer: Amazon is my current employer, but I don't work on AI or anything else mentioned in the article.
Yes, this is a result of copyright laws. The other commenters are wrong/uninformed.
If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc. There are essentially zero advantages, other than it is what is required under US copyright law (or at least, it is what their highly paid lawyers believe is required under US copyright law).
Isn’t it being destroyed because it makes the scanning process easier?
No.
A judge a while ago decided that as long as the physical copy is destroyed, and "transformed" into an electronic copy, you can do the upload. But if you preserve the physical copy after scanning it, you are in violation of copyright because you "copied" the book.
That's literally the only reason they are trashing them. It's a legal requirement.
Ah, so I can buy and scan a dvd, destroy it and then legally distribute the legal copy via torrents. Good to know, because that is what the LLM thieves are doing.
Distributing an exact copy of the text would be illegal; if you could get an LLM trained on a book to output the exact copy of the text from the book then that would be illegal as well.
They still made a bunch of copies and reused it to train multiple models after making the first copy. Unless they copy and destroy a book each time they use it as as training sample
Can you find a citation for this? I have heard this claimed rule recently from other people, and I haven't seen this decision (nor do I know what level of court or jurisdiction it might be). This is not a rule that I heard many years ago when working on and adjacent to copyright issues (including book scanning!), although of course the issue has been newly litigated again recently, so there may be new interpretations coming out.
Edit: Someone else linked to an order in Bartz v. Anthropic which appears to emphasize that destroying the original copies improved the defendant's position with respect to the fair use analysis. Is that the decision you're thinking of?
Perhaps partially? I assume that cutting the pages out of the spin makes them easier to scan at least partially. That said I have no insider knowledge of this type of operation so I don't know how much easier that actually makes it.
But ultimately, it's a moot point, because the legal requirement means the books must end up destroyed. Even if the people at Amazon wanted to scan the books in a way that required no destruction at all, it's not currently (legally) possible for them to do so, so they might as well take the easy way out today.
Non-destructive scans are at least 10x as much, and tend to lower quality.
https://software.annas-archive.gl/AnnaArchivist/annas-archiv...
Cutting the pages out makes them machinable. Non-destructive scans involves gently turning pages, and paying a lot of attention to the state of the spine. Destructive scans involve guillotine cutting the spine off, scanning the covers by hand, putting the pages into a hopper, clamping them in and hitting a button. While that book is scanning, you're already cutting the spine off the next book. If the machine jams, try to work the jam out gently, scan the pieces, and let the computer stitch it together.
Perhaps this is a use case that's worth investigating, to invent better, less-destructive scanning processes! Or some way to rebind them afterwards.
They could just not do the evil thing.
(This is why I will never be a billionaire)
Those laws are lobbied for by large corporations, these are not just laws that exist outside of that context. They can also be changed, or Amazon could just incur the fines.
Large corporations will move fast and break things when it’s convenient; they don’t care much about the law - just about profit.
I support physical media and doing whatever you want with said physical media. If you want to overpay for a copy of Sharepoint 2007 For Dummies and destroy it, knock yourself out. Just because media has been printed and is "rare" doesn't mean it has any practical value.
You could easily write an article about how Goodwill and the Salvation Army dump millions of "rare" books that no one would ever buy for $0.50. And thrift stores will dump multiple copies of said books. AI companies would only ever need one!
The only thing I disagree with is AI companies feeling entitled to freely use every piece of copywritten work without restriction, but that applies just as much to web scraping, pirated books, etc.
Discussions:
2 days ago https://news.ycombinator.com/item?id=49310725
21 days ago https://news.ycombinator.com/item?id=49068738
Why is the textual data bound up in these paper books worth the trouble?
These LLM training runs have already ingested essentially the whole public internet. What marginal value is to be gained from scanning and destroying obscure books?
For me the biggest functional issues with LLMs don't seem to have any connection with "I wish they had read this obscure community cookbook from 1946". Is that going to get Claude to stop saying "honestly" to me? Is it going to get LLMs to stop making up sources that don't exist? What is in it for Amazon or any LLM company to chase more obscure data like this.
I'm not an expert in LLM training, but I think we can all agree that the writing on the Internet is generally very low-quality and surface level compared to the depth of books. Most books in the past were even edited by a separate person from the writer!
Granted that that is indeed the case, surely the quantity of high-quality writing available online (including things like Project Gutenberg) still dwarfs the quantity remaining in obscure unscanned books.
Humans manage to be pretty intelligent with only reading perhaps a few thousand books in their lifetime. It seems unlikely that AGI will appear but only once it has read that last out of print 1983 book on knitting patterns.
Frontier researchers have found that dumping more and more data into training is effective at improving LLM capabilities. Nobody has a gears-level understanding of how training on some particular kind of data leads to some particular behaviors, so they generally take the attitude that more is better.
I have to agree; even if they are getting regular 20th century out-of-print books, is that going to add a significant percentage to their training data?
I can only think of it being a 'low-background steel' situation where they want to locate original, non-digitized text for validation or knowledge bases.
We really need to be thinking about the long-term implications of what Amazon, Anthropic, OpenAI, and others are doing here.
Because it's not just one company doing this, it's dozens if not hundreds of them, all with huge budgets and all competing over the same dwindling supply of older books. What those talking about library disposals and similar measures are missing is the sheer scale that destructive book scanning for AI ingestion is operating with. It's unprecedented.
It's not impossible that in a few years the number of older physical books for sale anywhere will plummet to almost nothing and that in many cases that'll include the last copies anywhere of particular titles. Second-hand book stores will close, and a ton of niche knowledge might be lost forever. It'll also all have been done in relative secret, without any transparency or public record kept of what was lost.
Worse, it's something that can only be done once. If our societies don't do something about this now there isn't going to be a chance for a do-over. Once a physical book is destroyed it's gone forever, and we'll be lucky if there's a digital copy left. But even then digital copies don't provide the same level of forensic verifiability that physical copies do. Future researchers looking for material that might've been contained in books like these will be out of luck.
If our societies do nothing to stop this, we'll all be poorer for it.
I don't understand why Amazon still bothers to blow money on their in-house AI effort (outside AWS Bedrock). What have all the scraping, scanning and training led to? Nobody is using their models.
The dataset is valuable on its own without the model.
Complete stab in the dark: AWS "training as a service" (further split into multiple microservices) for companies looking to train their own models.
AWS has nova and titan models that nobody seems to use
I didn't even know AWS had in-house models, that's how seldom they are mentioned.
A lot of the rare books seem to be books that are most relevant to specific professions within specific countries. I would expect university or national libraries in those countries to have copies of those books, but they often do not digitize them because controlled digital lending is seen as a legal grey area at best. So these ai companies are doing digitization for them and not publishing it, duplicating work for these libraries once they do decide to digitize.
Many other rare books are just rephrasings of other books on a given topic and this can be done with synthetic data generation now. Same goes for websites
Honestly this feels like a US exclusive problem. Like shipping from overseas would be way too expensive for the cheap books they want. The Chinese just use Anna's archive instead
A couple thoughts:
Would that an enhancement to the Library of Congress or such?
A couple answers…
Why do you think for even a moment that they would buy only one copy, rather than every available copy of something rare and difficult to get? Anyone doing this stands to gain from destroying the only copies of something that has scarcity, to stop competitors from ever getting access to it. That's a serious concern.
Why would they do that? What is this 'should' to a company that would do this in the first place? Are you not suggesting the opposite of what they aim to produce?
That seems really speculative. I think you're assuming the maximum possible evil intent for no reason, mostly because you hate AI and anyone involved in it.
I remember when google was scanning a bunch of rare books, I mean they might still be doing that? Either way, that was cool.
I have a few "rare books" and have read many, you'd be suprised at what is publicly available on google books since like ~2010ish.
Years ago I saw an article on this. For Google Books, they had two processes.
The first was destructive. This was for mainstream books currently being published so they had no value. It's (I believe) where you cut off the spine and scan the pages.
For rarer books, there was a non-destructive process. Basically the book was opened to each page and scanned. This was slower but didn't destroy the book.
I don't understand why these companies haven't just licensed the scans Google has already done. Why is each company doing this rather than just scanning the books once and sharing the scans?
> I don't understand why these companies haven't just licensed the scans Google has already done.
Because for the vast majority of the books, Google doesn't have the legal right to license those scans. It would be legal if the books were out of copyright, but despite the connotation of "rare books" those generally aren't the books we're talking about. Further, in the cases where Google didn't destroy the original of the book they scanned, their scanned copy may be considered infringing under the new standard, so Google doesn't want call undue attention to what they have.
yeah, I think most of this really stems from people streatching the use of "rare books". The cool rare books google scanned one page at a time were borrowed from a library because they are rare and or valuable and usually old (out of copyright by a long time).
If I remember correctly, it wasn't a distinction between common and rare books.
It was a distinction between library books and others. They partnered with libraries, and obviously, libraries didn't want their books destroyed, so they devised a non-destructive book scanner. Some of the books they scanned were indeed rare, but rare or not, you don't destroy books you borrowed from a library!
Here’s a clip of The Internet Archive’s nondestructive process [39 seconds]:
https://nitter.net/internetarchive/status/135809098218971955...
Wonder the additional cost to add automated page turning. More than Bezos can afford, impoverished chap he is.
Google probably doesn't have the scan, and now, if they do, it might be proprietary.
While we can quibble (and are) over the details of this particular occurrence, I'm beginning to take the opinion that used/rare booksellers need to implement some form of "KYC" to help guard against the epistemic threats posed by this practice.
The booksellers are probably so ecstatic to get some value out of books otherwise destined for the landfill that they wouldn't lift a finger for "epistemic threats".
According to the article, the scanning facility has been there for a while, so it may well be that they have been destructively scanning books for a while, even before AI. It makes sense for the major book retailer to have digital copies of every book so that it can begin selling those as soon as it can negotiate the rights (if there are any).
In any case, all the indignation about destructive digitization misses the point that rare books takings space in a warehouse for years without being bought will eventually be destroyed anyway.
Am I too old now, expecting someone to make a Rainbows End reference? Vernor Vinge predicted this 20 years ago.
(Also the person who coined Singularity, though Ray Kurzweil really wanted everyone to think it was his idea.)
Yep, unpleasantly close prediction. Probably shouldn't give them any ideas.
That is certainly an interesting case for thinking things but not typing or saying them out loud.
Problem is, they all read the same books we do.
Something I've been thinking about is that a physical book is a physical artifact, and it's age and authenticity can be verified by treating it as such.
In contrast a digital book has no such marks. It's impossible to tell if it was redacted to hide inconvenient passages, or even completely rewritten or fabricated (which can now easily be done at scale with AI). If the physical book is a primary source then the digital version must be considered a secondary source: A potentially biased retelling.
For this reason, a digital book can never be a perfect substitute for the physical book it was created from.
I watched Short Circuit yesterday. All I can think now is "need input".
Great article and work to reverse engineer who was mass buying used books.
>It'll live on if they publish those scans or contribute them to a national archives or something
I'm not really aware of many benevolent acts Amazon has taken in the past decade. Are you?
I'm less concerned with how these rare books didn't rot in large lots of unused books and more concerned with whether or not this helps preserve the content longer.
The content is never made available. It would be illegal to.
It will be legal to share in some decades. Not that Amazon will bother, though.
I don't expect an amazon.com/checkoutrarebooks page but I also don't really think the legality of sharing the content has been much of a concern for these companies either. Ironically, one of the few things Meta got in trouble about with building their AI was helping share the training data.
The content is preserved obviously, and I hope that sometime in the future it will be made available in its original form.
What worries me about the trends is inevitable sanitization of content or straight out falsification.
... because now it can be done at scale.
>What worries me about the trends is inevitable sanitization of content or straight out falsification.
That sounds really speculative, and not inevitable at all.
I think you're just trying to invent things to be worried about because you don't like AI and don't trust AI companies.
Well, 'history is written by the victors' is not just a toss.
I seriously doubt any of the companies doing this will preserve the scans for very long. It costs money to store stuff, and none of these companies are have any concern for anyone who isn't them, so they're not going to spend the money or lift a finger unless they work out a way to make it profitable.
archive.org estimates it costs $2/GB to store data in perpetuity (well, at least until decades of storage scaling trends stop of course). That's probably less than it costs to acquire and scan the data in, I'd be surprised if they just tossed it out at the end when it's so cheap to store. The same thing happened at the healthcare data lakes I worked on where the trend switched from the usual "how long are we legally required to store this information" to "how long can we legally hold this information".
https://web.archive.org/web/20260817141552if_/https://www.40...
So... do the materials that it is printed on make it rare? Or is it the combination of ideas, paragraphs, and words that comprise the text make it rare and therefore valuable? If they're releasing digital versions of this text, i feel that is better off anyway.
Well, I hate to be the one to break it to you but they're not releasing digital versions, they're just destroying them.
I look forward to the day that AI companies pay large advances for new books.
It'll take a while before publishing collapses due to the availability of the same information via an LLM, but once this happens new books will get a lot more expensive for AI companies and a lot cheaper for everyone else.
I've been lead author and co-author on several books on programming. My first book sold tens of thousands of copies even though the writing was not just trash but a crime against humanity against anyone who had to read it. This was back in the day when the source code for the book was provided on request on a floppy, which was actually quite rigid because it was a Mac floppy.
My writing got better and better, and I got better publishers who actually hired editors, but the books sold fewer and fewer copies. Even crappy self-serving poorly written stackoverflow posts are often good enough. Then LLMs killed stackoverflow. Is that real ironic or Alanis ironic?
But I'm not holding my breath waiting for a book deal from OpenAI.
> My first book sold tens of thousands of copies even though the writing was not just trash but a crime against humanity against anyone who had to read it.
I’m sure I never read your book, but there is something cathartic about reading an admission like this. I have no doubt people got value out of your books, but I distinctly remember as a kid saving up to buy a couple expensive programming books from the bookstore and then being sorely disappointed in the way they were written. At the time I thought I was too young to understand adult writing, but when I went back to the books later as an adult it was obvious they were just written by amateurs. I still cherished them and learned a lot, but I will always remember the struggle of trying to follow along with what was probably some first-time writer’s attempt to learn how to write as they went along.
> But I'm not holding my breath waiting for a book deal from OpenAI.
Was this title eligible for class action status in the recent Anthropic scanning case? I know people who popped up on the list decades after writing obscure or forgotten technical titles. The estimated payout is $3k per title, usually split between publisher and author.
https://www.authorsalliance.org/2025/09/07/the-anthropic-set...
There was some specific thing about which ID number, not ISBN as I recall, your book had to have to be part of the settlement. Some of my books qualify some don't.
Copyright office registration. If you or your publisher didn't do it, you don't qualify. Some publishers, even big ones, were exposed as neglecting this basic step when the Anthropic case was settled.
I would check the settlement db just to make sure: https://secure.anthropiccopyrightsettlement.com/lookup
You may entitled to a share assuming the publisher registered with the class action and there is a copyright registration in your name.
> Then LLMs killed stackoverflow.
I think the common consensus is that stackoverflow killed stackoverflow, quite a few years before LLMs became entrenched.
They are only destroying the books because they are required to by copyright law. They obviously wouldn't do so if they were allowed to merely copy the book and preserve the original.
If it was $1 cheaper to destroy the books even if they didn't have to, they probably would. Storing books is expensive and selling it on means it could be picked up and scanned by a competitor.
There's no copyright law that requires the owner of a copy of a book to destroy it. What are you talking about?
From the ruling of Bartz v. Anthropic:
“And, the digitization of the books purchased in print form by Anthropic was also a fair use but not for the same reason as applies to the training copies. Instead, it was a fair use because all Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies.”
I’m not a lawyer, so it’s possible I’m missing something. But it seems to me like the ruling here implies fair use only holds so long as no net new copies, digital or otherwise, are created. If that’s the case, then it necessitates the destruction of the original.
Not true. To the extent that it is fair use to digitize a work for various purposes, it is also fair use to keep the original. It is only if they wanted to resell the original that they would have to delete their digitized copy. The reason they are destroying them because removing the binding is the most efficient way to scan them, they have no use for the originals after they have been scanned, and don't want to spend money storing them.
No, they aren't, for training AI, at least not based on anything but pure speculation.
The recent trial court decision that keeps being pointed to to support that:
(1) Found that for training AI, digitizing and copying works was fair use, period, with no requirement to destroy.
(2) For creating a centralized digital library for general use, digitizing works while destroying the hardcopy was fair use even with no intent to use them for training AI.
It seems like most bookshop owners today are hoarders-turned-entrepreneur. They rent a storefront, adopt 2-5 cats, and they set up shop with their hoard of books so that they can hoard on a commercial level that would never be possible in a private home.
It also seems that American society has been swept away by the worship of books, rather than literary appreciation or promotion of education. "Banned Books Week" as the primary exhibit here. A tug-of-war over books that are supposedly "banned" when the verb itself has been twisted beyond recognition. I see posters and television shows and public service announcements that promote "Books!" and "Reading!" for no other purpose. Many books are trash and they will fill your head with garbage as sure as social media can, so why the indiscriminate worship of books?
It is conspicuous that, aside from established booksellers, and perhaps the LFL owners, many of the people crying out and moaning over book-destruction seem unwilling or unable to actually take those books in and store them. That is the point, that these books are unworthy of taking up storage space, which is an ongoing cost, and maintenance, and therefore, it is fiscally responsible to sell them, scan them, and destroy them. If you want to dedicate a room full of shelves to books that nobody will ever read, then go ahead and buy out your local bookseller! They will not stop you or cancel your order! Tell them you are saving those innocent books from the Big AI Bugaboo! They'll grovel at your feet!
Also now, I'm seeing small booksellers who feel "suspicious" or "skeptical" about large book orders. There was one Down Under who said "oh we got an order for over 70 and we can't physically handle that!" I feel like it's becoming a "shut up and take their money" situation. These booksellers, as professional business owners, should know that they may never liquidate stock, and this is their golden opportunity to simply get rid of some of that excess hoarded inventory.
But if you're merely cosplaying as a bookseller, and deep down you're a hoarder and a worshipper of books, you may really be reluctant to do transactions or earn money that sells away your books.
I remember when hackers believed in the doctrine of first sale. You can do whatever with the stuff you own.
In a strictly libertarian sense, sure.
In more liberal sense, because you might be destroying something unique that later generations might actually like to see, the life rule that states "Don't be a dick" probably applies trumps even first sale doctrine.
If there was a public interest in these books then it was already attached, and it was an outrage for the sellers to exclusively possess these books, and it was an outrage to sell them to any other private party.
More and more, I am glad that I stopped giving Amazon my (formerly) enormous amount of business.
The headline implies at some level that Amazon loves or cares about ..... stuff.
Nothing but a drawing of an alligator eating a book. Where is the article?
Link to URL? There is no article at the above link.
Does anyone else get pretty tired with 404s style of writing? I hate to make this comparison but it feels like a Fox News for the left.
I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it. It is an absolute shame for books to be destroyed but what does it really mean to be rare here? I know they kind of tried to differentiate but it sounds like this could be John Doe’s self help book that never sold well. If you ever are connected to a library you will start to realize how many books simply get thrown out or sold for nothing because nobody wants them.
To me the problem is partly the copyright law. I think it’s. Hard problem but I always lean more towards books having short copyright shelf lives and making it legal for digital copies to be shared after which I think would eliminate I good part of this problem. Not to mention 99% of the books published are probably garbage but that is highly subjective.
So it is bit of a meta rant but I think there a couple holes to go down that could be extremely interesting but they always write these informationally light articles. Like scrolling through a NYT visualization for just some shipping datapoints. Don’t really dig deep on anything and then end with a trust me bro these are rare books that Amazon is destroying for AI when I cannot be that upset with the amazons of the world. There are a lot of reasons for a business to digitize books, most books are worthless and it makes sense to cut the bindings for scanning. I would rather talk about how could you fix copyright to make this less an issue but is it even an issue with how many books get thrown out?
Not really. It's not to my taste, but I've found their journalism to generally be lefty but focused on important events.
Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
The difference is huge between 404 and Fox. Fox is out here trying to tell people there's a trans agenda, and that Biden was a lunatic leftist. They are just making up stories and publishing them because they know their audience engages. 404 definitely make editorial choices about which stories to pursue but I've largely found them to be grounded in real depictions of stuff that is happening.
Again, what rare books? Like they pointed to a book with low published volume or foreign language is “rare”. If you have ever been to annual library book sales you start to realize just how many books get thrown out. I can go buy a stack of books from the 1880s and 1890s on eBay right now for $25. $4 each. Most people don’t realize how worthless books are, even “rare” ones.
I even imagine that the market price for these “rare” books is helping filter out anything truly valuable and rare. It just reads as a rage bait tmz article. The quantity of used books including “rare” books that get thrown into the dump is astronomical.
I worked in a bookshop. I know all about this. I've seen pallets of books with covers torn off so they can be reported destroyed.
Here's a bookseller describing some of the books, which he believes are going to middlemen who are playing a sort of arbitrage with the AI companies by buying obscure titles, reselling them, and tossing away anything that doesn't sell. https://charliebecker.substack.com/p/is-an-ai-company-buying...
The titles mentioned there are not rage-inducing: "The Insider’s Guide to Metro Denver from 1995" and "How to Use Corel WordPerfect 1991". 404 writes and article about book destruction, and then explicitly neglects to mention a single title.
I suspect that's because folks wouldn't react the same way if they realized it was computer manuals for Windows 3.1. Even more obviously, Amazon doesn't want to be paying a lot for these books, so it seems outlandish that they'd be buying valuable rarities, since the booksellers would know the worth of those copies.
Good reporting on this would have included sales prices, volumes, and titles.
You don't know the books are being destroyed. There are automated scanners with page turners.
Most reporting says they are being destroyed. Page flipping is slow compared to cutting the spine and scanning the pages.
They destroy regular books. This was an intentional order from a rare book dealer. There is no reason to believe they use the same destructive process for books they should know to have inherent value and aren't purchasing in significant volume.
I don’t think we read the same article. None of these books were identified and they said rare can be low print volume or foreign language.
> Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
No, destroying collectible books would be a shame, not just any rare worthless books. But these are not collectible. The article tried to dance around it by saying maybe some books have a sentimental value to someone somewhere. But that doesn’t mean any library or collector wants it. Don’t fall for manufactured outrage!
Edit: Here’s an example of an extremely rare book. My great great grandfather published a book of sermons around 1920. That book has zero value to anyone other than my dad. Would I be outraged if it ended up at someone’s estate sale, then a used bookstore, and then an LLM consumed it to learn to read? No; I would have expected it to have been discarded by humans before the LLM even got to it. Most of what we leave behind is discarded.
This was my issue as well. It seems like they tip toed around it in one paragraph saying usually these are books with isbn numbers so not truly collectible or very rare and then in another paragraph plainly stated that it could be foreign language or low volume books that are rare. Rare expresses a different meaning to me and I think it amounts to exactly what you said, manufactured outrage.
We recently had a post highly upvoted here that collected bus tickets from an earlier era. It showed the unique moment in time where such fares were highly detailed, unique pieces of art. Ultimately destined to be single use and largely extremely pedestrian.
We've also seen a website that collects old paper restaurant placemats from across the ages. Literally disposable, zero value items that were made to be discarded by the thousands.
And yet, when someone has an interesting idea that ties them together, makes us recollect or think about our past in some novel way, these useless uncommon things suddenly become quite interesting.
A book on sermons, o its own, from the 1920s is maybe not that interesting. A book of sermons selected from each decade? A collection that compares regional books of sermons? A compare and contrast of the 2020s and 1920s? I can imagine many interesting thesis where a book like that becomes interesting because of the context of other books that are juxtaposed to it.
A book on Detroit motorways and bus schedules from the 30s isn't interesting or valuable on its own. But when you contextualize it, suddenly it might be a way to understand our history, our path through development and redevelopment. Connecting our present moment to the past.
Yes, I collected a couple decades of Muni fast passes too and finally gave them to an artist to destructively turn into some project. I was sad that my “rare” tickets probably won’t be kept forever, but I’m glad that somebody got some use out of it. I feel the same way about LLMs. I’m glad that they are learning from old published materials. In my opinion, 404media is cynically inciting anti-tech anger; they don’t otherwise have any interest in the preservation ecosystem.
You could donate that to the church archives of his denomination or the religious studies department of a university or at least some sort of local historical society.
I was in charge of a "lending library" ministry at my church for a few years. Let me tell ya.
We got started when another ministry moved out, and sort of from zero. My pastor's clear instructions to me were: make sure everything we carry is doctrinally sound.
So we inherited several full collections of books in rapid succession. Some had even belonged to priests and religious. Those gave me a fascinating time, because I could basically rubber-stamp every title that a priest had in his personal collection. But slowly the balance began to tip into rather esoteric volumes that normal laypeople couldn't really use. Literally books full of sermons and other arcane subjects!
I was tasked with discarding/recycling all the rejects. There were tons of rejects, believe me! So with every session when we had boxes full of donation, it was imperative to cull the bad stuff very fast, shelve the rest, and then find somewhere to dump the trash. The manager was encouraging me to recycle, or at least not tip them all into the Dumpster, but it turned out to be a logistical nightmare to find anyplace that would recycle books like that. Having no vehicle, I had to continually figure out ways to cart around heavy loads, just to get them out of church and into the trash somewhere. That was the worst part of my job.
Now it was clear that books were not a very hip or current medium, but there were plenty of elderly parishioners who did appreciate the resource and did compliment my work, but our church was not free of prejudice or judgementalism, and let me just say, there was an angel or entity whose purpose was only to jumble all the books while I wasn't watching, and leave a deliberately unorganized mess for me to confront every week. This made the task distinctly Sisyphean, in addition to the need to constantly discard rejected books.
I finally threw in the towel when large boxes of Spanish-language books were donated; there was no way at all for me to vouch or determine their orthodoxy, and the shelves were full anyways, and I was just tired of propping up a legacy ministry anyway. But it really drove home my opinions about books, hoarders, and that is why I have no troubles with the way books are currently being treated.
So it turns out your organization didn’t need so many donated books. The ideal outcome for your discards should still be sent somewhere to someone who wants to read or possess them, or barring that, to be scanned so their contents are at least preserved somewhere. It belongs in a library, even if incorporeal.
You don’t realize how many books are out there, most are worthless and nobody wants.
It still belongs in a library! At least a digital one.
Something can be "rare" while also having no value. One only needs to take a look at their local Facebook Marketplace listings to see this in action.
Rare invokes images of limited edition runs of well loved books, when in reality it's probably extremely outdated software guides, how-tos, technical manuals, etc.
> I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it.
I was going to say a very similar thing, and this is something I strongly disagree with when it comes to HN's moderation, and it goes like this:
The product is the outrage.
And HN should know better and mods should actively discourage, warn and prevent accounts (who karma farm, among other things) from even being able to post rage-bait articles. These aren't "hacker curiosities", they're just insipid bullshit. We wasted time and learned nothing.
I agree with you. And it sucks, because 404 writes about interesting topics! But they're so negatively polarized it often obscures what's actually interesting.
Should have let them torrent.
Just another reminder that piracy remains the absolute best archival strategy we have.
you do realize "rare books" in this context most likely means random technical manuals nobody cares about and not collectors items, right?
If most but not all, what number are ones that people and collectors do care about?
Any bookseller worth their salt will ensure that a truly valuable book will not languish on their shelf for 1 minute longer than it takes to find a buyer willing to pay a fair price.
If AI-scanners are somehow bid-sniping bona fide collectors and wealthy aficionados, there may be cause for concern. But that is most certainly not happening here.
How do you have confidence in that? These buyers are reportedly exceptionally non price conscious.
Personally when seeking out rare books with few copies in existance and extremely scarce availability, I have often found listings that have lasted for quite some time. Not all rare books go to auction. You might be thinking only of some extreme of notable works and not a wider spectrum of desired but scarce publications that does indeed exist contrary to your confident assertion.
because that ended up being the case the last couple times this topic came up and now we've gotten wolf-cried
the part where the article avoids mentioning what book this was is a tell
If collectors cares about them, the sellers wouldn’t have sold it for pennies.
These buyers are reportedly exceptionally non price conscious. What evidence is there that all books are sold for pennies, not just most?
Because they want quantity, and they don’t need specific titles, so they have no reason to pay more than the cheapest bulk prices.
Unless we know this to be the case it's reasonable to assume it might be more.
It's reasonable to assume the clickbait article with no information is actually important?
The books.
Why would that be a reasonable assumption? This is so blown out of proportion. The imagery invoked by the narrative is one of huge corporations destroying the final copies of literary treasures. That's just not happening.
> That's just not happening.
And that is a reasonable assumption?
I believe that is the rational default assumption, yes. Presumably there is some known list of books that they purchased and ran through these scanning machines. If there are true literary treasures that are genuinely hard to get access to on that list, then I would join the ideological crusade against.
Where is that list so we can be sure? On some internal systems we can't access. So I'd prefer rare books to not be destroyed if there's a risk that it could be the last remaining copy.
Oh, that's really interesting, I'd love to see the list of books that they've digitized too. Where did you find it?
follow the discussion around this on X, i think just a few minutes of research on this topic / reading past the headline you'll find out that this is pretty much a nothingburger
No, what I mean is that I want to know how to figure out what books they are digitizing so I will know what they're trying to improve their models on.
I didn't realize that the list of books was published somewhere, so that you know with confidence what they're adding to the collection.
How do I browse the list you're using for tracking this?
It is interesting to know! My assumption would be it's technical or scientific publications, maybe things like the Springer back catalog (assuming they've not already licensed these). Some of these are "rare" as in they are essentially published PHD theses with very little in the way of sales / print runs.
Of course I could be wrong and they are destroying 14th century monastic scrolls or out of print Mills & Boon editions.
They are rare because nobody cares about them otherwise. Why is everyone acting as if they are trashing Gutenberg Bibles or first edition LOTR copies?
Carnegie was building libraries and concert halls, Silicon Valley CEOs buy girlfriends with big tits, cheat in computer games and destroy books.
Bad enough to steal IP but destroying books…that’s just evil
To me, the hate against "destroying books" is kinda silly. The reason we're against destroying books is because the Nazis did it to destroy information.
This situation has absolutely no relation to that. This is the exact opposite, and the only reason this information isn't available publicly and is in risk of getting lost is copyright law.
TLDR: Amazon isn't the nazis in this story, copyright law is.
Making the verbatim content of the book lost forever is not equivalent to burning them? I'd even wager that the CO2 emissions of the scan,shred, and train process is even higher than burning.
> Making the verbatim content of the book lost forever is not equivalent to burning them?
The verbatim content are the words, not the paper.
Books are lost all the time because the last book ended up in a landfill. But if an AI lab digitizes it, now it's stored in an extremely redundant storage lake in a datacenter and the company has huge incentives to make sure they don't ever lose that data.
So is it moral of me to destroy a copy of "Windows 95 for Dummies"?
What if a person with small children and an elderly, incontinent pet with a penchant for peeing on books wants to buy it - can I sell this book to such a dangerous purchaser who might destroy it?
I love book as much as the next person but the hyperbole about "rare books" is absurd. Nobody is buying the Gutenburg Bible and destroying it for AI. The books in question are certainly not rare enough to be in museums - without titles there's no proof these are anything of real value.
But rich people want to violate copyright now! So judges ruled that it's now legal. But, because "the law is the same for everyone" they needed some excuse. NOT because, you know, obviously this demonstrates that very rich companies get to violate the law and you get hit by $30000 per infraction when you do it, when Anthropic ... doesn't even have to stop violating copyright when they do it.
Anyway ... this case is:
Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA, U.S. District Court for the Northern District of California, decided by Judge William Alsup
“the purchased print copy was destroyed and its digital replacement not redistributed, this was a fair use.”
https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...
Obviously this violates precedent, I had an internal LLM (probably a frontend for Claude or ChatGPT) find them (just like the convictions for file sharing in the 2000s required counter-to-the-law reasoning by judges, fair use was almost never accepted as a valid excuse, even when it obviously was, but of course Sony was a billion dollar company and needed to be in the right. In fact that this had to happen was explicitly given as a reason to create the DMCA)
Anyway, some precedents:
Hotaling v. Church of Jesus Christ of Latter-Day Saints, 118 F.3d 199 (4th Cir. 1997)
“Although the Church acknowledges that its sole remaining copy is not the one it originally acquired … it maintains that the remaining copy does not infringe Hotaling's copyright because it is a replacement copy…”
(this reasoning was rejected by the court)
https://law.justia.com/cases/federal/appellate-courts/F3/118...
Atari, Inc. v. JS & A Group, Inc., 597 F. Supp. 5 (N.D. Ill. 1983)
... defendant sold a device for making backup copies of copyrighted Atari cartridges and argued that §117 permitted replacement/archival copying. The court rejected the broad replacement theory.
https://law.justia.com/cases/federal/district-courts/FSupp/5...
This very court has clearly declared that making a copy of a copyrighted work for replacement purposes is illegal, on multiple occasions.
I would like to point out that this isn't Anthropic's only extreme-WTF law violation. When the original judgement against them was made against them, they were forced to admit that using books to train models was illegal if acquired illegally AND THEN WERE ALLOWED TO KEEP DOING IT (thankfully the court never mentioned that part in the judgement so at least they can claim that was never decided when it becomes a huge problem in future cases as it obviously will). But that's not how this works. In my opinion Anthropic and OpenAI and everyone else need to at minimum take training material that was acquired in violation of copyright out of their training data unless and until they have a separate licensing agreement with the copyright holders. As long as Claude knows more about Harry Potter than is said in the promotional summaries it is obviously in violation.
Because of the copyright-filesharing court wars of the 2000s, which were also handled dishonestly by courts (whether we're talking US or EU courts), and the absurd copyright extensions, they had to now make some new excuse, and settled on this very sad, very destructive option. It's not even defensible legally, imho, but of course the biggest wallet must win. I don't understand. It's such a sad joke at this point, and it's not like the courts even still had credibility after the file sharing cases.
I wonder which sad excuse will be forthcoming from the courts when we have someone release a movie made by an AI model that is obviously a direct ripoff from some high-budget studio movie, and Disney needs to be protected from ... say ... "Scorched: Brothers of Aridelle" — In a vast desert kingdom, the royal brothers Elias and Anders grow up together, but Elias secretly possesses dangerous fire magic and isolates himself after accidentally hurting Anders as a child; years later, at Elias’s coronation, Anders announces his engagement to the seemingly charming Princess Hanna, provoking an argument that exposes Elias’s powers and sends him fleeing into the dunes, where he accidentally unleashes an endless heatwave that dries the kingdom’s wells and turns the capital into a furnace ...
I can think of a couple of solutions: 1. Glue a US $1 bill to the spine and take them to court when they destroy it. 2. Send them a license to use the book, with the condition that if they fail to return it within 30 days, they owe $1M. (Hey, if e-books can be licensed, why not physical books?)
In the end its not who can train the better model its who owns the better stack of training data. Acquiring old texts then destroying them makes their information proprietary.
How many AI companies are doing this? Say there are only 5 copies of a book left out there but 100 different companies want to shred it. Rare items could effectively vanish from the market. "Rare" meaning inclusive of high-quality items, disregard the bot opinion that mass book destruction is fine because most books are worthless anyway. And are any safeguards in place to make sure scanned books get saved in their original form for posterity before being recycled into trainingslop? Aside from Google's quasi-legal/ethical mass book scanning operation of course.
> "Rare" meaning inclusive of high-quality items,
This is the citation needed that is missing from every report so far, including this one which deliberately refuses to reveal anything about these books.
You’d think if there were examples of actually valuable, rare books being shredded that the journalists would at least be able to name one such example. Instead it’s always vague posting about the destruction without ever naming any examples.
I think it’s because if they named some example titles, everyone would see that they don’t care about these books being shredded.
On the contrary these companies should be publicly listing the name/info of every book they destroy and use as training data.
If the books are really worthless as you say then their case would be proved transparently.
This article presumably had to maintain confidentiality to protect the seller who agreed to place a tracking device in the shipment.
I would say the burden of proof is on the companies destroying human cultural heritage en masse, not the handful of journalists calling for attention to the matter.
> I would say the burden of proof is on the companies destroying human cultural heritage en masse, not the handful of journalists calling for attention to the matter.
You're literally assuming the conclusion. The exact topic under contention is whether they are "destroying human cultural heritage".
I do work adjacent to the AI book scan-shred pipeline. There are definitely significant books that aren't "Windows 95 for Dummies" which are getting down to single-digit remaining copies.
I just looked for one novel, which wasn't a fantastic book, but it is the first use of a pithy and fun phrase that is so ubiquitous that you'll probably read it a couple of times today. I argued with Claude, GPT and Gemini for ten minutes just now, even knowing the title of the book, to even prove the book exists. It took me years to find a copy originally and then I lost it in a move. I found one more copy today from a rare book seller, but it just sold (to Amazon?).
Is the book valuable? Not particularly, but I feel it's noteworthy and important. I don't want to name it either, because now I have some searches out and the next copy that pops up I'll scan and put on IA. There can only have been a few thousand copies originally published in 1947, it's only in hardcover. I know of a couple of other copies in private hands, so it's not zero copies, but it has to be single-digits.
I have one periodical issue that I know of only one other existing copy (Worthpoint only shows one copy ever sold in their database) and if you look on collector sites there is a blank because nobody even knows what the cover looks like. I can't explain it, since the publication routinely printed hundreds of thousands of copies of each issue, but here we are. Perhaps all the copies were withdrawn and pulped immediately after publication for some reason? It's in my scan pile, so I'll have it uploaded soon. Is it significant? Not hugely, but every other issue of this title has been scanned already, so it's scratching an itch to get this one done.
There's definitely rare stuff getting scanned and shredded. Someone in a comment above said it's not like Nazi book-burning since they were trying to destroy information. But it is like that if you consider there remain no other physical copies and all the electronic copies are locked up in a way that nobody can access except to trick an LLM to spit out a paraphrased copy from its training data.
TechCrunch is trying too hard to make a connection in the title there. Amazon was trying to make money then as it is now. Their selling of books at the beginning was no more principled than the destruction of books now.
This is the legal loophole that allows them to do what they need to do, and the benefit of doing it outweighs the modicum of outrage this title will generate.
Edit. Because I see my statement confused all the HN experts:
Anthropic's version of this was Project Panama [1]. The destruction of the original allows them to keep a digital copy in the AI training dataset. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
[1] https://en.wikipedia.org/wiki/Project_Panama
The worst part of Amazon destroying priceless old books in the quest to build the torment nexus is the hypocrisy
all due respect but if they buy them they can do whatever they want with them. not like they are people and not like you have a right to them either. not even like you would have been able to acquire them yourself.
You're not really responding to the post, which is about the hypocrisy.
"They're allowed to be a hypocrite" doesn't mean they aren't a hypocrite.
Okay, let's say some books deserve the equivalent of UNESCO status. How would you, personally, go about this?
They obviously aren’t priceless if Amazon is buying them for pennies.
Is it a “loophole” that you can buy something and burn it? That’s how property ownership works. If the books were really so rare and valuable then the original owner should have given them to a museum rather than sell them as scrap. Yet no one who is complaining right now cared about them before Amazon got involved.
No, that's not the loophole. The loophole is that a US federal judge recently ruled that scanning a book for the purpose of training an AI is "fair use" and hence legal _if and only if_ the book is destroyed in the process. If you retain the original, you would be creating an unauthorized copy, but if the original is destroyed, then it is permissable as format shift: https://fingfx.thomsonreuters.com/gfx/legaldocs/jnvwbgqlzpw/...
Sometimes original owner does not know that it is a first edition single copy. For them it is an old book. There have been many cases where owners did not know true value of their antiques.
Well in this case the original owner doesn’t know, the purchaser doesn’t know, the reporter doesn’t know, you and I don’t know. So what actually makes these books rare?
> Is it a “loophole” that you can buy something and burn it? That’s how property ownership works.
It sounds shocking if you're missing the context and only rely on the blogspam which is this TechCrunch article.
They aren't buying the books only to destroy them but to build an AI training dataset, which means they're making a digital copy. This becomes a copyright issue. The destruction is the legal loophole that allows them to keep the digital copy as the only copy in circulation and be considered fair use as decided by a judge (see below).
The concept of ownership means you can do anything you want with the object, the book in this case. Not with the content, like make or distribute copies. The scanning machines are cutting the spine of the book, feeding the pages to scanners, and then destroying the physical copy.
Anthropic's version of this was Project Panama [1]. Quoting from the page:
> Judge William Alsup ruled that the destruction and digitization of legally purchased books constituted fair use
[1] https://en.wikipedia.org/wiki/Project_Panama
Its somewhat ironic.
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
So books that would probably have ended up as trash. These AI training facilities are actually doing these book a service. Not only they are probably going to keep the scans safe (for future training), but having the book end up in an AI model may be the only way it is going to have any use at all.
Let's say for instance that the book in question is about woodworking, and it is not great, lots of mistakes and inaccuracies, unoriginal content, etc... except for a single thing, maybe a trick for making a certain measurement or something like that. Who would read such a crappy book for this single good trick he doesn't know is there, well, an computer will, computers process terabytes of crap without tiring and complaining, that's what they are for, and with a well designed LLM, that one trick may resurface, waiting for someone to ask about that specific measurement.
The problem here is not that rare books end up in AI training facilities, it is that these AI facilities are owned by for-profit companies keeping the data to themselves. These books should go to public libraries instead, for everyone to access, it would be better if these books weren't destroyed in to process too. But the question becomes: why didn't public libraries didn't do that in the fist place? And maybe in a more respectful way. The AI companies would just have had to license the database to libraries, probably simpler and cheaper than having their own scanning facilities.
To me, this mess is a failure of the copyright system. One one hand, large scale book digitization projects intended to preserve and make the original text accessible get lawsuits by publishers, while AI training is "fair use". It means we have built a system that encourages destroying rather than preserving these books!
Yes, we've inadvertently set up a system of incentives that were designed to preserve information and monetize it. And we've instead set up a set of incentives to make it scarce and destroy it.