AI companies buy and destroy physical books to train models, says Anna's Archive
AI companies destroy physical books – let's scan rare books before it's too late
AI companies are secretly buying millions of secondhand books, scanning them, and destroying the originals to train their models, according to a guest post on Anna's Archive. The post cites Anthropic's 'Project Panama,' exposed in a $1.5 billion copyright settlement, as an example. Anna's Archive calls on volunteers worldwide to scan rare books before they are lost forever.
It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.
- ezfe
I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
- HedonicEscal8r
The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art.
I support Anna's Archive, by the way. Information wants to be free.
- skeledrew
Maybe begging the question here. If a physical book is rare, doesn't that mean it wasn't available to many in the first place? It seems to me providing its knowledge via LLM, even if it's a private company, benefits more people than if it were sitting in a library somewhere maybe read by a few, or worse in some private collector's set.
I can't help feeling there's some hypocrisy or something here with this call to be outraged at AI companies and scan books now. What about before when they were still mostly locked away from the world? It's only when they're actually being made available to - at least a part of - the broader world that they're a "cultural heritage" worth preserving. Shame.
- tgsovlerkhgsel
Many countries require a copy of each book published there to be submitted to their national archives/library. The US has had this requirement since 1790 as far as I can tell.
All countries I checked (Germany, France, Canada) seem to have similar requirements. WIPO claims that this is the case in the majority of countries (https://www.wipo.int/documents/d/copyright/docs-en-registrat...).
Thus, most books should not be at risk of getting permanently lost due to these practices.
This is Anna's Archive (ab)using a current controversy (which I think is based around emotional appeals and distorted facts to create outrage over a non-issue) to tongue-in-cheek advertise their open pirate library.
- sieve
Physical books and digital content is special in that you can mostly archive their content almost permanently for cheap. Buildings, paintings, idols, living things, natural features of the environment ... not so much.
So the solution is:
- mandatory copyright registration and renewal with links to where the work can be acquired
- a blanket carve out for any non-commercial trust-style org so that they can scan books etc and keep the data on their servers. They should be able to issue digital membership cards for a fee so that patrons can access the archives. Any work that is "live" based on the registration database will be locked. All "dead" material can be shared with members.
In this way, a hundred digital preservation societies can bloom.
- glimshe
The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
- xvxvx
Pretty funny that they just took Anna’s archive and ingested it.
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
- jupp0r
I highly doubt they destroy digital copies of the books after scanning. They will want to train their future models on the same content. So what prevents them from making these digital copies available to the public? Copyright!
- shakna
I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
- thuruv
I am baffled at these practices and somewhere confused on what's the end game here? monopoly on information? altering data? exclusive subscription based knowledge? Feels like we have welcomed the AI era with open hands hoping( at-least assuming) that data democracy will be there, yet feels like its a long road!