Tag Archives: Judge William Alsup

Flame on

That is the setting I saw a few days ago, for the most I ignored it for obvious reasons and I will get to that. But ‘suddenly’ the non-book readers are in a bind, someone is destroying books and Anthropic is pointed at as the guilty party. So here is the first thing. When you acquire a book, it is your property and it is up it you what you do with it. Anthropic bought millions of books, they cut of the back of the book and then scanned the book, after which they destroyed the book they had acquired. So far so good and it leaves me with a few questions. And whilst the (so called) book lovers go for the Fahrenheit 451 scenario, the people who oppose AI are in a fritz. I have different questions. 

  1. Is an owner allowed to set a book to digital format?
  2. How rare is the book?

The second question is linked to the rarity of the work. We can assume that hell breaks open if Anthropic does this to the Magna Carta (1215) or the Gutenberg bible (1454), but how many will cry havoc if someone does this to Joop ter Heul (1889) or a famous five book (1942) or even a Harry Potter (1997) book? There is a setting that the editor is responsible for keeping the book available and if they do not, it is what there is left. And when this happens it means no one is interested in that book. The first question is linked to the rights set in the settings of copyrights. Can a book be replicated digitally? And if so, what is the problem? 

So here comes the article (at https://mashable.com/tech/anthropic-ai-book-training-destroy) in Mashable with the headline ‘AI companies keep destroying old books. Here’s why.’ And as I see it there is so much white noise in all this, that the bottom question is ignored. Does Anthropic have the right to replicate a work into digital format, if so, what is everyone crying about? If they are not allowed to do that, the law is broken, but merely in the digital setting. It is still their book and they could shred it for all they cared for. It is the cold reality of commerce taken legally out of context. And you all know this (especially the previous generation). How many have recorded an album to tapes? I know I have. I prefer the actual CD, but before 1980 I had no income and for the most no music. So when we see this setting, how many have copied a book? (I admit I have copied a few manuals in that past) but that went away when Borland released its products with manuals and it felt really good to have Turbo C with a manual and all for $249. But that was then and now we do not see ‘value’ in books and it is often rejected with the Fahrenheit 451 label, but the reality is there. This leads me to a simple question. Ask yourself, how often have you been to a library in the last month? That should give you the part you need to know, so whilst the article gives us “AI companies looking to build better AI models are hungry for fresh data from any source that isn’t the internet. Books — generally better edited and more cogent than your average Reddit thread — make AI sound smart. (There’s a premium on books published before 2022, ironically because we can’t be sure if books were written by AI after that date.)” as well as ““legally-binding nondisclosure agreement” that would hide an AI company client’s “identity and strategy.” Why? ISBNdb explained: “Destroying millions of books evokes images of burning libraries … the optics problem is real. ‘AI company destroys two million books’ is not a headline that generates sympathy.”” But in all this, are they breaking any law? If that is not the case, why get fussy about it? For the sold book is revenue for the writer and the publishing house, so the issue remains, is there permission to reset a book to a digital format? Because that hurts the writer and the publishing house in their pocket and lets be clear. Writing is a commercial enterprise. In doubt ask JK Rowling. Apparently “She earns an estimated $60 million to $80 million per year from book royalties and digital sales alone, with total worldwide sales for the Harry Potter series surpassing 600 million copies and grossing over $7.7 billion globally” giving us a clear view that she made more than the writers of the bible, which is said to be “a collection of 66 books written by about 40 different human authors over a span of roughly 1,500 years” and I leave you to wonder how many of those 40 writers ended their lives in the poor house (just to make a point).

And then we get to the part that (kinda) impacted me “But Anthropic was caught, and has admitted to using millions of pirated books to train Claude. That class action lawsuit, in which a record $1.5 billion copyright settlement was just approved, also revealed the scale of Anthropic’s physical book-destruction operation. Anthropic “became convinced that using books was the most cost-effective means to achieve a world-class LLM,” wrote U.S. District Judge William Alsup in a lengthy legal order dated June 2025. The company was “not so gung-ho” about using pirated material from 2024 onwards, to quote a memorable internal email in evidence, but still wanted to train Claude on “all the books in the world.”” As such, as I did write over 4000 literary (an exaggeration to be sure) works. Not the number, I am now at 4100 articles, so where is my money ($5M post taxation would be decently nice and highly appreciated) and there is clear view of the transgression, no one reads 1700 stories in an hour, it just doesn’t go well for the brains (not even my brain and I wrote the stuff) But the issue remains, and there is no clear setting. Is there ‘policing’ on the scanned works? Is there a copyright issue? Because that matters. If that issue does not exist, why are we crying over a book no one has read in years, optionally not read in decades?

Makes you wonder, doesn’t it?

So have a great day and feel free to dream about burning all the books you have to keep warm today, or even warm up your mother in law to 451F. 

Leave a comment

Filed under Law, Finance, IT, Media, Science