Copyright wars

When there is any conversation about AI, it always involves the word ‘billions’. Billions invested, billions spent or billions made. Now, there are billions to settle.

This came on July 20th — ‘US judge approves Anthropic’s $1.5 billion settlement of copyright lawsuit’ and it wasn’t the first lawsuit and definitely not the last one.

This latest lawsuit was about Anthropic downloading a dataset containing copyrighted books and training its models on it. Anthropic was arguing that it was a Fair Use because the training was ‘transformative’. The issue was that Anthropic downloaded the content — over 7 million books — from what one can call with euphemism ‘shadow libraries’, without paying a dime for it. Anthropic also bought vast amounts of books, took them apart — aka, destroyed them — scanned them and used them for training.

The authors, while sad that the physical books were destroyed, complained that Anthropic used the pirated copies for the training without permission from these authors.

You can read the class action settlement here or visit the Copyright Settlement website here for all the information about the lawsuit and the settlement.

Anthropic found that it was cheaper to settle than to admit to any wrongdoing. But, at the same time, it created a conundrum for itself at the same time. Since Anthropic claims that the work was ‘transformative’ it is now legally responsible for the new content, which includes the case when the transformed content is incorrect. Presenting wrong information by AI is commonly described as hallucination.

You would be very correct to argue that yours truly is not a lawyer nor a legal expert to come to this conclusion. And you would be correct! That would put you on the same side of argument with Google, where Google argues against the ruling at the German court. In June of this year, ‘Google to challenge German ruling saying it is liable for AI-generated false claims’.

That case was brought by two publishers, who claimed that the AI overview linked them — wrongly — to scammy business practices. Google was arguing that it can’t be held responsible and that most of the AI Overviews are accurate. Except when it’s not.

All these companies using content to train their models are claiming that it is a fair use but when the ‘transformative’ output is wrong, they deny any responsibility for the accuracy and claim that it is meant only for entertainment

To make it even more absurd, Anthropic is complaining that Chinese AI companies are stealing the capabilities of its models. And to add to the pyramid of absurdity some people justify the use of these Chinese models with the argument that these models are ‘open’, versus Anthropic or OpenAI’s closed models. The question I would ask the same people is ‘Where do you think that the training content came from?’ #chokingonirony

And since we are on this topic, let’s review another lawsuit. This one was brought by Google against SerpApi. I am sure that very few people know about the existence of that company. What does it do? SerpApi is in the business of scraping the search results from Google search pages. Say you want to conduct an automated, repetitive search about a topic of interest. You want to learn about advancement in your industry or your competitors. You have a choice — either repeat the same search several times a day manually or use SerpApi to automate the process.

Google didn’t like it because it claimed that by doing this automated scraping it loses the opportunity to make money from advertising. Google argued that a) some of the content which was showing was copyrighted and b) it created measures to prevent SerpApi from doing the scraping and because of that SerpApi violated the Digital Millennium Act.

You can read the judgment here

The judge dismissed the lawsuit and agreed with SerpApi that “for a technological measure to ‘effectively control access to a work’ it must, among other things, ‘require the application of information, or a process or a treatment, with the authority of the copyright owner, to gain access to the work.’”

The judge allowed Google to amend the lawsuit to demonstrate that the owners of the content expressly asked Google to create a protection of their work. I am sure that Google will now contact the owners of every piece of content ever crawled and indexed, get a confirmation that it legally included this copyrighted content in its index and the owner of the content asked Google to protect it. Pigs will fly before that happens.

The recurrent pattern. All these ‘AI companies’ would like to get free content, use it to make money and yet to avoid any responsibility for the actual product they are providing to their customers. I think their level of denial prevents them from appreciating the cognitive dissonance they are experiencing. The upcoming lawsuits will help them with that.

Next
Next

The Wand of Wonders