Home AI Sorry, LLMs – Congress Might Make It A Whole Lot Harder To Train On Copyrighted Content

Sorry, LLMs – Congress Might Make It A Whole Lot Harder To Train On Copyrighted Content

SHARE:

AI bots are hungry.

They’re scraping information found on passports and credit cards and training on novels without authors’ consent. Even fanfiction has been used to train some (presumably quite nerdy) bots.

But a bill proposed in the Senate a few weeks ago could change that.

In late July, Sens. Josh Hawley (R-Mo.) and Richard Blumenthal (D-Conn.) introduced a bipartisan bill called the AI Accountability and Personal Data Protection Act. If passed, the bill would mandate stricter regulations for AI training on copyrighted material, including establishing a federal tort (i.e., a harmful act with legal liability) for misuse of personal data and determining specific remedies for damages.

While the bill protects data and materials that belong to individuals, rather than larger entities, publishers and other businesses that work with creators would be affected, too, if they had published any work copyrighted by an individual.

The material in question ranges from personal data, like browsing history and IP addresses, to copyrighted materials, like books and paintings.

You can read the full text of the bill here.

I’m just a bill

The question at the heart of the bill is whether using copyrighted material for LLM training is fair use, said Chris Mammen, a partner at law firm Womble Bond Dickinson LLP, told AdExchanger.

The bill doesn’t seek to displace previous rulings or suggest that there’s no such thing as fair use of copyrighted material, Mammen explained.

“Fair use” is a broad term referring to any permissible use of copyrighted works without needing a license or explicit consent from the creator, often for purposes like research and reporting.

Now, lawmakers want to create a clearer definition of what is and isn’t deemed fair use, as well as a baseline for penalties if personal data or copyrighted material is used unjustly. The bill would call for compensation equal either to the actual financial loss suffered, three times any profit made from exploiting the data or $1,000 – whichever of those three is greatest.

Rules and regulations

The good news is that there’s already a four-factor test laid out in the Copyright Act of 1976 to determine whether a given way of using someone else’s work is fair use – which means it should be simple enough to figure out how to apply this to individual use cases, right?

Apparently not. As it turns out, Mammen said, “it’s not a very easy question, given the way the four factors are articulated in the statute.”

Those four components include:

  • the nature of the use (whether it’s commercial or personal, or taken verbatim from the source or paraphrased);
  • the nature of the work (creative or factual);
  • how much of the copyrighted work is used (which can be interpreted to mean how much of a work was input into an LLM or how much of an original product was used in the output);
  • and the market impact (i.e., whether it’s basically a knockoff and displacing a preexisting work).

Interpreting these factors isn’t exactly cut and dried. In June, Mammen said, two judges “reached starkly different conclusions” regarding the market impact of LLMs training on copyrighted works.

In Bartz v. Anthropic, Judge William Alsup, a district judge for the Northern District of California, determined that the training was fair use, since the LLM in question had not generated a knockoff or imitation of the original books in question.

However, in Kadrey v. Meta Platforms, Inc., Judge Vince Chhabria (who also practices in the Northern District of California) proposed several ways that LLM training could potentially harm the market, including the fact that if AI models eventually create similar works due to training on the originals, that would inherently create competition and the potential to replace them.

Bills to pay

But although almost everyone believes that content creators are “entitled to some sort of compensation,” Mammen said, and have the right to require permission before their content is used to generate new, similar works, the existence of a vague moral obligation isn’t enough to establish legal rights or protections in court.

Still, some progress is being made on the compensation front. In June, the IAB Tech Lab proposed a new initiative that would offer publishers more control over how LLMs use their content and how they would be paid.

Around the same time, Cloudflare implemented a new model to block AI crawlers from accessing content without express permission and a preset form of payment.

But what about data that shouldn’t be used at all, even for a fair price?

The sincerest form of flattery?

The way that generative AI processes data isn’t really comparable to the way that humans, or even other machines, have used existing content in the past – hence the need for new regulations.

Historically, data has been primarily used for analytics or automated decision-making, said Mammen, rather than generating new content.

While creating art from someone else’s creative outputs isn’t a new concept – “like cover bands,” he said, “or people who make new art in the style of somebody else” – generative AI brings it to a new level.

What’s really “giving us some pause,” said Mammen, “is the fact that AI can do it at scale with great fidelity and with great speed.”

Still, despite the hesitations voiced by attorneys and lawmakers alike, the future of this bill remains to be seen.

Once a bill is introduced, it’s a “long, long journey to the capital city” – and there’s no guarantee it will become a law, especially considering the current administration’s goal to eliminate “bureaucratic red tape” around AI development.

But although only a first step, this bill takes the widely held belief that creators deserve control over the use of their work and turns it into a concrete call to action.

Must Read

New WBD Report Makes The Case For Getting The Measurement Basics Right

Warner Bros. Discovery has a new white paper analyzing the data and methodologies of five top video measurement providers: VideoAmp, iSpot, Comscore, Innovid and Samba.

Monopoly Man looks on at the DOJ vs. Google ad tech antitrust trial (comic).

Google And The DOJ Filed Their Proposed Final Judgments In The Ad Tech Case – Here’s What They’re Still Arguing About

Google and the Department of Justice filed the next round of paperwork that will determine what Google’s punishment will look like in the ad tech antitrust case.

T-Mobile Brings Its Mobile Data Exclusively To Vistar To Scale Up DOOH Targeting

Advertisers can now use Vistar to activate both off-the-shelf and custom audiences built on T-Mobile’s first-party location and app data.

Privacy! Commerce! Connected TV! Read all about it. Subscribe to AdExchanger Newsletters

Apple Has Far-Reaching Plans To Block Hundreds Of Programmatic Data Companies From iOS

Apple’s WebKit crackdown appears to extend well beyond The Trade Desk, putting hundreds of ad tech, data and identity vendors on a mysterious, dynamically updated block list.

Josh Reed, Zoom's VP of brand and content, speaking at AdExchanger's Programmatic IO event in New York City (September 28, 2006)

Zoom’s Marketing Challenge Is That It’s Too Well Known For Its Own Good

Zoom has 99% unaided brand awareness, which sounds great on paper. But there’s a catch: Most people still think it’s just a video-call app.

Why Agencies Think They Shouldn’t Own Agentic AI Tools Or The Data Used To Build Them

Agencies are differentiating their tech stacks by building custom agentic AI tools for their clients. And they’re rethinking owning those AI tools – particularly since licensing them creates new revenue streams.