Tech

Firms like Meta and A16z admit having to pay billions for training data would ruin their generative-AI plans as they fight new copyright rules

Mark Zuckerberg and Marc Andreessen
Meta CEO Mark Zuckerberg and the Andreessen Horowitz cofounder Marc Andreessen. Reuters
Read in app

The world's biggest tech companies really do not want to have to pay for the enormous amount of copyrighted data needed to train the models underlying their generative-AI tools.

During the ongoing comment period opened by the US Copyright Office as it considers new rules for generative AI, firms such as Meta, Microsoft, Google, Apple, OpenAI, and Andreessen Horowitz, along with news organizations, media agencies, and related individuals, were among nearly 11,000 commenters. The Copyright Office asked in its notice for input on creating a licensing regime or some other process that would "remunerate copyright owners and/or creators for the use of their works in training AI models."

Most tech companies seemed to agree that being required to pay for the huge amounts of copyrighted material scraped from the internet and used to train large language models behind AI tools like Meta's Llama, Google's Bard, and OpenAI's ChatGPT would create an impossible hurdle to develop the tech.

"Generative AI models need not only a massive quantity of content, but also a large diversity of content," Meta wrote in its comment. "To be sure, it is possible that AI developers will strike deals with individual rights holders, to develop broader partnerships or simply to buy peace from the threat of litigation. But those kinds of deals would provide AI developers with the rights to only a minuscule fraction of the data they need to train their models. And it would be impossible for AI developers to license the rights to other critical categories of works."

Google, Microsoft, and OpenAI made similar arguments — that the amount of data used to train their models is so massive there is no way they could figure out a way to pay for it. None of the companies denied using copyrighted material without authorization from rights holders. Instead, they generally argued that putting copyrighted material on the internet makes it "publicly available" and therefore fair game for use. Using that data to train an LLM constitutes "fair use" under current copyright law, the companies added.

Google referred to the copyrighted material it uses to train AI tools like Bard as "knowledge harvesting," contending that current copyright law is intended to allow such harvesting to happen. Holding a developer like Google responsible for the use of copyrighted material in training "would impose crushing liability on AI developers," the company argued, adding that generative AI was about the "free flow of ideas."

In addition, as far as Andreessen Horowitz, the venture-capital firm also known as A16z, is concerned, the billions of dollars it and other investors have pumped into the AI craze should be reason enough not to create any new rules meant to benefit copyright holders.

This investment has been "premised on an understanding that, under current copyright law, any copying necessary to extract statistical facts is permitted," A16z wrote. Upending that understanding, or assumption, "will jeopardize future investment" in AI, the firm said. It also argued that any kind of licensing regime for the use of copyrighted work in AI didn't make sense because of the potentially huge amount of money that would be owed to content owners.

"Under any licensing framework that provided for more than negligible payment to individual rights holders," A16z wrote, "AI developers would be liable for tens or hundreds of billions of dollars a year in royalty payments."

Meanwhile, most entities and individuals involved in creating material being used in AI model training, like News Corp., Getty, WME, and even the "Breaking Bad" creator Vince Gilligan, argued in favor of updated copyright rules to offer protection and payment from AI tools.

Currently, there is almost no way to prevent copyrighted content from being crawled from the internet and used to create an LLM; copyright law doesn't address the issue. Authors, visual artists, and even developers are already suing the likes of OpenAI, Microsoft, and Meta because their original work was used without their consent to train the companies' AI tools.

Are you a tech employee or someone else with insight or a tip? Contact Kali Hays at khays@insider.com, on the secure messaging app Signal at 949-280-0267, or through DM on X/Twitter at @hayskali. Reach out using a nonwork device.

Read next

Kali Hays was a Tech Correspondent at Business Insider covering the major social media platforms like Meta, Twitter, and Snap. Her reporting covered major changes and the internal culture at these companies, the founders and executives who run them, and business developments and products. Hays also wrote frequently about AI and emerging trends and shifts in the tech industry overall. Her work has been widely cited, including by the FTC in an investigation into Elon Musk’s takeover of Twitter, and she has appeared as an expert on NBC, CBS, the BBC and elsewhere. Her exclusive reporting and scoops include:Meta's Facebook Messenger hit with layoffs amid ongoing 'efficiency' pushLayoff angst looms over Meta employees as they face tough performance reviews and ongoing reorgsMeta aiming to reveal and demo Orion, its first true AR glasses, at its fall developer conferenceMeta's Responsible AI team shrinks amid layoffs and restructuring, even as the company goes all-in on AIMeta updates RTO policy with stricter mandate, saying workers may lose their jobs if they don't show up 3 days a weekLeaked documents from Mark Zuckerberg and Priscilla Chan's charity include a tacit admission that their biggest bet on education reform was a flop'He is in war time': Mark Zuckerberg's desperate, last-ditch attempt to remake himself — and MetaOpenAI is expected to release a 'materially better' GPT-5 for its chatbot mid-year, sources sayOpenAI's employees were given 2 explanations for why Sam Altman was fired. They're unconvinced and furious.AI is killing the grand bargain at the heart of the web. 'We're in a different world.'Jack Dorsey warns Block employees of coming job cuts: 'The growth of our company has far outpaced the growth of our business.'Elon Musk is considering taking X out of Europe amid EU compliance investigationLeak: Elon Musk said he wants X to be a dating app, too, in an all-hands meeting on the anniversary of his Twitter takeoverLinda Yaccarino, Elon Musk, and the most difficult CEO job on earthElon Musk's Twitter races to build a live video service as it woos right-wing media personalitiesElon Musk is moving forward with a new generative-AI project at Twitter after purchasing thousands of GPUsSnap begins a new round of layoffs with staffers expecting more next weekEvan Spiegel proclaims 'social media is dead' in leaked memo, predicts Snap is about to 'transcend' the smartphoneSnap workers say they're being closely 'tracked' to enforce compliance with the RTO mandateHow Snap misread big threats from TikTok and Apple and lost its chance at becoming an advertising giant