AI

OpenAI and Anthropic are ignoring an established rule that prevents bots scraping online content

Sam Altman, wearing a blue jacket and holding a beer bottle, chats with audience members at the Technical University of Munich (TUM) after a panel discussion
Sam Altman, CEO of OpenAI. Sven Hoppe/picture alliance via Getty Images
Read in app

The world's top two AI startups are ignoring requests by media publishers to stop scraping their web content for free model training data, Business Insider has learned.

OpenAI and Anthropic have been found to be either ignoring or circumventing an established web rule, called robots.txt, that prevents automated scraping of websites, according to a person with knowledge of the analytics of TollBit, as well as another person familiar with the matter.

TollBit is a startup that's aiming to broker paid licensing deals between publishers and AI companies. It found several AI companies are acting in this way and informed certain large publishers in a Friday letter, which was reported earlier by Reuters. The letter did not include the names of any of the AI companies accused of skirting the rule.

OpenAI and Anthropic have stated publicly that they respect robots.txt and blocks to their specific web crawlers, GPTBot and ClaudeBot.

However, according to TollBit's findings, such blocks are not being respected, as claimed. AI companies, including OpenAI and Anthropic, are simply choosing to "bypass" robots.txt in order to retrieve or scrape all of the content from a given website or page.

A spokeswoman for OpenAI declined to comment beyond pointing BI to a corporate blogpost from May, in which the company says it takes web crawler permissions "into account each time we train a new model." A spokesperson for Anthropic did not respond to emails seeking comment.

Robots.txt is a single bit of code that's been used since the late 1990s as a way for websites to tell bot crawlers they don't want their data scraped and collected. It was widely accepted as one of the unofficial rules supporting the web.

With the rise of generative AI, startups and tech companies are racing to build the most powerful AI models. A key ingredient is high-quality data. The thirst for such training data has undermined robots.txt and the unofficial agreements supporting the use of this code.

OpenAI is behind the popular chatbot ChatGPT. The company's largest investor is Microsoft. Anthropic is behind another relatively popular chatbot, Claude. It's largest investor is Amazon.

Both chatbots serve up answers to user questions in the tone of a human. Such answers are only possible because the AI models they are built on include massive amounts of written text and data scraped from the web, much of it under copyright or otherwise owned by creators.

Several tech companies last year argued to the US Copyright Office that nothing on the web should be considered under copyright when it comes to AI training data.

OpenAI has struck a few deals with publishers for access to content, including Axel Springer, which owns BI. The US Copyright Office is set to update its guidance on AI and copyright later this year.

Are you a tech employee or someone else with a tip or insight to share? Contact Kali Hays at khays@jkmperu.com or on secure messaging appSignal at +1-949-280-0267. Reach out using a non-work device.

Read next

Kali Hays was a Tech Correspondent at Business Insider covering the major social media platforms like Meta, Twitter, and Snap. Her reporting covered major changes and the internal culture at these companies, the founders and executives who run them, and business developments and products. Hays also wrote frequently about AI and emerging trends and shifts in the tech industry overall. Her work has been widely cited, including by the FTC in an investigation into Elon Musk’s takeover of Twitter, and she has appeared as an expert on NBC, CBS, the BBC and elsewhere. Her exclusive reporting and scoops include:Meta's Facebook Messenger hit with layoffs amid ongoing 'efficiency' pushLayoff angst looms over Meta employees as they face tough performance reviews and ongoing reorgsMeta aiming to reveal and demo Orion, its first true AR glasses, at its fall developer conferenceMeta's Responsible AI team shrinks amid layoffs and restructuring, even as the company goes all-in on AIMeta updates RTO policy with stricter mandate, saying workers may lose their jobs if they don't show up 3 days a weekLeaked documents from Mark Zuckerberg and Priscilla Chan's charity include a tacit admission that their biggest bet on education reform was a flop'He is in war time': Mark Zuckerberg's desperate, last-ditch attempt to remake himself — and MetaOpenAI is expected to release a 'materially better' GPT-5 for its chatbot mid-year, sources sayOpenAI's employees were given 2 explanations for why Sam Altman was fired. They're unconvinced and furious.AI is killing the grand bargain at the heart of the web. 'We're in a different world.'Jack Dorsey warns Block employees of coming job cuts: 'The growth of our company has far outpaced the growth of our business.'Elon Musk is considering taking X out of Europe amid EU compliance investigationLeak: Elon Musk said he wants X to be a dating app, too, in an all-hands meeting on the anniversary of his Twitter takeoverLinda Yaccarino, Elon Musk, and the most difficult CEO job on earthElon Musk's Twitter races to build a live video service as it woos right-wing media personalitiesElon Musk is moving forward with a new generative-AI project at Twitter after purchasing thousands of GPUsSnap begins a new round of layoffs with staffers expecting more next weekEvan Spiegel proclaims 'social media is dead' in leaked memo, predicts Snap is about to 'transcend' the smartphoneSnap workers say they're being closely 'tracked' to enforce compliance with the RTO mandateHow Snap misread big threats from TikTok and Apple and lost its chance at becoming an advertising giant