The bots already won the front door
The bots are winning. That’s my clearest takeaway after reading the first section of the latest State of the Bots report from TollBit, which builds payment rails between publishers and AI crawlers . In it, People Inc.’s Chief Innovation Officer, Jonathan Roberts , outlines the company’s approach to AI bots scraping content from its many media properties: aggressively block unauthorized bots while…
The bots are winning, according to the State of the Bots report from TollBit. Chief Innovation Officer at People Inc., Jonathan Roberts, outlines their approach to blocking unauthorized AI bots from scraping their content. However, Roberts admits that it's becoming increasingly difficult to identify and block these unauthorized crawlers.
Some of these "bad" bots even try to mimic legitimate ones like Google's. To make things worse, industrial-scale scraping companies have started using networks of devices in homes to camouflage their traffic as coming from real people. Even People Inc., with its substantial resources and expertise in AI search crawlers, hasn't entirely succeeded in blocking all unauthorized bots.
This highlights the challenges faced by media companies of all sizes in combating bot access. The future of publishing in the AI era should focus on content usage rather than just preventing scraping. Instead of allowing all crawlers unrestricted access to their content, publishers should reconsider how their content surfaces for users and implement measures to preserve value, encourage good behavior, and build their business around accurate and up-to-date information retrieval.
The training vs. retrieval debate is crucial. While AI companies have been using crawled information for training their models, this is not the primary concern anymore. The real threat to media business models lies in information retrieval, where AI is used as a discovery surface. This requires accurate and timely information, which is different from training large language models.
Publishers should block training bots but stop expecting payment for training data, as few can afford the scale needed for valuable corpora. Licensing deals have shifted from training to retrieval, as observed by Media and the Machine Substack's Kelly. Monetizing AI answers remains a challenge, with most users not clicking through to the original sources.
While there is some value in appearing as the authoritative source when answering a query, it is not monetizable in itself. The future lies in policing answers rather than just preventing scraping. By ensuring legitimate access to content through licensing or other models, publishers can better control how their content is used in AI-generated answers.
Written by urgent.news from Fast Company's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.