We need better than Pay-to-Crawl
The data-hunger of current AIs is reviving an interesting old idea: pay-to-use internet (pay-to-crawl in this case). This development could have very positive ramifications in principle, but current proposals worry me. What is happening? Both the training and operation of popular AI models use vast quantity of data scraped from the internet. This comes at a cost to the provider of said data, who…
The escalating appetite for data among contemporary AI systems is rekindling a previously intriguing concept: pay-to-use internet, colloquially known as "pay-to-crawl." While this notion holds promise in principle, many current proposals leave me uneasy. The training and running of prevalent AI models necessitate substantial data, which is typically sourced from the internet.
This incurs costs for data providers who must cover infrastructure expenses and the electricity required to sustain the data. However, these costs are justified when the data reaches human consumers but not when the data is consumed by AI models. Internet traffic is now predominantly driven by bots, a fact that website owners often seek to curb or restrict.
The challenge lies in distinguishing these bots from genuine human users, as some bots cunningly masquerade as humans by disregarding the robots.txt file, altering their user-agent, and often operating through botnets composed of unsuspecting consumer PCs. In response to this issue, some companies and non-profits are attempting to make these bots pay for data access via a pay-to-crawl system.
While this idea holds merit, the current implementation leaves much to be desired. The notion of paying for access was once integral to the internet's inception, but payment systems have since given way to advertisements, which have proven detrimental. Despite numerous attempts to develop alternative payment systems, none have garnered significant traction.
The absence of a network effect within AI bots suggests that pay-to-crawl could potentially serve as the motivation for a more equitable and productive internet. However, the current proposals fail to instill confidence. They are voluntary, automated licensing frameworks that do not effectively prevent bots from falsely representing themselves as human.
Furthermore, the legal necessity of the provided license remains uncertain, as US courts have consistently ruled that AI training falls under the umbrella of fair use. Consequently, the primary application of pay-to-crawl seems to be in establishing exclusive data libraries for AI use, akin to a Wikipedia accessible only through Gemini.
Although pay-to-crawl cannot differentiate between humans and bots, the pursuit of such a system has spurred a technological arms race. Cloudflare, the largest provider of bot-exclusion services, is also testing a pay-to-crawl system, thereby granting it undue influence over a significant portion of the internet. A more equitable approach would be to demand a micropayment from all users, regardless of whether they are bots or humans, to cover the hosting costs.
This nominal fee would be imperceptible to humans, who consume data at a relatively slow pace, while simultaneously denying bots access to the data, thereby maintaining internet integrity. Creative Commons has expressed concerns over pay-to-crawl, but their perspective is largely idealistic, as ensuring content remains free to humans inevitably renders it accessible to bots posing as humans.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.