Urgent.News

What's breaking now, across thousands of outlets.

AI

AI crawlers from Meta and Alibaba almost destroyed a volunteer-run LGBT history archive

Jonathan Harborne founded his LGBT history encyclopedia 15 years ago to document some of the struggles his peers had encountered in their lives. He worried that even as changing social values made life better for LGBTQ+ people, the struggles past generations had endured for recognition and respect could be wiped from memory, along with the pain caused by the AIDS pandemic. The site, the LGBT…

AI crawlers from Meta and Alibaba almost destroyed a volunteer-run LGBT history archive

Jonathan Harborne, the founder of the LGBT History Project, a volunteer-run online encyclopedia documenting LGBTQ+ history, recently faced a significant threat to his site. Over the past three weeks, AI crawlers from Meta, Alibaba, and other companies have been overwhelming his server with automated requests, causing the site to become sluggish and costly to maintain.

The project, which has garnered over 50 million views since its launch in 2011 and is archived by the British Library, required Harborne to double the size of his server to handle the excessive traffic. He noticed the issue when he began modernizing the site and migrating it to a new AWS server, hoping for improved performance and security.

However, the opposite happened when Claude Code, a Meta AI tool, analyzed his server logs and detected a massive influx of automated traffic. On a single day, Meta's crawler made approximately 26,000 requests, pulling around 12GB of data from the site, including published articles, editing histories, login pages, and other elements.

Harborne, who funds the project personally, expressed frustration at having to bear the financial burden of these scraping activities, which were performed without respecting his site's robots.txt file, a standard file that guides crawlers on what parts of a site are off-limits. This incident is not unique to Harborne's site; it reflects a broader issue of increasing bot traffic on the internet, with reports indicating that 30% of all web activity is now bot traffic.

Independent website operators are struggling to cope with this surge in automated traffic, which can be costly and sometimes even damaging. Catherine Flick, an ethics professor, points out that running a web server has become increasingly challenging due to the dominance of a few large companies controlling internet infrastructure.

Harborne, however, remains committed to running the LGBT History Project, tightening controls to manage bot access. He emphasizes the broader implications of this issue, warning that AI companies might be damaging the very ecosystem of human-created information that supports their products' functionality. He likens the situation to a virus that threatens to destroy the host it is supposedly benefiting.

Harborne believes that small website owners have little recourse when dealing with such issues, as blocking crawlers falls on their shoulders, and holding these corporations accountable is nearly impossible. The problem, he argues, stems from profit being the primary value in the internet's current landscape, rather than human rights.

AI ethicist Carissa Veliz from the University of Oxford echoes Harborne's concerns, highlighting the precarious state of the internet, which seems on the verge of becoming a Wild West dominated by corporate interests.

Written by urgent.news from Fast Company's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at fastcompany.com →

More in AI

More from Thursday 13 August →