984 Requests Said They Were Perplexity. None Could Prove It.
This morning I told my own website that I was ClaudeBot. It took one line of curl and a user agent string copied out of Anthropic's own documentation. Three requests to one article page. Then I opened my dashboard: Anthropic, ClaudeBot, training — 1,698 requests had become 1,701, last seen at 09:22. None of them was ClaudeBot. They came from a laptop in Japan, over ordinary home broadband, from…
984 requests arrived at a website claiming to be various crawlers, but none of them could be verified as legitimately sent by those entities. A user agent string is not concrete proof of identity, as anyone can create one. Verification is only possible when a vendor publishes the IP addresses their crawlers use. Some vendors provide such lists, while others do not.
In this case, Meta, ByteDance, and Amazon do not publish any IP ranges or reverse DNS schemes, making verification impossible for their crawlers. For Anthropic and Common Crawl, the verification lists were outdated. Anthropic began publishing a range list on August 18 but I was still using a snapshot from July 20. Common Crawl published a list on August 11, but it did not get updated in the thirty-day analysis period.
The Perplexity requests were verified against a list that was last updated in October 2025, showing that the list was incomplete and not current. This highlights the importance of having fresh and accurate verification lists to accurately identify the crawlers sending requests.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.