Of Japan's 248 websites, 23 have robots.txt set to allow but are returning 403 errors for AI crawlers
We counted how many Japanese sites have a discrepancy where robots.txt allows AI crawlers but the server (CDN or WAF) returns a 403 error. On the morning of October 7, 2026, we checked 300 Japanese sites listed on Tranco's list one by one. Out of 248 sites that returned pages normally as a browser, 23 sites returned a 403 error for AI crawlers' User-Agent that were allowed in robots.txt. This article provides the numbers, breakdown, examples by layer, and what this measurement method does not reveal. First, the limitations (as of October 7, 2026): We only claimed the User-Agent.
A study found that 23 out of 248 Japanese websites returned a 403 error to AI crawlers, despite allowing them in their robots.txt files. The crawlers affected were GPTBot, ClaudeBot, PerplexityBot, and OAI-SearchBot. The study analyzed 300 Japanese websites listed by Tranco on October 7, 2026. The discrepancies were found to occur at various layers, including content delivery networks (CDNs) and web application firewalls (WAFs).
Written by urgent.news from Dev.to's report — not a translation of it. Machine-written — may contain errors; check the original before relying on it.