{
  "id": 12504232,
  "title": "Of Japan's 248 websites, 23 have robots.txt set to allow but are returning 403 errors for AI crawlers",
  "url": "https://urgent.news/2026/10/07/248-23-robots-txt-ai-403",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-07T00:20:09.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/mahirhir/ri-ben-nosaito248jian-zhong-23jian-robotstxthaxu-ke-demoaikuroraniha403-fdo"
  },
  "original_language": "ja",
  "account": "A study found that 23 out of 248 Japanese websites returned a 403 error to AI crawlers, despite allowing them in their robots.txt files. The crawlers affected were GPTBot, ClaudeBot, PerplexityBot, and OAI-SearchBot. The study analyzed 300 Japanese websites listed by Tranco on October 7, 2026. The discrepancies were found to occur at various layers, including content delivery networks (CDNs) and web application firewalls (WAFs).",
  "summary": "This report discusses a discrepancy in web crawling policies for Japanese websites. The study analyzed 300 Japanese websites listed on Tranco, and found that 248 of these sites allowed AI crawlers through their robots.txt file, but returned a 403 error when accessed through CDN or WAF layers. The 23 websites where this occurred represent the main focus of the report. The study utilized Google's GPTBot, OpenAI's ClaudeBot, PerplexityBot, and OAI-SearchBot to test the servers, and found that 14 of the 23 sites returned a 403 error for all four bots. The findings suggest that there may be a difference in how websites handle bot requests at the server level versus their robots.txt files, and that this discrepancy is not always clearly defined.",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}