Urgent.News

What's breaking now, across thousands of outlets.

AI

日本のサイト248件中23件、robots.txtは許可でもAIクローラーには403

日本のサイトで、robots.txtはAIクローラーを許可しているのに、サーバーの手前(CDNやWAF)ではそのクローラーに403を返している、という食い違いがどれくらいあるのかを数えました。2026-10-07の朝、Trancoのリストに載っている日本のサイト300件を1件ずつ確認したところ、ブラウザとして普通にページが返ってきた248件のうち23件で、robots.txtでは許可されているAIクローラーのUser-Agentに403が返りました。 この記事は、その数字と内訳、層ごとの例、そしてこの測り方でわからないことです。 先に限界(2026-10-07時点) User-Agentを名乗っただけです。…

Original Japanese Read in English

This report discusses a discrepancy in web crawling policies for Japanese websites. The study analyzed 300 Japanese websites listed on Tranco, and found that 248 of these sites allowed AI crawlers through their robots.txt file, but returned a 403 error when accessed through CDN or WAF layers. The 23 websites where this occurred represent the main focus of the report.

The study utilized Google's GPTBot, OpenAI's ClaudeBot, PerplexityBot, and OAI-SearchBot to test the servers, and found that 14 of the 23 sites returned a 403 error for all four bots. The findings suggest that there may be a difference in how websites handle bot requests at the server level versus their robots.txt files, and that this discrepancy is not always clearly defined.

Brief written by urgent.news from Dev.to's own syndicated text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Changing what my phone agent was shown beat changing the model

This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked On 18 September 2026, two turns after its own lookup had returned a booking's real email, my AI phone receptionist told…

  • Changing notes to AI phone agent improved performance significantly
  • Incorrect bookings dropped from 17 of 216 runs to none
  • Weakest model after changes outperformed strongest model before changes

More from Wednesday 7 October →