MikeTrendsTrends right now

Mmastodon TechnologyAI first seen 14 h ago, last 14 h ago, peak #6

Japanese sites block AI crawlers despite robots.txt permission

Original: Trancoの日本のサイト300件を2026-10-07に確認。普通に読めた248件のうち23件で、robots.txtが許可しているAIクローラーのUser-Agentに403。内訳、形の違い、確かめ方、限界。 # geo # seo #

A check of 300 top Japanese websites from the Tranco list on 7 October 2026 found that 23 of the 248 pages that loaded normally returned HTTP 403 errors to AI crawler user-agents, even though the sites' robots.txt files permitted those bots. The writeup details which sites are affected, how the blocking patterns differ, how to verify it, and the limits of the method.

Why now: It highlights a mismatch between what robots.txt declares and what servers actually serve, a practical concern for SEO and AI data collection.

TrancoAI crawlersrobots.txt

Open on mastodon →

Evidence

API: https://socialmediatrends-api.osmike.com/v1/trends/1300408