Tag archive

#robots.txt

8 published storiesCanonical archive

Stories tagged Robots.txt

Topic archive
Lina VossJuly 25, 202611 min

AI Crawler Response Codes Guide for Technical QA

A precise, source-backed guide to reading 200, 301/302, 403, 404, 429, and 5xx responses through the lens of GPTBot, ClaudeBot, and PerplexityBot — plus the QA process to catch AI-crawler access failures before they cost you visibility.

AI Crawler Response Codes Guide for Technical QA
Ada O'BrienJuly 25, 202611 min

AI Crawler Blocks: Diagnose Why Bots Can’t Reach Your Site

A precise, source-backed framework for telling apart robots.txt disallows, WAF/firewall rules, rate limiting, and CDN bot-management blocks — with the exact crawler identities, IP verification methods, and diagnostic steps to confirm which one is stopping GPTBot, ClaudeBot, or PerplexityBot from reaching your pages.

AI Crawler Blocks: Diagnose Why Bots Can’t Reach Your Site
Ada O'BrienJuly 25, 202611 min

Robots.txt Rules for AI Crawlers: A Governance Guide

A working robots.txt for AI crawlers isn't a file you write once — it's a policy you maintain as OpenAI, Anthropic, Perplexity, Google, and Apple keep splitting one bot into three with different rules for training, search, and live answers.

Robots.txt Rules for AI Crawlers: A Governance Guide