Why robots.txt is critical
Your robots.txt file configuration directly impacts your store’s accessibility to AI crawlers. A misconfigured robots.txt can make your product pages harder to retrieve, cite, validate, or recommend.
Why it matters for AI
The robots.txt file is the first thing a crawler visits before exploring your site. If it contains a blocking rule, a respectful crawler stops immediately. The important part is role separation: search bots, user-action fetchers, training bots, and ads-validation bots do not have the same business impact.
AI Crawlers to Know
| Crawler | Operator | Usage |
|---|---|---|
GPTBot | OpenAI | ChatGPT training and general AI knowledge |
OAI-SearchBot | OpenAI | ChatGPT search and citation (no training data collected) |
OAI-AdsBot | OpenAI | ChatGPT ads landing-page validation and relevance |
ChatGPT-User | OpenAI | ChatGPT with real-time browsing |
PerplexityBot | Perplexity | Perplexity Search & Shopping |
Perplexity-User | Perplexity | Fetch triggered by a direct user action |
ClaudeBot | Anthropic | Claude with browsing (traffic doubled Q3 2025 - Q1 2026, SE Ranking, 2026) |
Google-Extended | AI usage/training control, not Google Search crawling | |
Amazonbot | Amazon | Amazon search & Alexa |
Bytespider | ByteDance | TikTok AI features |
Why OAI-SearchBot matters: Unlike GPTBot (used for training), OAI-SearchBot is OpenAI’s retrieval crawler for search and shopping discovery. OAI-AdsBot is different again: it validates ad landing pages and should be treated as paid-media readiness, not organic GEO.
PerplexityBot is not Perplexity-User: PerplexityBot serves Perplexity’s search results and is not used to train foundation models. Perplexity-User is triggered when someone asks Perplexity to visit a specific page, and it generally ignores robots.txt. Allowing one does not allow the other, and blocking PerplexityBot does not stop a user-initiated fetch. Separate the two questions: is the crawler allowed, and is the page reachable when a user asks for it. (Perplexity crawler docs)
What to check
Your robots.txt can contain several types of blocking rules:
- Global block:
Disallow: /underUser-agent: *means everything is blocked for every crawler - Specific blocks:
User-agent: GPTBotplusDisallow: /means this crawler is explicitly targeted and blocked - Partial blocks:
Disallow: /policies/means only your policy pages are blocked (less severe, still damaging)
If a crawler is not named explicitly, the rules under User-agent: * apply to it.
Verity Score analyzes your robots.txt automatically and flags any AI crawler you are blocking.
The Shopify default case
Shopify’s default robots.txt contains:
User-agent: *
Disallow: /admin
Disallow: /cart
Disallow: /orders
Disallow: /checkouts/
Disallow: /*/checkouts
Disallow: /carts
Disallow: /account
These rules are correct: they only block admin and checkout pages. Your product pages, collections and content pages stay accessible.
Watch out: some Shopify apps or custom setups add extra rules that can block AI crawlers. Check regularly.
Most Common Mistakes
Mistake 1: global disallow
User-agent: *
Disallow: /
This blocks everything on your site for every crawler. It is the worst possible configuration.
Mistake 2: blocking AI bots specifically
User-agent: GPTBot
Disallow: /
User-agent: PerplexityBot
Disallow: /
Some guides recommend blocking AI crawlers to “protect your content”. It is counterproductive if you want these systems to recommend you.
Mistake 3: confusing Disallow: / with Disallow: /policies/
Disallow: / blocks the ENTIRE site. Disallow: /policies/ blocks only the /policies/ directory. The difference is a single character, the impact is total.
How to fix it on Shopify
Check your current robots.txt
Go to https://your-store.com/robots.txt and read the contents.
Edit it on Shopify
On Shopify, robots.txt is controlled through the robots.txt.liquid file in your theme:
- Shopify Admin, then Themes, then Actions, then Edit code
- Look for
robots.txt.liquidin the Templates folder - Confirm that no AI crawler is blocked
- If the file does not exist, Shopify serves the default (which is correct)
Recommended Configuration
User-agent: *
Disallow: /admin
Disallow: /cart
Disallow: /orders
Disallow: /checkouts/
Disallow: /account
# OpenAI crawlers (training + search/citation)
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
# Other AI crawlers
User-agent: PerplexityBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: Google-Extended
Allow: /
Sitemap: https://your-store.com/sitemap.xml
Resources
- OpenAI bots documentation
- Perplexity crawlers documentation
- Anthropic crawler documentation
- Shopify robots.txt documentation
- Google crawlers overview
Related articles
- Blocked AI Crawlers: the exact step-by-step audit procedure (robots.txt + CDN)
- llms.txt for Shopify: Useful AI Discovery File, Not a Primary Ranking Signal
- Agent Readiness Score: the 11 /.well-known/ files to publish in 2026
- Understanding Your GEO Score: 9 Factors Explained
- GEO vs SEO: What’s the Difference for E-commerce?
- Schema.org Product: Why and How on Shopify
- Sell on ChatGPT: The Complete Shopify Guide for 2026