# 3d3d.ca # # THE LINE THIS FILE DRAWS # # Owner direction, 2026-08-17: "not so scrapable, but really full and complete." # # So the rule is about PURPOSE, not about who owns the crawler: # # ALLOWED · fetching a page to answer somebody's question and cite this site # as the source. That is a referral, and it is what this business # wants. Search indexes and user-triggered assistant fetches both # sit here. # BLOCKED · fetching in bulk to build a training corpus. That takes the work # and returns nothing, and the site gets no attribution from it. # # Several vendors run BOTH, under different tokens, and getting them the wrong # way round is the common mistake: blocking OAI-SearchBot removes you from # ChatGPT's cited results while doing nothing about training, and blocking # Google-Extended does NOT affect Google Search or AI Overviews, which are # governed by Googlebot. # # HONEST LIMIT: robots.txt is a request, not a fence. It is obeyed by the # well-behaved and ignored by everyone else, and nothing in this file is # enforced. Enforcement has to happen at the edge. See # docs/site-pass/04-BOT-POLICY.md. # ── Search indexes. Allowed: they send people here. ────────────────────────── User-agent: Googlebot User-agent: Bingbot User-agent: Applebot User-agent: DuckDuckBot User-agent: PetalBot Disallow: /api/ Allow: / # ── Answer engines that cite. Allowed for the same reason. ─────────────────── # OAI-SearchBot indexes for ChatGPT's search results; ChatGPT-User fetches a # page because a person asked about it. Neither trains. Same split for # Perplexity, Anthropic, DuckDuckGo, You.com and Meta. User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Claude-User User-agent: Claude-SearchBot User-agent: DuckAssistBot User-agent: YouBot User-agent: meta-externalfetcher Disallow: /api/ Allow: / # ── Bulk training and corpus collection. Not allowed. ──────────────────────── # GPTBot trains, OAI-SearchBot above does not. ClaudeBot trains, Claude-User and # Claude-SearchBot above do not. Google-Extended governs Gemini training only # and has no effect on Search. Applebot-Extended is the training opt-out that # leaves Applebot's Siri and Spotlight indexing untouched. User-agent: GPTBot User-agent: ClaudeBot User-agent: anthropic-ai User-agent: Google-Extended User-agent: Applebot-Extended User-agent: CCBot User-agent: Bytespider User-agent: meta-externalagent User-agent: FacebookBot User-agent: cohere-ai User-agent: cohere-training-data-crawler User-agent: Diffbot User-agent: omgili User-agent: omgilibot User-agent: Webzio-Extended User-agent: ImagesiftBot User-agent: Timpibot User-agent: AI2Bot User-agent: Kangaroo Bot Disallow: / # ── Amazonbot: judgement call, and it is documented here rather than hidden. ─ # Amazon documents it as serving Alexa answers to user requests AND improving # their services. It is genuinely dual-purpose under one token, so there is no # way to take the first without the second. Blocked, because the citing half is # small and the corpus half is not. Reverse this line if Alexa referrals ever # matter. User-agent: Amazonbot Disallow: / # ── Everything else. ───────────────────────────────────────────────────────── # Endpoints are excluded because crawling them is pure waste: they are not # pages, they are not in the sitemap, and nothing links to them. User-agent: * Disallow: /api/ Allow: / Sitemap: https://3d3d.ca/sitemap.xml # Feed discovery. The purpose-based crawler policy above remains unchanged. # RSS feed: https://3d3d.ca/rss.xml (also declared by the HTML alternate link)