# ------------------------------------------------------------------ # DELIBERATELY ALL-ALLOW. DO NOT ADD Disallow RULES. # This SPA replaced a legacy WordPress site. ~900 legacy URLs # (/events/, /event/, /tag/, /category/, /author/, /wp-content/, # and ?tribe-bar-date / ?eventDisplay=past / ?et_blog / ?s= URLs) # are still in Google's index. They are being de-indexed via HTTP # 404 + a noindex meta tag on the SPA 404 route — NOT via robots.txt. # Blocking them here would re-create GSC's "Indexed, though blocked # by robots.txt" state: a Disallow stops crawling but NOT indexing # from external links, so Googlebot could never crawl the page to # see the noindex. Also NEVER Disallow JS/CSS/assets — Googlebot # must fetch them to render this client-rendered SPA. # ------------------------------------------------------------------ User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / User-agent: Twitterbot Allow: / User-agent: facebookexternalhit Allow: / # AI / LLM crawlers — explicitly allowed for Generative Engine Optimization (GEO) User-agent: GPTBot Allow: / User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: ClaudeBot Allow: / User-agent: Claude-Web Allow: / User-agent: anthropic-ai Allow: / User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / User-agent: Google-Extended Allow: / User-agent: Applebot-Extended Allow: / User-agent: CCBot Allow: / User-agent: Meta-ExternalAgent Allow: / User-agent: Bytespider Allow: / User-agent: cohere-ai Allow: / User-agent: * Allow: / Sitemap: https://rescuedogwines.com/sitemap.xml