# robots.txt - North American Coating Solutions # https://nacoatingsolutions.com # Generated 2026-07-31 by bin/build-optimization-files.php. Do not hand-edit. # # Policy: this site WELCOMES AI, answer-engine and search crawlers. There is no # blanket Disallow. The only excluded paths are POST handlers, build tooling and # template fragments, none of which are pages. # # NOTE: robots.txt is a crawl signal, not an access control. # ---------------------------------------------------------------- default -- # CORRECTION 2026-09-22: the comment that used to sit here said "a first-match parser # reads rules in file order". That is not how robots.txt works. RFC 9309 and every # major crawler use MOST-SPECIFIC-PATH match, not file order, so the ordering of # Allow and Disallow inside a group does not decide anything. The order below is # kept because it reads well, not because it is load-bearing. User-agent: * Disallow: /api/ Disallow: /admin/ Disallow: /includes/ Disallow: /bin/ Disallow: /research/ Disallow: /data/ Disallow: /404.php Disallow: /error_log Disallow: /*.bak Disallow: /*.orig Disallow: /*.old Disallow: /*.save Disallow: /*.swp Disallow: /*.tmp Disallow: /*.sql Disallow: /*.log Allow: / # Crawl-delay: 1 <-- DISABLED 2026-09-22, not deleted. # It sat inside 'User-agent: *', the group every crawler that is not named # elsewhere falls through to, so it was throttling exactly the agents we most want # reading a 288-URL site. Google ignores Crawl-delay entirely; Bing and Yandex obey # it, and at 1 second a full crawl of this site takes five minutes it does not need # to take. optimization.md PART 9.7: no Crawl-delay in the default group. # To restore it, uncomment the line above. Nothing else has to change. # ------------------------------------------- AI, answer engine and search -- # ONE stacked group. A crawler obeys the single most specific group that names it # and ignores 'User-agent: *' completely, so the same rules are repeated here. # Without that repetition, naming an agent would silently WIDEN its access. # # There is NO BLANK LINE between the User-agent lines below. A blank line ends a # group for the classic parsers, which would turn each block into an empty group # carrying no rules at all. Comment lines are ignored and are safe. # OpenAI User-agent: GPTBot User-agent: OAI-SearchBot User-agent: ChatGPT-User # Anthropic User-agent: ClaudeBot User-agent: anthropic-ai User-agent: Claude-Web User-agent: Claude-User User-agent: Claude-SearchBot # User-agent: ClaudeWebCrawler <-- not a real token. Anthropic uses ClaudeBot, # Claude-User, Claude-SearchBot and anthropic-ai, all named above. Disabled 2026-09-22. # Google - Google-Extended is the Gemini training control, separate from Googlebot User-agent: Googlebot User-agent: Googlebot-Image User-agent: Googlebot-News User-agent: Google-Extended User-agent: GoogleOther User-agent: Storebot-Google # Microsoft / Bing / Copilot User-agent: Bingbot User-agent: Bingbot-Image User-agent: msnbot # User-agent: Copilot <-- not a documented crawler token. Copilot reads the Bing # index through bingbot, which is named above. Disabled 2026-09-22. # User-agent: Copilot-User <-- same: not a documented token. Disabled 2026-09-22. User-agent: BingPreview # Perplexity User-agent: PerplexityBot User-agent: Perplexity-User # Apple - Applebot-Extended is the Apple Intelligence training control User-agent: Applebot User-agent: Applebot-Extended # Meta AI User-agent: facebookexternalhit User-agent: FacebookBot User-agent: meta-externalagent User-agent: Meta-ExternalAgent User-agent: Meta-ExternalFetcher # User-agent: MetaInternalBot <-- does not exist. Meta crawls as meta-externalagent, # Meta-ExternalAgent, Meta-ExternalFetcher and facebookexternalhit, all named above. # Disabled 2026-09-22. # Common Crawl - a training corpus many models are built on User-agent: CCBot # Other answer engines and assistants User-agent: Amazonbot User-agent: Bytespider User-agent: YouBot User-agent: PhindBot User-agent: cohere-ai User-agent: cohere-training-data-crawler User-agent: MistralAI-User User-agent: DuckDuckBot User-agent: DuckAssistBot User-agent: Slurp User-agent: Diffbot User-agent: Timpibot User-agent: Webzio-Extended # Link preview and social unfurl User-agent: Twitterbot User-agent: LinkedInBot User-agent: WhatsApp User-agent: Pinterestbot User-agent: Slackbot User-agent: Slackbot-LinkExpanding User-agent: Discordbot User-agent: TelegramBot User-agent: redditbot Disallow: /api/ Disallow: /admin/ Disallow: /includes/ Disallow: /bin/ Disallow: /research/ Disallow: /data/ Disallow: /404.php Disallow: /error_log Disallow: /*.bak Disallow: /*.orig Disallow: /*.old Disallow: /*.save Disallow: /*.swp Disallow: /*.tmp Disallow: /*.sql Disallow: /*.log Allow: / # --------------------------------------------------- SEO tooling crawlers -- # Blocked for crawl budget only. No opinion is expressed about these products. User-agent: AhrefsBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / User-agent: BLEXBot Disallow: / User-agent: DataForSeoBot Disallow: / # DISABLED nacoatingsolutions.com by appwt-ai-crawler-allowlist.sh (Tony: allow all AI/voice/search): User-agent: PetalBot # Disallow: / <-- ALSO DISABLED 2026-09-22. The allowlist script commented out the # User-agent line above and left this rule behind with no group to belong to. Under # RFC 9309 a rule outside a group is ignored, but a lenient parser could attach it to # whichever group it read last, which is not a thing to leave to a parser's mood. User-agent: SeekportBot Disallow: / User-agent: serpstatbot Disallow: / User-agent: ZoominfoBot Disallow: / User-agent: MegaIndex Disallow: / # ------------------------------------------------------------- discovery -- # SpaceXAI - added 2026-09-04 by owner directive (optimization.md PART 110.0) User-agent: SpaceXAI Allow: / Sitemap: https://nacoatingsolutions.com/sitemap.xml Sitemap: https://nacoatingsolutions.com/sitemap-products.xml Sitemap: https://nacoatingsolutions.com/sitemap-images.xml Sitemap: https://nacoatingsolutions.com/sitemap-video.xml # Four files rather than one index, added 2026-09-22. The pages, the 250 product detail # URLs, the images and the hero video are listed separately so each set can be watched # and rolled back on its own in Search Console. Every URL in all four returned 200 with # a matching canonical at generation time; redirects were not followed. # AI-readable summaries of this business (llmstxt.org convention): # https://nacoatingsolutions.com/llms.txt # https://nacoatingsolutions.com/llms-full.txt # https://nacoatingsolutions.com/ai.txt # https://nacoatingsolutions.com/.well-known/security.txt # ---- AI, VOICE AND SEARCH CRAWLERS — ALLOWED (appwt-ai-crawler-allowlist.sh) ---- # Tony 2026-09-21: allow every AI search, voice search and search engine across # the USA, Canada, Australia, the UK, India and Pakistan. Timpibot named explicitly. # An answer engine that cannot read the page cannot cite it. ALLOW, never Disallow. # Regenerate, never hand-edit: appwt-ai-crawler-allowlist.sh block # OpenAI training crawler, feeds ChatGPT User-agent: GPTBot Allow: / # OpenAI search index, feeds ChatGPT Search User-agent: OAI-SearchBot Allow: / # ChatGPT browsing on a users behalf User-agent: ChatGPT-User Allow: / # OpenAI Operator agent User-agent: Operator Allow: / # Anthropic crawler, feeds Claude User-agent: ClaudeBot Allow: / # Anthropic web access User-agent: Claude-Web Allow: / # Claude browsing on a users behalf User-agent: Claude-User Allow: / # Anthropic search index User-agent: Claude-SearchBot Allow: / # Anthropic legacy agent string User-agent: anthropic-ai Allow: / # Perplexity index User-agent: PerplexityBot Allow: / # Perplexity browsing on a users behalf User-agent: Perplexity-User Allow: / # Gemini and Vertex grounding User-agent: Google-Extended Allow: / # Google non-search product crawls User-agent: GoogleOther Allow: / # Vertex AI grounding User-agent: Google-CloudVertexBot Allow: / # Firebase product crawl User-agent: Google-Firebase Allow: / # Search Console live test User-agent: Google-InspectionTool Allow: / # Bing index, also feeds ChatGPT and Copilot User-agent: Bingbot Allow: / # Bing snapshot renderer User-agent: BingPreview Allow: / # Microsoft legacy agent string User-agent: msnbot Allow: / # Google Search index User-agent: Googlebot Allow: / # Google Images User-agent: Googlebot-Image Allow: / # Google News User-agent: Googlebot-News Allow: / # Google Video User-agent: Googlebot-Video Allow: / # Google Shopping User-agent: Storebot-Google Allow: / # Siri and Spotlight User-agent: Applebot Allow: / # Apple Intelligence grounding User-agent: Applebot-Extended Allow: / # Alexa answers User-agent: Amazonbot Allow: / # DuckDuckGo User-agent: DuckDuckBot Allow: / # DuckDuckGo AI answers User-agent: DuckAssistBot Allow: / # Timpi decentralised index — NAMED BY TONY 2026-09-21 User-agent: Timpibot Allow: / # Common Crawl, the base corpus for many models User-agent: CCBot Allow: / # Cohere retrieval User-agent: cohere-ai Allow: / # Cohere training corpus User-agent: cohere-training-data-crawler Allow: / # Mistral Le Chat browsing User-agent: MistralAI-User Allow: / # Meta AI training User-agent: Meta-ExternalAgent Allow: / # Meta AI on-demand fetch User-agent: Meta-ExternalFetcher Allow: / # Meta language corpus User-agent: FacebookBot Allow: / # Facebook and WhatsApp link previews User-agent: facebookexternalhit Allow: / # X link previews User-agent: Twitterbot Allow: / # LinkedIn link previews User-agent: LinkedInBot Allow: / # Slack link previews User-agent: Slackbot-LinkExpanding Allow: / # Telegram link previews User-agent: TelegramBot Allow: / # ByteDance and TikTok User-agent: Bytespider Allow: / # TikTok search User-agent: TikTokSpider Allow: / # Huawei Petal Search and Celia voice, material in Pakistan and India User-agent: PetalBot Allow: / # Yandex, used across Asia User-agent: YandexBot Allow: / # Naver User-agent: Yeti Allow: / # Baidu User-agent: Baiduspider Allow: / # Sogou User-agent: Sogou web spider Allow: / # Seznam User-agent: SeznamBot Allow: / # Qwant User-agent: Qwantbot Allow: / # Mojeek, an independent UK index User-agent: MojeekBot Allow: / # Brave independent index User-agent: Brave-Search Allow: / # legacy, harmless to allow User-agent: Neevabot Allow: / # You.com answers User-agent: YouBot Allow: / # Allen Institute User-agent: AI2Bot Allow: / # Allen Institute Dolma corpus User-agent: Ai2Bot-Dolma Allow: / # Diffbot knowledge graph User-agent: Diffbot Allow: / # Imagesift User-agent: ImagesiftBot Allow: / # Webz.io corpus User-agent: omgili Allow: / # Webz.io corpus User-agent: omgilibot Allow: / # Webz.io extended User-agent: Webzio-Extended Allow: / # Kangaroo LLM User-agent: Kangaroo Bot Allow: / # Huawei PanGu User-agent: PanguBot Allow: / # Firecrawl agent retrieval User-agent: Firecrawl Allow: / # Semrush AI corpus User-agent: SemrushBot-OCOB Allow: / # generic agent retrieval User-agent: Scrapy Allow: /