# House of McPolin — Aaron McPolin # Full terms: https://mcpolinhouse.com/terms-and-conditions/ # # Search engines are welcome. Being findable is the point. # Ingestion for machine learning training is not licensed. The distinction # below is deliberate: crawlers that only index are allowed, crawlers whose # job is to collect training data are not. Where a company runs both, the # search crawler is allowed and the training crawler is named and blocked. # ---- search: allowed, and wanted -------------------------------------- User-agent: Googlebot Allow: / User-agent: Googlebot-Image Allow: / User-agent: Bingbot Allow: / User-agent: DuckDuckBot Allow: / User-agent: Slurp Allow: / User-agent: Applebot Allow: / User-agent: DuckAssistBot Allow: / # ---- AI search: allowed ------------------------------------------------ # These index pages so the work can be found and cited in AI search answers # ("fine art nude photographers", "erotic fine art photography"). Each # company states in its crawler documentation that these agents are not used # to train its models; the training crawlers are separate, and blocked below. User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # ---- AI training: not licensed ---------------------------------------- # Blocking the training crawler while allowing the search crawler is exactly # what Google-Extended and Applebot-Extended exist for. User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Claude-Web Disallow: / User-agent: anthropic-ai Disallow: / User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Meta-ExternalFetcher Disallow: / User-agent: FacebookBot Disallow: / User-agent: Amazonbot Disallow: / User-agent: cohere-ai Disallow: / User-agent: cohere-training-data-crawler Disallow: / User-agent: Diffbot Disallow: / User-agent: omgili Disallow: / User-agent: omgilibot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: img2dataset Disallow: / User-agent: Timpibot Disallow: / User-agent: Scrapy Disallow: / User-agent: AI2Bot Disallow: / User-agent: Kangaroo Bot Disallow: / User-agent: PanguBot Disallow: / User-agent: Webzio-Extended Disallow: / User-agent: YouBot Disallow: / # ---- everyone else ----------------------------------------------------- User-agent: * Allow: / # Facts about the artist, editions and representation, from the studio: # https://mcpolinhouse.com/llms.txt Sitemap: https://mcpolinhouse.com/sitemap-index.xml Sitemap: https://mcpolinhouse.com/sitemap.xml Sitemap: https://mcpolinhouse.com/sitemap-images.xml