# Peridot — https://peridot.noetisphereai.com # # Search engines are welcome. Training crawlers are not. # # The rule below is the machine-readable half of a reservation made in words at # /terms.html and in full in the application's licence: Noetisphere expressly # reserves the rights of reproduction and extraction for text and data mining, # including for the training of artificial intelligence systems, for the # purposes of Article 4(3) of Directive (EU) 2019/790 and its equivalents. # Permission is withheld and can be granted only in writing, in advance: # info@noetisphereai.com # # This file asks. It cannot enforce — a crawler that ignores robots.txt is not # stopped by robots.txt, and the honest place to say so is here rather than in # a summary that implies otherwise. What it does is remove the defence that no # reservation was expressed, which is what makes the legal route usable. User-agent: * Allow: / # The coupon console. Nothing on this site links to it and the page carries # its own noindex, but a Disallow is the half a well-behaved crawler reads # BEFORE fetching — the meta tag only works on a page already retrieved. # # This is not what protects it. Access is decided by Cognito group membership # on every request (amplify/functions/coupons/handler.ts); naming the path here # tells a crawler not to index it and tells anybody reading this file that it # exists, which is a trade worth making because the URL was never the secret. Disallow: /admin.html # ── Model training and dataset collection ────────────────────────────── # OpenAI User-agent: GPTBot Disallow: / User-agent: ChatGPT-User Disallow: / # Anthropic User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: Claude-Web Disallow: / # Common Crawl — the corpus many others are built from User-agent: CCBot Disallow: / # Google and Apple publish these as TDM opt-out tokens rather than as crawlers: # they do not fetch anything, they signal that fetched pages must not be used # for model training. Search indexing is unaffected by them, which is why # Googlebot and Applebot themselves are left alone above. User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / # Others User-agent: PerplexityBot Disallow: / User-agent: Bytespider Disallow: / User-agent: Amazonbot Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: FacebookBot Disallow: / User-agent: cohere-ai Disallow: / User-agent: Diffbot Disallow: / User-agent: omgili Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: YouBot Disallow: / User-agent: AI2Bot Disallow: /