# https://shopit.se/robots.txt # Last updated: 2026-08-12 # # POLICY # Shopit content may be indexed, retrieved and cited to answer users' # questions in real time. It may not be used to train or fine-tune AI models. # # Vendors that ship separate user-agents for training vs. retrieval have the # training agent denied below; their retrieval agents are permitted under the # wildcard group. Do NOT add retrieval agents to a named group — a named # group makes that agent ignore "User-agent: *" entirely. # # CONTENT SIGNALS (https://contentsignals.org) # search : building a search index and returning links + short excerpts. # Does not include AI-generated summaries. # ai-input : inputting content into AI models — RAG, grounding, generative # search answers. This is the signal that gets you cited. # ai-train : training or fine-tuning AI models. # # ANY RESTRICTIONS EXPRESSED VIA CONTENT SIGNALS ARE EXPRESS RESERVATIONS OF # RIGHTS UNDER ARTICLE 4 OF EUROPEAN UNION DIRECTIVE 2019/790 ON COPYRIGHT # AND RELATED RIGHTS IN THE DIGITAL SINGLE MARKET. # =========================================================================== User-agent: * Content-Signal: search=yes,ai-input=yes,ai-train=no Disallow: /api/ Disallow: /go-to-store Disallow: /free/ Disallow: /service/ Disallow: /landing-ehandel-se Disallow: /user/ Disallow: /recover/ Disallow: /register/ Disallow: /search$ Disallow: /search? Disallow: /search/ Disallow: /compare/ Disallow: /*?f= Disallow: /*&f= Disallow: /*?color= Disallow: /*&color= Disallow: /*?text= Disallow: /*&text= Disallow: /*?code= Disallow: /*&code= Disallow: /*?expand= Disallow: /*&expand= Disallow: /*?expandedBrands= Disallow: /*&expandedBrands= Disallow: /*?secondary_expand= Disallow: /*&secondary_expand= Disallow: /*?show_secondary= Disallow: /*&show_secondary= Disallow: /*?fold= Disallow: /*&fold= Disallow: /*?foldedBrands= Disallow: /*&foldedBrands= Disallow: /*?group= Disallow: /*&group= Disallow: /*?sort= Disallow: /*&sort= Disallow: /*?of= Disallow: /*&of= Disallow: /*?unpf= Disallow: /*&unpf= Disallow: /*?size= Disallow: /*&size= Disallow: /*?c= Disallow: /*&c= Disallow: /*?category_id= Disallow: /*&category_id= Disallow: /*?lang= Disallow: /*&lang= Disallow: /*?cashback= Disallow: /*&cashback= # =========================================================================== # TRAINING CRAWLERS — DENIED # # Each agent below collects content into a persistent corpus used to train or # fine-tune models, or redistributes it in bulk to others who do. None is the # path by which a user receives an answer citing shopit.se. # =========================================================================== # OpenAI — training crawler. # Permitted above: OAI-SearchBot (ChatGPT search index), ChatGPT-User (browse). User-agent: GPTBot Disallow: / # Anthropic — training crawler. # Permitted above: Claude-SearchBot (search index), Claude-User (browse). User-agent: ClaudeBot Disallow: / # Google — Gemini Apps and Vertex AI model improvement. # Google states this token does NOT affect Google Search. Organic ranking and # AI Overviews eligibility are governed by Googlebot, permitted above. User-agent: Google-Extended Disallow: / # Apple — Apple Intelligence training opt-out token. # Permitted above: Applebot (Siri, Spotlight, Apple search). User-agent: Applebot-Extended Disallow: / # Meta — training and AI product corpus. # Permitted above: meta-externalfetcher (user-initiated fetch). User-agent: meta-externalagent Disallow: / # Amazon — a single user-agent covers crawling, Alexa and Rufus, with no # separation between retrieval and model training. User-agent: Amazonbot Disallow: / # Common Crawl — bulk redistribution into third-party training corpora, # with no attribution path back to the source. User-agent: CCBot Disallow: / # ByteDance — training crawler, aggressive crawl rate, no citation path. User-agent: Bytespider Disallow: / # Allen Institute for AI — builds the Dolma training corpus. User-agent: AI2Bot Disallow: / User-agent: AI2Bot-Dolma Disallow: / # webz.io — commercial resale of crawled content into training datasets. User-agent: Omgilibot Disallow: / User-agent: Omgili Disallow: / # Image training corpus. User-agent: ImagesiftBot Disallow: / # =========================================================================== Sitemap: https://shopit.se/sitemap.xml