# robots.txt for math.iisc.ac.in # Updated: 2026-05-25 # # Legitimate search engine crawlers are welcome. # AI training/scraping crawlers are blocked to protect bandwidth and content. # --- Blocked: ByteDance / TikTok (caught in infinite redirect loop) --- User-agent: Bytespider Disallow: / # --- Blocked: OpenAI crawlers (bulk PDF downloading) --- User-agent: GPTBot Disallow: / User-agent: ChatGPT-User Disallow: / User-agent: OAI-SearchBot Disallow: / # --- Blocked: Amazon (bulk crawling) --- User-agent: Amazonbot Disallow: / # --- Blocked: Anthropic --- User-agent: ClaudeBot Allow: /~arvind/ Disallow: / User-agent: Claude-User Allow: /~arvind/ Disallow: / User-agent: Claude-SearchBot Allow: /~arvind/ Disallow: / User-agent: Claude-Web Allow: /~arvind/ Disallow: / User-agent: anthropic-ai Allow: /~arvind/ Disallow: / # --- Blocked: Meta --- User-agent: FacebookBot Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Meta-ExternalFetcher Disallow: / # --- Blocked: Common Crawl (primary AI training data source) --- User-agent: CCBot Disallow: / # --- Blocked: Perplexity --- User-agent: PerplexityBot Disallow: / # --- Blocked: Apple AI training --- User-agent: Applebot-Extended Disallow: / # --- Blocked: Miscellaneous AI/scraper bots seen in logs --- User-agent: SleepBot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: cohere-ai Disallow: / # --- Opt out of Google AI training (Bard/Gemini) while keeping search indexing --- User-agent: Google-Extended Disallow: / # --- Allow: Legitimate search engine crawlers --- User-agent: Googlebot Allow: / User-agent: bingbot Allow: / User-agent: Slurp Allow: / User-agent: DuckDuckBot Allow: / User-agent: Baiduspider Allow: / # --- Default: allow all other crawlers, but request a polite crawl rate --- User-agent: * Crawl-delay: 5