# ============================================================================ # robots.txt — History of Market · 美股编年史 # ---------------------------------------------------------------------------- # Rights & usage (short form): # Reuse with attribution is welcome; copying the work is not. Data # (/api/**.json): CC BY 4.0. Charts, share images, embeds and quotes: # free to reuse when credited "History of Market · 美股编年史 # (historyofmarket.com)" with a link. No derivative or look-alike # sites, no use of the brand names, no citation without the source. # Terms: https://historyofmarket.com/license/ # # Full terms: /license/ · Machine copy: /llms.txt § Rights & Usage Notice # # Crawling is welcome. This file explicitly allows every mainstream AI / LLM # crawler so the Chronicle shows up in modern answer engines — ChatGPT, # Claude, Perplexity, Bing Chat, Google AI Overviews, Meta AI, Apple # Intelligence, Amazon Alexa, DuckDuckGo Assist, You.com, Mistral, Cohere, # ByteDance, and so on — as well as traditional search engines. Permission to # crawl, index and summarise is not permission to rebuild the chronicle as # your own: see /license/. # # /api/ is intentionally crawler-allowed. Every dataset is declared as a # schema.org `Dataset` whose `distribution.contentUrl` points to the # corresponding /api/*.json — disallowing the directory would silently # break every Dataset declaration, Google Dataset Search ingestion, and # the explicit AEO contract written into /llms.txt and /api/profile.json # (which advertise the same JSONs as the canonical machine entry points). # Sitemap deliberately omits the raw JSONs so search engines still rank # the editorial pages above them, but no agent is blocked from fetching. # ============================================================================ User-agent: * Allow: / # Cloudflare Content Signals: search + AI input + AI training are all # permitted. Signalling is about crawling, not about the licence above. Content-Signal: search=yes, ai-input=yes, ai-train=yes # ── Traditional search crawlers (the usual allow-all default, made explicit) ── User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / User-agent: Slurp Allow: / User-agent: DuckDuckBot Allow: / User-agent: Baiduspider Allow: / User-agent: YandexBot Allow: / User-agent: Sogou web spider Allow: / User-agent: Applebot Allow: / # ── LLM training & answer-engine crawlers ── # OpenAI User-agent: GPTBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: OAI-SearchBot Allow: / # Anthropic User-agent: ClaudeBot Allow: / User-agent: Claude-Web Allow: / User-agent: anthropic-ai Allow: / User-agent: Claude-User Allow: / User-agent: Claude-SearchBot Allow: / # Google AI (covers Bard / Gemini / AI Overviews) User-agent: Google-Extended Allow: / User-agent: GoogleOther Allow: / # Microsoft / Bing Chat / Copilot User-agent: Bingbot Allow: / User-agent: msnbot Allow: / # Perplexity User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # Apple Intelligence (training opt-in) User-agent: Applebot-Extended Allow: / # Meta / Llama User-agent: FacebookBot Allow: / User-agent: Meta-ExternalAgent Allow: / User-agent: Meta-ExternalFetcher Allow: / # Amazon / Alexa User-agent: Amazonbot Allow: / # DuckDuckGo Assist User-agent: DuckAssistBot Allow: / # You.com User-agent: YouBot Allow: / # Mistral User-agent: MistralAI-User Allow: / # Cohere User-agent: cohere-ai Allow: / User-agent: cohere-training-data-crawler Allow: / # ByteDance (Doubao / Volcengine) User-agent: Bytespider Allow: / # Common Crawl (feeds many open-source LLMs) User-agent: CCBot Allow: / # Diffbot (Factual / AI training) User-agent: Diffbot Allow: / # Timpi (decentralised search) User-agent: Timpibot Allow: / # Webz.io / Webhose (news + forum ingest) User-agent: Webzio-Extended Allow: / User-agent: omgili Allow: / # ── Sitemaps ── Sitemap: https://historyofmarket.com/sitemap.xml