# As a condition of accessing this website, you agree to abide by the following # content signals: # (a) If a Content-Signal = yes, you may collect content for the corresponding # use. # (b) If a Content-Signal = no, you may not collect content for the # corresponding use. # (c) If the website operator does not include a Content-Signal for a # corresponding use, the website operator neither grants nor restricts # permission via Content-Signal with respect to the corresponding use. # The content signals and their meanings are: # search: building a search index and providing search results (e.g., returning # hyperlinks and short excerpts from your website's contents). Search does not # include providing AI-generated search summaries. # ai-input: inputting content into one or more AI models (e.g., retrieval # augmented generation, grounding, or other real-time taking of content for # generative AI search answers). # ai-train: training or fine-tuning AI models. # use: how AI systems may consume the content (immediate, reference, or full). # ANY RESTRICTIONS EXPRESSED VIA CONTENT SIGNALS ARE EXPRESS RESERVATIONS OF # RIGHTS UNDER ARTICLE 4 OF THE EUROPEAN UNION DIRECTIVE 2019/790 ON COPYRIGHT # AND RELATED RIGHTS IN THE DIGITAL SINGLE MARKET. # BEGIN Cloudflare Managed content User-agent: * Content-Signal: search=yes,ai-train=no,use=reference Allow: / User-agent: Amazonbot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Bytespider Disallow: / User-agent: CCBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: CloudflareBrowserRenderingCrawler Disallow: / User-agent: Google-Extended Disallow: / User-agent: GPTBot Disallow: / User-agent: meta-externalagent Disallow: / # END Cloudflare Managed Content # Lumify Robots.txt # Lumify — the agent-ready Sports Intelligence API (https://lumify.ai) User-agent: * Allow: / # Sitemap location — the only directive above that crawlers actually parse. Sitemap: https://lumify.ai/sitemap.xml # LLM-readable documentation (informational only — "LLMs:" is not a # standard robots.txt directive; crawlers ignore it, so these are kept # as comments rather than pseudo-directives). # LLMs: https://lumify.ai/llms.txt # LLMs-full: https://lumify.ai/llms-full.txt # AI agent discovery (also listed in sitemap.xml) # Manifest: https://lumify.ai/.well-known/agent.json # OpenAPI: https://lumify.ai/openapi.json # MCP: https://lumify.ai/mcp # AI setup: https://lumify.ai/docs/ai # Block legacy WordPress URLs (site migrated from WordPress) Disallow: /knowledgebase-category/ Disallow: /knowledgebase-tag/ Disallow: /tag/ Disallow: /category/ Disallow: /feed/ Disallow: /*/feed/ Disallow: /wp-content/ Disallow: /wp-includes/ Disallow: /wp-admin/ # Block Cloudflare internal paths Disallow: /cdn-cgi/ # Block URLs with query parameters (prevents duplicate content from tracking # params). NOTE: this is intentionally broad — any future public, indexable # route that relies on a query string (rather than a path segment) will be # blocked too. Re-check this rule before shipping such a route. Disallow: /*?* # Block authenticated app pages (require login) Disallow: /app/ Disallow: /admin/ # Block login/logout pages Disallow: /login Disallow: /logout # Block utility/health check endpoints Disallow: /health Disallow: /health-check # Block API endpoints from indexing Disallow: /api/ # Block auth flow confirmation pages Disallow: /verify-email Disallow: /verify-beta Disallow: /beta-confirmation Disallow: /reset-password Disallow: /forgot-password Disallow: /register # --------------------------------------------------------------------------- # AI crawlers — explicitly welcomed. Lumify is built to be read by agents. # Named blocks mirror the same disallow rules as User-agent: * above (a named # User-agent block replaces the wildcard block for that bot rather than # combining with it, so the rules must be repeated here). # --------------------------------------------------------------------------- User-agent: GPTBot User-agent: ChatGPT-User User-agent: OAI-SearchBot User-agent: ClaudeBot User-agent: anthropic-ai User-agent: PerplexityBot User-agent: CCBot Allow: / Disallow: /knowledgebase-category/ Disallow: /knowledgebase-tag/ Disallow: /tag/ Disallow: /category/ Disallow: /feed/ Disallow: /*/feed/ Disallow: /wp-content/ Disallow: /wp-includes/ Disallow: /wp-admin/ Disallow: /cdn-cgi/ Disallow: /*?* Disallow: /app/ Disallow: /admin/ Disallow: /login Disallow: /logout Disallow: /health Disallow: /health-check Disallow: /api/ Disallow: /verify-email Disallow: /verify-beta Disallow: /beta-confirmation Disallow: /reset-password Disallow: /forgot-password Disallow: /register # Google-Extended doesn't crawl independently — it's a Google-specific opt-in # token controlling use of Googlebot-crawled content for AI training (Gemini). User-agent: Google-Extended Allow: /