User-agent: * Disallow: /images/ Disallow: /photos/ Disallow: /TTLogin.aspx Disallow: /index_new.aspx Disallow: /PrintHelp.aspx Disallow: /*.axd Disallow: /*.ashx Disallow: /wp.aspx Disallow: /bc.aspx Disallow: /backup.aspx Disallow: /admin/ Disallow: /cpanel/ Disallow: /config/ # /showcase/profile.aspx is DELIBERATELY NOT DISALLOWED as of 2026-09-09, and this # is temporary. It is an athlete profile, and athlete profiles now emit # X-Robots-Tag: noindex plus a meta robots tag. A crawler cannot read a noindex on a # page it is forbidden to fetch, so the Disallow that used to sit here was keeping # already-indexed profiles IN the index rather than getting them out - the two # directives work against each other, and only one of them can drop a page. # # Re-block it once Search Console reports zero indexed profile URLs, NOT on a date. # Procedure and what to watch: docs/reference/seo-crawler-controls.md. Disallow: /olc/waivers.aspx Disallow: /team/practices.aspx # PDF/Excel export triggers. These POST now, so a well-behaved crawler has nothing # to follow anyway - these lines are for the ones that guess URLs from old links. # Each hit used to build an Aspose PDF: 30 a day became 5,238 in five days, 9,007 of # 9,015 anonymous, peaking overnight. The app bounces these to the schedule and # builds nothing; this just saves everyone the round trip. Disallow: /*export= Disallow: /*printbr= # --------------------------------------------------------------------------- # XML sitemap. Generated by sitemap.ashx and served at /sitemap.xml through the # Sitemap rules in UrlRewriter.config - NOT at its .ashx URL, which the # "Disallow: /*.ashx" line above would have blocked. /sitemap.xml is an index # pointing at /sitemap-.xml for each season in the published window - the two # ahead, the current one, and the two behind. # # THREE LINES, NOT ONE. This exact file is deployed to all three tournament App # Services (soccer, baseball and lacrosse - see deploy-sportsincollege-staging.yml) # and is served on every customer domain those events sit on as well, so a single # host-specific line would be wrong everywhere except one host. Each sitemap lives # on the host whose events it lists. # # Naming all three here is also what authorizes the customer-domain URLs inside # them: an event with its own domain canonicalizes to that domain (phase 7), so its # sitemap entry does too, and a search engine only trusts a sitemap for URLs on # another host when that host's robots.txt names the sitemap - which this file, # served on that domain, does. # # To be precise about what the handler does, since the line above invites the wrong # reading: the sitemap is NOT filtered by host. Every host serves the identical # complete list, customer domains included. Only these three sport-domain URLs are # advertised, so that is invisible in practice - but do not build per-host behavior # on an assumption this file appears to make and the code does not. # # The hosts are the ones web.SportCanonicalHost builds from the SPORT app setting. # A fourth sport means a fourth line. Sitemap: https://soccer.sincsports.com/sitemap.xml Sitemap: https://baseball.sincsports.com/sitemap.xml Sitemap: https://lacrosse.sincsports.com/sitemap.xml # --------------------------------------------------------------------------- # RECRUITING sitemap. A separate document for the OTHER product this codebase # serves - college recruiting on the {sport}incollege.com domains. Generated by # recruitingsitemap.ashx and served at /sitemap-recruiting.xml through the # "Sitemap recruiting" rule in UrlRewriter.config, again because of the # "Disallow: /*.ashx" line above. Athlete profiles are NOT in it - see below. # # SIX LINES, NOT THREE, and the reason the event block gives applies here twice # over. That block needs three because one robots.txt is deployed to three # tournament App Services and each sitemap lives on its own host. This product # has TWO host families per sport, not one: the recruiting domain # (www.{sport}incollege.com) and the sport domain ({sport}.sincsports.com), # because PageBase carries a redirect that may or may not move the recruiting # pages onto the sport domain depending on what the origin reports for HTTPS. # That is ledger question 23 and it is not settled, so the generated document # resolves the host at RUNTIME by re-evaluating the redirect's own condition - # and this static file, which cannot, has to name both possibilities. Whichever # one the code picks, its host is named here. # # Naming a host here is also what AUTHORIZES the URLs inside a sitemap served # from another host: a search engine only trusts a cross-host sitemap when the # target host's robots.txt names it, and this identical file is served on every # one of these hosts. # # The handler is a no-op on a host with no sport parameter - it answers 404 # rather than an empty urlset - so a line that turns out to name a host this app # does not answer on costs nothing. Sitemap: https://www.soccerincollege.com/sitemap-recruiting.xml Sitemap: https://www.baseballincollege.com/sitemap-recruiting.xml Sitemap: https://www.lacrosseincollege.com/sitemap-recruiting.xml Sitemap: https://soccer.sincsports.com/sitemap-recruiting.xml Sitemap: https://baseball.sincsports.com/sitemap-recruiting.xml Sitemap: https://lacrosse.sincsports.com/sitemap-recruiting.xml # --------------------------------------------------------------------------- # ATHLETE PROFILES ARE NOT DISALLOWED HERE, AND THAT IS DELIBERATE. DO NOT ADD # THEM. # # recruit.aspx, athleteprofile.aspx, /profile/*, default.aspx?ath= and the other # athlete-scoped pages emit a real noindex instead - X-Robots-Tag from # Global.asax and a on the recruiting shells, both decided # by web.AthleteNoIndexUrl. These are minors and recruit.aspx?id= is # sequentially enumerable, so a profile already sitting in an index has to be # DROPPED, not merely left out of the sitemap. # # A Disallow line here would prevent exactly that. It stops a crawler FETCHING # the page, so the crawler never reads the noindex, and an already-indexed URL # stays indexed indefinitely - typically as a bare URL with no snippet, which is # worse rather than better. Allowing the fetch is what makes the removal happen. # Revisit only once Search Console shows the profiles gone. # # (/showcase/profile.aspx was the one exception, disallowed long before this # decision. Jonathan opened it on 2026-09-09 so its noindex can actually be read. # It goes back behind a Disallow once the profiles are out of the index - see the # note at the top of this file and the procedure in # docs/reference/seo-crawler-controls.md.) # --------------------------------------------------------------------------- # Search engines: ALLOWED. Do not add a Disallow for one without reading this. # # Bingbot and Slurp (Yahoo) were disallowed here until 2026-09-09, alongside a # list of AI crawlers, as part of an effort to slow the rate at which bots # crawled the sites. They are not AI crawlers. Blocking them cost us Bing, # Yahoo, DuckDuckGo and Microsoft Copilot outright - Bing's index is what the # last three read from - for no reduction in the abuse this file was aimed at. # # Crawl rate is not controlled from here. Bingbot honors Crawl-delay and Bing # Webmaster Tools has a Crawl Control panel that schedules crawl rate by hour. # Googlebot honors neither: its Search Console crawl-rate control was retired in # early 2024, and it self-throttles on slow responses instead. The enforcement # that actually holds for anything else is Cloudflare - see # docs/reference/seo-crawler-controls.md. # --------------------------------------------------------------------------- # SEO tooling. Allowed at a slow crawl so we can audit our own site and see our # own backlink profile; neither needs to crawl fast, and both honor Crawl-delay. User-agent: AhrefsBot Crawl-delay: 10 User-agent: SemrushBot Crawl-delay: 10 # --------------------------------------------------------------------------- # The two crawlers that were never named here, added 2026-09-16. # # A per-crawler pull of the last 7 days settled a long argument about crawler # load in an unexpected direction. Of 439,000 crawler requests and ~9.9 GB, # THREE crawlers were 94% of the bandwidth - Claude-SearchBot (3.98 GB), Baidu # (3.38 GB) and Applebot (1.92 GB) - and not one of the three appeared anywhere # in this file. Meanwhile every AI crawler that was being argued over came to # 158.9 MB between them, 1.6% of the total. # # The lesson is worth more than the lines: the crawlers costing us something # were the ones nobody had thought to name, and the ones in the file were there # because someone had heard of them. Pull the numbers before writing a rule. # The full table and the three-ways-wrong rate-limit proposal it replaced are in # docs/reference/seo-crawler-controls.md section 2. # --------------------------------------------------------------------------- # Apple. Throttled rather than blocked: Applebot feeds Siri and Spotlight, which # is real distribution even though it is small, and it honors Crawl-delay. User-agent: Applebot Crawl-delay: 10 # Baidu. Blocked outright, which is a decision rather than an oversight, and one # line to reverse if an international case ever appears: 3.38 GB a week to a # Chinese search engine, for a platform serving US youth sports. # # Binary because it HAS to be - Baidu does not honor Crawl-delay in robots.txt. # It takes crawl pressure only from its own webmaster platform, so a delay line # here would look like a throttle and do nothing. The choice was block or allow. User-agent: Baiduspider Disallow: / # --------------------------------------------------------------------------- # Blocked crawlers. robots.txt is advisory - polite crawlers obey it, abusive # ones ignore it. The enforcement is Super Bot Fight Mode at the Cloudflare # edge; this list is the courtesy signal, not the control. # --------------------------------------------------------------------------- # Meta. FacebookBot is the speech-data crawler and meta-externalagent is the AI # crawler - neither is the link-preview fetcher, so blocking them does not # affect how our links render when someone shares one. User-agent: FacebookBot Disallow: / User-agent: meta-externalagent Disallow: / # Low-value data scrapers. User-agent: MJ12bot Disallow: / User-agent: BLEXBot Disallow: / User-agent: DotBot Disallow: / User-agent: omgili Disallow: / # --------------------------------------------------------------------------- # AI crawlers: ALLOWED, at a slow crawl. Decided 2026-09-09, having previously # been blocked outright as part of slowing down bot traffic. # # Being the answer when someone asks an assistant "what software runs youth # tournaments" is an acquisition channel, and we had opted out of all of it. # ChatGPT-User is the one that mattered most: it is not a scraper at all - it # fires when a person in ChatGPT asks about us or follows a link to us, so # blocking it turned away a prospect mid-research. # # The cost is bounded at the edge, not here. Cloudflare rate-limits these agents # per user agent; robots.txt is advisory and was never the control. See # docs/reference/seo-crawler-controls.md, which also has to be in place before # this file is deployed. # --------------------------------------------------------------------------- User-agent: GPTBot Crawl-delay: 10 User-agent: OAI-SearchBot Crawl-delay: 10 User-agent: ChatGPT-User Crawl-delay: 10 User-agent: ClaudeBot Crawl-delay: 10 User-agent: Claude-SearchBot Crawl-delay: 10 User-agent: PerplexityBot Crawl-delay: 10 # Common Crawl stays blocked: it is a bulk dataset harvester rather than a # product that sends anyone back to us, so it is all cost and no channel. User-agent: CCBot Disallow: /