AI crawlers are the robots that ChatGPT, Claude, Perplexity and Google's Gemini send to read websites, and a one-page file called robots.txt tells them whether they may come in. We read that file on 43 Sacramento-area brokerage websites on October 11, 2026. Not one blocks the crawlers behind ChatGPT, Claude, Perplexity or Gemini. Two block ByteDance's crawler, and that is the only AI block in the sample.
So if an AI assistant never mentions your brokerage, robots.txt is almost certainly not the reason. That is worth knowing before anyone sells you a fix for it. Below is what the file controls, the one choice in it worth making on purpose, and the one place a block can hide where robots.txt will never show it.
The sample is the same 43 brokerages we timed on a phone, read for their Google site name and checked for security. The method is at the end, and no firm is named.
What the AI crawlers are
Each AI company publishes the names of its crawlers, and most now run separate ones for training and for search. That split is the whole story, because blocking one does not block the other.
OpenAI (ChatGPT)
Anthropic (Claude)
Perplexity
Google (Gemini)
| Company | Training crawler | Search crawler | Fetches when a user asks |
|---|---|---|---|
| OpenAI (ChatGPT) | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic (Claude) | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | none listed | PerplexityBot | Perplexity-User |
| Google (Gemini) | Google-Extended (a setting, not a separate robot) | Googlebot, as for normal Search |
OpenAI's crawler page puts it plainly: "Each setting is independent of the others," so a site "can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot." It also says what blocking search costs: "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers."
Anthropic says the same of its search crawler: disabling Claude-SearchBot "may reduce your site's visibility and accuracy in user search results." Perplexity says PerplexityBot "is not used to crawl content for AI foundation models." And Google says Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search."
What we found on 43 brokerage sites
Serve a robots.txt file
Have no robots.txt (the address returns "not found")
File unreadable because the certificate expired in 2019
Block any OpenAI, Anthropic, Perplexity or Google crawler
Block ByteDance's crawler (Bytespider)
Name an AI crawler anywhere in the file
Have an llms.txt file
| What we checked | Sites |
|---|---|
| Serve a robots.txt file | 38 |
| Have no robots.txt (the address returns "not found") | 4 |
| File unreadable because the certificate expired in 2019 | 1 |
| Block any OpenAI, Anthropic, Perplexity or Google crawler | 0 |
| Block ByteDance's crawler (Bytespider) | 2 |
| Name an AI crawler anywhere in the file | 7 |
| Have an llms.txt file | 9 |
A missing robots.txt is not a problem. Under the internet standard for the file, a "not found" answer means a crawler "MAY access any resources on the server." The four sites without one are open to everyone, the AI crawlers included.
The seven files that name an AI crawler mostly do not block it. Four are Squarespace sites that list 26 AI crawlers by name, but put them in the same group as every other robot, with rules that only fence off a few back-office paths, so the homepage stays open to all of them. One names seven AI crawlers just to say Allow: /. The two that block ByteDance also block several SEO tools, which looks like a deliberate cleanup rather than a stance on AI.
One site adds a newer line, a content signal, reading search=yes, ai-input=yes, ai-train=yes: yes to all three. Nobody in the sample says no to training.
The site most closed to AI is the one nobody chose to close. Its security certificate expired in May 2019, so anything that insists on a secure connection, crawler or person, cannot get in at all. We covered that in our security check.
The one place a block can hide
robots.txt is a request the site publishes, not a lock. The standard says so in as many words: "These rules are not a form of access authorization." A block can also live somewhere robots.txt never shows: the network service in front of your site.
The big one is Cloudflare. On July 1, 2025 it announced that "every new domain will now be asked if they want to allow AI crawlers." Answer no, and the AI crawlers are turned away before they ever read your robots.txt. Twelve of the 43 brokerage homepages are served through Cloudflare. We cannot tell from outside whether any of them said no, because we read sites as an ordinary browser and never pretend to be a crawler. Only the account owner can see that setting.
If your site uses Cloudflare and you do not remember answering that question, it takes two minutes to check in the dashboard.
Do you need an llms.txt file?
Nine of the 43 have one. llms.txt is a proposal for a plain-text summary of a site written for AI tools, and two of the nine say in the file itself that an SEO plugin generated them.
It does no harm. But Google, the only AI company whose guidance we found addressing it, says you do not need it: "You don't need to create new machine readable files, AI text files, or markup to appear in these features," its AI features guidance says of AI Overviews and AI Mode. What Google does require is ordinary: the page must be "indexed and eligible to be shown in Google Search with a snippet." If someone is selling you llms.txt as the way into AI answers, ask them where an AI company says so.
The one choice worth making on purpose
Most brokerages should leave everything open, which is what all 43 have done. Being findable is the point of the website.
If you would rather your listings copy and blog posts not be used to train AI models, there is a clean way to say so without disappearing from AI search. Block the training crawlers by name and leave the search ones alone:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
That leaves OAI-SearchBot, Claude-SearchBot, PerplexityBot and Googlebot free to read your site, so you can still turn up in ChatGPT search, Claude, Perplexity and Google. What you should never block, unless you mean to vanish from those answers, is the search crawlers. OpenAI says a change takes about 24 hours to reach its systems.
Our own site
We ran the same check on www.flarebuilt.com. Our robots.txt names 14 crawlers, AI and search alike, and lets all of them read everything except our /api/ paths. We serve an llms.txt, generated from the site so it cannot fall out of date.
The gap, which we would rather name than leave out: allowing every crawler did not make AI assistants know who we are. Several unrelated companies share the Flare name, and an assistant needs evidence from other sites to tell them apart. robots.txt gives permission. It does not give presence.
Check your own site in five minutes
- Open yoursite.com/robots.txt. Look for
Disallow: /underUser-agent: *(that blocks everyone, Google included) or under any of the crawler names in the table above. - If you use Cloudflare, check its AI crawler setting. A block there never appears in robots.txt.
- Decide on training separately from search. If you want to opt out of training, block GPTBot, ClaudeBot and Google-Extended, and leave the search crawlers open.
- Skip the llms.txt sales pitch. Spend that effort on the things AI tools actually read: a clear site name, a Google Business Profile, and a site that loads on a phone.
- Ask an assistant about yourself. Type your brokerage's name and city into ChatGPT, Claude or Perplexity. If it does not know you, the fix is more evidence about you, not a change to robots.txt.
What we would build
A brokerage site that every search and AI crawler can read from the first day, with a robots.txt written on purpose rather than left to a template, a site name Google can confirm, and the license number on the page. A launch site starts at $900 and is live in 48 hours, and a full brokerage site starts at $4,500 and is live in about a week, with Flare Care keeping it current for the months after.
If you want this check run on your site, tell us about your brokerage, or build your package in about two minutes.
How we checked
The sites are every live website from our two earlier Sacramento-area audits: brokerage corporations licensed in Sacramento County since 2024 (residential) and Sacramento-region corporations with Commercial or CRE in their licensed name (commercial), the same 43 as our speed, site name and security checks.
On October 11, 2026 we requested two public files from each site, /robots.txt and /llms.txt, once each, as an ordinary desktop browser. Nothing else was requested. We read each robots.txt the way the standard tells a crawler to, and asked one question for each crawler name: may it read the homepage? Which sites run through Cloudflare comes from the response headers we recorded on October 8.
What it cannot see: blocks applied by a firewall or a service like Cloudflare, and whether any AI assistant actually mentions these firms. These files change whenever a site owner edits them, so this is a snapshot of one day.
