OpenAI Search Crawler OAI-SearchBot Documentation
Official OpenAI docs: OAI-SearchBot is for ChatGPT Search. GPTBot is training. ChatGPT-User is a user fetch. We log crawls in Promptwatch Agent Analytics.
OpenAI search crawler OAI-SearchBot documentation lives on Overview of OpenAI Crawlers. This is a technical SEO ticket, not a mention-dashboard ticket. OAI-SearchBot is for search. It surfaces websites in ChatGPT's search features. Sites opted out will not be shown in ChatGPT search answers, though they can still appear as navigational links. OpenAI recommends allowing OAI-SearchBot in robots.txt and allowing requests from published IP ranges (searchbot.json). The two halves of the recommendation are the robot file and the CDN allowlist, and both have to be done or the fetch still fails.
Example user-agent, with the caveat that the version may change: compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot. Robots.txt fetches may include an extra robots.txt marker in the UA.
We put this in the audit before anyone argues about GEO copy. If the search bot gets a 403, the mention program is theater, because no amount of content work will surface a page the engine cannot fetch. The 403 is the first thing to fix, and the content is the second, because the content only matters once the fetch succeeds.
Do not mix the bots
GPTBot is training data. ChatGPT-User is used when a user, or a Custom GPT, fetches a page. Robots.txt may not apply to ChatGPT-User the same way. It is not the Search opt-out control. Use OAI-SearchBot in robots.txt to manage Search. The three crawlers have three jobs, and the ticket has to name all three or the client gets a policy that does the opposite of what they asked for. A blanket "block OpenAI" rule is the most common way that policy goes wrong, because the rule blocks the search surface the report is meant to measure.
The publishers FAQ repeats the cost of getting this wrong. Block OAI-SearchBot and you hurt summaries and snippets. utm_source=chatgpt.com tags clicks when they happen. Referrals are a different post, and the referral column only fills if the search bot was allowed to begin with, so the referral and the crawl are two logs that depend on the same allow.
CDN WAF presets labeled "block AI bots" often catch all three crawlers. We read the access log for the real UA, then we change the preset. Robots.txt intent is not ground truth, because the CDN sits in front of the robots file and can override it. The log is the source, and the robots file is the intent, and the two can disagree.
Google-Extended is a Google control. It does not govern OAI-SearchBot. Do not copy a Google robots line into the OpenAI ticket, because the two publishers run separate crawlers with separate opt-out controls, and a Google line in an OpenAI ticket is a line that does nothing useful.
After the bot is allowed
Allowing the bot does not guarantee a citation. It removes a self-inflicted 403. Passage work is still how to get cited by ChatGPT. Google Overviews stay on Search Central: our reading. The unblock is the ticket that lets the rest of the program run, and it is not the ticket that earns the citation on its own.
Promptwatch Agent Analytics is where we log whether AI crawlers fetched the URL and whether that path later showed up as a citation. Professional is the first brand plan with crawler logs and includes 25 million at $245/mo. Kick-off includes 10 million logs at $199/mo, with unlimited projects and 10 seats. Business is $579/mo with 100 million logs. Essential is $95/mo and stores daily ChatGPT Search answers, citations, and visitor analytics, but it has no listed crawler-log allowance. Explore is ChatGPT prompts only, not the crawl stack we brief. The plan tier is the part that decides whether the fetch row is in the product, and the fetch row is the evidence the client is paying for.
Otterly, Peec, and Profound do not replace the robots.txt change. Ahrefs and Semrush stay for the rest of the crawl graph, because the link graph is still the input that decides whether a page can be cited at all, and the crawl graph is the context the visibility report needs.
What we will not do
We will not block GPTBot and tell a client they opted out of ChatGPT Search. We will not promise citations from a robots allow. We will not skip log verification. Each of those is a small discipline, and the log verification is the one that catches the CDN preset that quietly undid the robots change, because the preset is the layer that runs in front of the robots file.
FAQ
Is OAI-SearchBot the same as GPTBot?
No. Search versus training. They are separate crawlers with separate opt-out controls, and a block on one is not a block on the other.
Can we block GPTBot and allow OAI-SearchBot?
Yes. Those are different decisions. Document both in the ticket, so the next person to touch the robots file knows which surface each line governs.
Does Promptwatch fetch as OAI-SearchBot?
Promptwatch monitors the UI and, on Professional, Business, or an agency plan, logs crawler hits you already received. It is not a replacement for allowing OpenAI's bot, because the bot has to be allowed before there is a hit to log, and the log is a read on the bot, not a substitute for it.
What to do this week
- Read the bots page.
- Allow OAI-SearchBot on pages you want in ChatGPT Search.
- Allow the published IP ranges at the CDN. Confirm in logs.
- Connect Agent Analytics on Professional or an agency plan with a crawler-log allowance.
- Load the prompts those pages should win, then email hello@1001seomedia.com if you want the robots and CDN pass done on the account.