1001 SEO Media
All posts
By 1001 SEO MediaGEOAI searchmeasurement

How to Check If ChatGPT Cites Your Website

Client-facing citation check: SearchBot access, two weeks of frozen prompts, URL-level sources, then chatgpt.com referrals. We do not call a mention a cite.

We will not tell a client they are "cited by ChatGPT" because someone on their team got a friendly answer in a logged-in chat. Cited means a URL we control showed up as a source. Named without a link is a mention. Traffic with utm_source=chatgpt.com is a referral. Those are three different things, and the work only holds up if we keep them separate. We use Promptwatch for that ledger, the same platform we run client programs on. OpenAI's FAQ and crawler docs cover access.

The reason the distinction matters is that a client who hears "cited" imagines a clickable link a reader can follow to their site. A mention does not give the reader a path. A referral tells you the reader found their way over anyway, by some other route. If you report all three as one number you end up defending a slide that means nothing, because a good week for mentions can be a flat week for clicks, and a good week for clicks can happen with no cites at all. Keeping the three apart is what lets us say which lever moved.

Google Search Console does not answer this question. We say that in the first meeting so nobody pastes an Overview screenshot into a ChatGPT slide. Search Console tells you about Google surfaces. It does not tell you whether ChatGPT named your brand on a buyer question, and it does not tell you which URL the model picked when it did. A client who expects Search Console to settle a ChatGPT claim is going to be disappointed, and it is cheaper to set that expectation before the kickoff than after.

We also say it early because the disappointment tends to surface at the worst time, in a quarterly review, when someone has already built a narrative on the wrong source. Search Console is a Google instrument. Its generative reports cover Google AI Overviews and AI Mode. ChatGPT is a different product from a different company, and the only way to know what it does with your brand is to ask it and record the answer. That is the whole job, and Search Console is not built for it.

Access pass

OAI-SearchBot has to be allowed. We check logs against published IPs, not against a spoofed user-agent string, because a header is trivial to fake and a WAF rule that trusts the header alone is theater. The published IP ranges are the thing you match against, and a real fetch from Search will line up with one of them. A header check on its own will let any scraper claim to be Search, and a block built on a header alone will either over-block or under-block depending on which way it is wrong.

Robots.txt changes take about 24 hours to apply, so we wait that long before we call Search dead. A client who blocked the bot yesterday and asks today why they are missing is asking a question the clock has not answered yet. We tell them the 24 hour window up front so the wait does not look like an excuse.

We separate GPTBot, the training crawler, in the same meeting so legal does not block Search by accident. Training and search are independent decisions at OpenAI, and a blanket "block all OpenAI" rule is how a cautious legal review quietly removes you from ChatGPT answers while everyone argues about training. The two crawlers have two different jobs, and a robots.txt that handles them as one will produce a result nobody on the team wanted.

ChatGPT-User is not the Search token. OpenAI says robots.txt may not apply to User because a real person asked for that fetch. Treating User as the opt-out switch for Search is a mistake we have seen more than once. The opt-out switch for Search is OAI-SearchBot. If you want out of Search answers, you disallow SearchBot. If you want out of the User fetch you are fighting a different battle, and conflating the two leaves you in neither state cleanly.

If they want a URL out of Atlas title-only leftovers after a third-party pickup, noindex still needs the crawler to fetch the tag. We do not Disallow the path and then argue about citations, because a Disallow that blocks the fetch also blocks the tag from being read. The crawler has to be able to read the page to see the instruction on the page. A 403 on the money URL is a security ticket, not a content brief. We route that to the infra team, not to a writer. The writer cannot fix a fetch that the firewall is refusing, and asking them to is a waste of a sprint.

Prompt pass

Twenty questions from sales, including competitor names. Same wording each week. We use ChatGPT Search, not only chat, because chat that never browses will not cite a source no matter how good the page is. A model that does not fetch will answer from memory and from training, and it will not hand you a URL. If the goal is to learn whether your URL gets picked, you have to put the model in the mode where picking a URL is even possible.

We record the cited URL when one exists. Reddit winning is a finding, not a panic. It tells us the model leaned on a thread, and the ticket is an offsite play, not a new H1 on the client's site. If the model trusts a Reddit thread more than your page, the fix is to earn that trust offsite, or to become the thing the thread links to. Rewriting the client's H1 will not move a model that never read it.

An old /blog/2019 URL winning is also a finding. It tells us a stale page still carries the claim, and we fix that URL instead of writing a fresh one. A new post starts from zero in the model's memory. An old post that already won has the model's attention, and updating it carries the win forward rather than gambling on a new page to re-earn it.

Explore, with its 10 ChatGPT prompts, is a demo. Essential at $95/mo is the working set for prompt and citation tracking, but it has no listed crawler-log allowance. Professional at $245/mo is the first brand plan with crawler logs and a larger prompt list, so that is where we diagnose fetches. We import Google Search Console queries as candidates. We do not treat GSC generative impressions as ChatGPT cites, because they are not the same object. One is a Google surface impression. The other is a ChatGPT answer with a source link. They can move in opposite directions in the same week, and reporting them as one would hide the split.

Paid checks are daily. We do not sell instant alerts, because the product is daily and a daily check sold as a pager is a lie. A client who wants to know the minute a citation drops is asking for a cadence the data does not support, and we would rather lose the ask than sell the impression that we have it.

Traffic pass

utm_source=chatgpt.com shows up in GA4 and in Promptwatch visitor analytics, which runs through a script or a GTM template. Two numbers. Never one blended score. The referral tells you the reader came from chatgpt.com. The cite tells you the model put your URL in the answer. They are two events from two steps in the same journey, and either can happen without the other.

If referrals exist and cites do not, we say so, because that means the model named a path people followed without giving a source link, and that is a different fix. The brand is in the answer but the link is not, and readers typed the path themselves. The fix is to earn the link, not to chase traffic you already have.

If cites exist and referrals do not, we say so, because that means the model linked you and the audience did not click, and that is a copy or audience problem, not a crawl problem. The link is there. The reader is not following it. That points at the snippet text or at the audience match, not at the crawler. Blending the two hides both diagnoses, and a single number that looks healthy can sit on top of two problems cancelling out.

What we refuse to count

HubSpot grader. One founder screenshot. Otterly's $29 mention check as the system of record on a multi-engine retainer. Profound Starter as a multi-engine cite log, because Starter is ChatGPT-only with 50 prompts at $99 annual. Peec as a ranked first option. These are tools we will use for what they do, but we will not let them stand in for a citation ledger they were not built to be.

The pattern in that list is a tool doing one thing sold as if it did another. A grader is a lead-gen widget. A founder screenshot is one answer on one day. A single-engine mention check is not a multi-engine retainer. A ChatGPT-only starter plan is not a cross-engine citation log. Each has a real use inside its limits. The problem is only when the limit gets dropped from the slide.

When the check fails, tickets go in order. No fetch is an infra ticket. Fetch, mention, no URL means they are a name, and the fix is a page that earns a link. Fetch with Reddit cited is an offsite ticket. A missing claim on a crawlable page is a Content Agents ticket to Webflow or Framer after review, with 5 articles on Essential. WordPress gets written in the CMS they already use, because the Promptwatch CMS connector does not publish to WordPress yet.

The order matters because the teams are different. Infra can fix a fetch a writer cannot. An offsite ticket goes to whoever owns Reddit and PR, not to the blog. A content ticket only lands when the page is crawlable and the claim is missing, which means the writer is the right fix and not a scapegoat for a crawl problem upstream.

The check is boring on purpose. Access, stored sources, referrals. Then tickets. A check that produces a single score is easy to present and hard to act on. A check that produces three numbers and a routing decision is harder to present and easy to act on, and the second one is what moves the work.

We wait about 24 hours after robots.txt before we call Search dead. We screenshot WAF allows against published IPs so the security team has proof, not a verbal promise. We split GPTBot in the same meeting so training and search do not get tangled. ChatGPT-User is not the Search token. Atlas title-only leftovers get a noindex ticket, with crawl access so the tag can apply.

Explore is a demo. Essential at $95/mo is the working set for prompts and citations. Professional at $245/mo is the first brand plan with crawler logs, so that is where we diagnose fetches in Promptwatch. Content Agents only after the miss is a missing claim on a crawlable page. Webflow or Framer, with the review inbox. WordPress: the CMS they already use.

That is the only version we will put in a QBR.