1001 SEO Media
All posts
By 1001 SEO MediaGEOAI searchmeasurement

How to Connect Cloudflare Crawler Logs to Promptwatch

How we connect Cloudflare crawler logs to Promptwatch with Logpush on Enterprise or a Worker on any plan, then read training versus search fetches.

If ChatGPTBot never fetched the page, we do not rewrite the H1, because rewriting a page the bot never saw is copywriting to an empty room. On Promptwatch Professional, Business, or an agency plan, we connect Cloudflare to Agent Analytics and read the log. Essential has no listed crawler-log allowance, so a team on Essential does not have the pipe and should not be promised it. This is the same order as how to get cited by ChatGPT: crawlers first, because a crawl miss explains a citation miss, and a citation miss without a crawl miss is a content problem instead of a robots problem.

Two install paths. Cloudflare Enterprise can use Logpush, which is the Enterprise-grade path and the one a security team will prefer. Any plan can use the Worker, which is the path that works on Pro and Free zones. The Worker uses a one-time token that Promptwatch does not store, so the credential does not sit in a vendor database. We do not leave that token in a wiki, because a token in a wiki is a token in a place that gets indexed and forgotten.

DNS has to be orange-clouded. If the hostname bypasses Cloudflare, there is no log to ship, because the traffic never touches Cloudflare's edge and Cloudflare cannot log traffic it does not see. We check that before we file a ticket with their IT, because a ticket about a missing log on a bypassed hostname is a ticket that wastes a week. Confirm the hostname is orange-clouded. If it is not, stop, because nothing else in this setup will work until that is true.

The site can be on Cloudflare and still fail this, and that is the trap. If the apex is grey-clouded, it will not work for that hostname, because the apex is the record that matters and a grey apex is a bypassed apex. Orange-cloud the host you care about, or we connect a different CDN integration if they actually terminate TLS elsewhere, because the question is where TLS terminates, not where the DNS record points. We will not pretend a bypassed record is sending logs, because a pretense here becomes a wrong diagnosis later.

Logpush or Worker

Logpush is the Enterprise path. If they already pay for Logpush elsewhere, we reuse the habit, because a habit that already exists is cheaper than a new one. Worker is what we use on Pro and Free zones when Enterprise is not on the table, and it works on the plans most clients actually run. Same destination: Promptwatch crawler logs. We do not invent a third method, because a third method would be a third thing to maintain and the two that exist cover the cases that matter.

Does the Worker store Cloudflare credentials? No. The Worker path uses a one-time token that is not stored, so the token is used and then it is gone. If their security team wants the Logpush path instead, that is an Enterprise conversation on the Cloudflare side, and it is a real conversation because Logpush is an Enterprise feature. Enterprise: Logpush into Promptwatch. Any other plan: deploy the Worker with the one-time token, because the Worker is the path that works without Enterprise.

After logs flow, we care about three kinds of hits, because they are not the same ticket and treating them as the same ticket is how the wrong team gets invoiced. Training crawls (GPTBot and friends eating the corpus), which is the bot reading the web to train a model. OAI-SearchBot and other search-time bots, which is the bot reading a page to answer a question right now. Citation fetches, when a model retrieves a URL to support an answer, which is the moment a citation is born.

A block on training is a robots/CDN conversation, because training is governed by robots and by the CDN's bot rules. A miss on citation fetch after a successful read is a content or authority conversation, because the bot saw the page and chose not to cite it, which is a content problem. Mixing them is how we invoice the wrong team, because a training block goes to the robots/CDN owner and a citation miss goes to the content owner, and the two owners are often different people. Wait until ChatGPTBot, ClaudeBot, PerplexityBot, or GoogleOther show up in the log before you diagnose, because a diagnosis without a bot in the log is a guess. Filter to the URLs in the SOW, because the SOW is the scope and the scope is what you report. Split training vs OAI-SearchBot vs citation fetch before you file a content ticket, because a content ticket filed against a training block is a ticket that fails.

Page inventory from sitemaps and GSC URL import sits next to these rows, because a log without a page list is a log without context. We want the URL list and the fetch list in one project, so a fetch can be matched to a page. Put sitemaps or GSC URLs in inventory so the log has a page list to sit next to, because the log alone tells you what was fetched and the inventory tells you what was supposed to be fetched. See import sitemaps and Search Console URLs.

What we still need from Cloudflare

Orange cloud on the hostname we claim to measure, because without it there is no log. A Worker or Logpush that actually runs, because a deployed Worker that does not run is the same as no Worker. Bots allowed in robots.txt when the SOW says they should be allowed, because a robots block will look like a citation miss and a robots block is not a content problem. Promptwatch will show the 403, because a 403 is a fact in the log. It will not fix the WAF rule, because the WAF rule is Cloudflare's and fixing it is a Cloudflare conversation.

We keep GA4. Crawler logs are not visitors, and visitors are not crawlers, and the two pipes measure different things. Visitor analytics is a different tag, and a team that wants both runs both, because the two are not substitutes.

Professional is the first brand plan with crawler logs and includes 25 million at $245/mo, which is the plan where a single-brand program gets the full pipe. Kick-off includes 10 million at $199/mo, with unlimited projects and prompts plus 10 seats, so an agency runs many clients on one plan. Business is $579/mo with 100 million logs, for a larger property or portfolio. Essential at $95/mo has no listed crawler-log allowance, so we do not use it for this setup, because using it would be promising a pipe it does not have. Explore is free ChatGPT prompts only, which is a learning tier. G2: 4.7/5.

No instant alerts on errors. We read the log on the same cadence as the rest of the work, because a cadence is a habit and a habit is what gets read.

FAQ

Does the Worker store Cloudflare credentials?

No. The Worker path uses a one-time token that Promptwatch does not store, so the token is used and discarded. If their security team wants the Logpush path instead, that is an Enterprise conversation on the Cloudflare side, because Logpush is an Enterprise feature.

What if the hostname is grey-clouded?

It will not work for that hostname, because a grey-clouded host bypasses Cloudflare's edge. Orange-cloud the host you care about, or we connect a different CDN integration if they actually terminate TLS elsewhere, because the question is where TLS terminates.

Is Logpush required?

Only on Cloudflare Enterprise. Any other Cloudflare plan deploys the Worker with the one-time token, because the Worker is the path that works without Enterprise. The destination is Promptwatch crawler logs, so the Promptwatch project still needs Professional, Business, or an agency plan, because that is where the log pipe lives.

If you want us to wire it, hello@1001seomedia.com. Send the zone name and whether the plan is Enterprise, because the plan decides the path.