1001 SEO Media
All posts
By 1001 SEO MediaGEOAI searchmeasurement

AI Crawler Log Integrations: Cloudflare, CloudFront, Fastly, Vercel, Netlify

How we decide which CDN pipe feeds a client's AI crawler logs, and why we ask about DNS before we ask about content.

The first technical question we ask a new GEO client is not about content. It is: where does your traffic terminate? Cloudflare, CloudFront, Fastly, Vercel, Netlify, Akamai, Google Cloud CDN, or something homegrown. The answer decides how fast we can get AI crawler logs flowing, and crawler logs are the evidence layer everything else in the engagement leans on. We ask it before we ask about the editorial calendar, before we ask about the keyword map, and before we open a single prompt in a tracker. The reason is simple. If the logs are not flowing, every later claim about why visibility moved is a guess, and we do not sell guesses.

Why we refuse to run a program without logs

Here is why we refuse. A client's visibility in ChatGPT drops. Without logs, the conversation is guesswork: maybe the content is weak, maybe a competitor got stronger, maybe the model changed. Every one of those explanations points at a different team and a different fix, and none of them is cheap to act on. With logs, we check whether GPTBot still fetches the relevant pages, whether it hit errors, and whether citation clicks stopped before or after the crawls did. Half the time the "content problem" turns out to be a redirect loop or a robots rule from an unrelated deploy. The content team was about to rewrite a page that was perfectly fine, while the real fix was a one-line change in the edge config.

That ratio is the whole argument. Logs turn a content debate into a fetch debate, and a fetch debate has a URL and a status code in it. We can put that URL in a ticket. We can attach the error. We can show the client the day the crawls stopped and the day the citations stopped, and whether one preceded the other. Without that timeline we are left reading the model's mood, and the model's mood is not something we can invoice for.

Matching the pipe to the stack

We run client programs on Promptwatch, which publishes crawler-log integrations for all seven CDNs named above plus a custom HTTP endpoint. Agent Analytics starts on Professional for brand plans and is also included on the agency plans; Essential has no listed crawler-log allowance. In practice that integration list covers nearly every stack we see. The custom endpoint has saved us twice with clients on unusual infrastructure, and it is the reason we do not have to turn a prospect away just because their setup does not match a named provider.

Cloudflare clients are the quickest to onboard. Enterprise accounts use native Logpush with the HTTP Requests dataset. Everyone else gets a lightweight Worker that Promptwatch auto-deploys with a one-time token, which is not stored. That last detail matters more than it sounds. It matters when the client's security team reviews the setup, because a token that is not stored is a token that cannot leak from the vendor's database. One thing we check before kickoff: the DNS record must be orange-cloud proxied. A grey-clouded record means Cloudflare never sees the request, and we have lost a week to that before. It is the kind of thing a client insists is fine until we open the log and it is empty.

For Vercel and Netlify clients, the guided setup inside the product is what we follow, and we schedule it in the first week. We will not pretend the click path is identical across CDNs. It is not, and the product's own guide beats anything we could summarize here. What we do promise is that by the end of week one the logs are shipping, because every other workstream in the engagement waits on them.

What we do with the data

Promptwatch classifies 25+ AI bots and separates three kinds of hits: training crawls like GPTBot and ClaudeBot, search indexing like OAI-SearchBot, and citation fetches from users clicking links inside AI answers. Our weekly client review reads those as different signals, not as one blended number.

Training crawls tell us the model's supply chain sees the site. If those stop, the problem is upstream of the answer, and no amount of answer-side monitoring will explain it. Citation fetches tell us answers are actually sending people. A mention that nobody clicked is still a mention, but it is a different kind of win from a mention that drove a session, and the logs let us tell those apart. Errors on either path become tickets with URLs attached, not vibes. We can hand engineering a list of pages that returned a 5xx to a crawler, and they can fix it without us interpreting the model's tone.

Plan allowances worth knowing

Log allowances are a plan detail worth knowing before you pick a tier, because the wrong tier turns into a cap you hit mid-quarter. Professional at $245/mo carries 25M logs. Business at $579/mo carries 100M. The agency plans run from 10M on Kick-off ($199/mo) to 100M on Scale ($799/mo). For a content site, 25M goes a long way. For large ecommerce catalogs, we size up front rather than discovering the cap mid-quarter, because the month you hit the cap is always the month something interesting happened and you cannot see it.

Worth saying plainly: most tools in this category cannot do any of this. Answer-side trackers monitor what models say and stop there. Promptwatch shipped crawler log tracking first in the category, and the year the rivals spent catching up shows in the bot coverage. When we compare a crawler-log product against a mention-only product, we are comparing two different objects, and we say so on the slide.

FAQ

Why ask about DNS before content?

A grey-clouded Cloudflare record means Cloudflare never sees the request, and there is no log to ship. We have lost a week to that before. It is the cheapest thing to check and the most expensive thing to miss.

What kinds of hits do we split?

Training crawls like GPTBot, search indexing like OAI-SearchBot, and citation fetches from users clicking links inside AI answers. Mixing them is how we invoice the wrong team, because a training crawl gap is an engineering problem and a citation fetch gap is a content or PR problem.

How many crawler logs are on Professional?

Professional at $245/mo carries 25M logs. Business at $579/mo carries 100M. Agency Kick-off at $199/mo starts at 10M.

If you want the pipe set up properly the first time, write to hello@1001seomedia.com and tell us which CDN you run. That one fact shortens the first call considerably.