1001 SEO Media
All posts
By 1001 SEO MediaGEOAI searchmeasurement

Technical Optimization for AI Crawlers: Schema, Crawlability, and Robots.txt

How we use Promptwatch technical recommendations for AI crawler access, schema that matches the page, and robots.txt without treating llms.txt as a Google ranking file.

Technical GEO is still crawlability. Schema is a helper. A new text file is not a ranking system. We use Promptwatch technical recommendations, then we change robots, CDN rules, and structured data the way we already would for Googlebot. When the work also needs Agent Analytics, we use Professional or Business on a brand account, or a self-serve agency plan. The framing is the point: technical GEO is not a new discipline, it is the old crawlability discipline applied to a new set of bots, and treating it as new is how teams buy tools they do not need.

Google's own AI features and AI optimization notes still say: useful pages, crawlable, indexed, structured data that matches visible text. They do not say Google reads llms.txt to build AI Overviews. We will not claim that. The absence of an llms.txt claim in Google's documentation is the part that decides how we talk about the file, and we talk about it as optional documentation, not as a lever.

Crawlability before schema

Promptwatch crawlability analysis flags whether AI systems can fetch and parse the page. On Professional, Business, or an agency plan, we pair that with Agent Analytics errors. A block in robots.txt or a bot fight on the CDN shows up as a failed fetch. That is the first ticket. See how to get cited by ChatGPT for the bot list we check by hand: GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot, Google-Extended, GoogleOther. The bot list is the thing we check first, because a page that cannot be fetched is a page that cannot be cited, and no amount of schema fixes that.

A free robots.txt generator and a bots directory live on promptwatch.com if we need a first draft. The live test is still the log. We do not ship a generated robots.txt that allows a bot the security team has not approved. A generated draft is a starting point, and the log is the proof, and the gap between the two is where the security review happens.

Schema suggestions in the product are enhancement ideas, not a new vocabulary. We mark up what the user can see: Product, FAQ, HowTo, Organization, LocalBusiness when those are true. We do not add FAQ schema for questions that are not on the page. Invisible FAQ is how you get a manual action and still fail the assistant. The rule is simple and it is the rule Google itself states: the markup matches the page, and markup that does not match the page is a liability, not a ranking signal.

llms.txt can exist. Promptwatch even has a generator. We will publish one if the client wants a machine-readable summary. We will not tell them Google Overviews will move because the file landed. Promptwatch's own research page says llms.txt has no proven impact on AI search. We treat it as optional documentation, like a nicer about page for bots. The optional documentation framing is the honest one, and it is the framing that keeps a client from paying for a file that does nothing measurable.

What we change in practice

  1. Allow the crawlers in scope. Confirm in logs.
  2. Fix rendering so the claim we want cited is in the HTML, not only in a client-side widget.
  3. Add or fix schema that restates visible facts.
  4. Then rewrite the passage.

The order is the work. Crawl first, because crawl is the prerequisite. Render second, because a claim that lives only in a client-side widget is a claim the model never sees. Schema third, because schema restates what is already on the page. Copy last, because rewriting copy on a page that cannot be fetched is rewriting into a void. The order is what makes the pass technical rather than editorial, and the order is what makes it stick.

Otterly at $29 will not do this pass. Peec at $95 is mentions. Profound Starter at $99 annual is ChatGPT only. We still run Screaming Frog and Search Console. Promptwatch supplies the answer log on ChatGPT, Gemini, Claude, Perplexity, Grok, Llama, DeepSeek, Mistral, Copilot, AI Overviews, and AI Mode. The AI-crawler half starts on Professional for brand plans and is also included on the agency plans. The split between the answer log and the crawler log is the split that makes the pass diagnosable, and only Promptwatch gives you both halves in one project.

Essential is $95/mo and includes technical optimization, but it has no listed crawler-log allowance. Professional is $245/mo and is the first brand plan with Agent Analytics. Business is $579/mo. Kick-off is $199/mo with crawler logs, unlimited projects and prompts, and 10 seats. Explore is free, 10 ChatGPT prompts. G2: 4.7/5. Daily paid refresh. No instant alerts. No $29 Promptwatch plan. The plan ladder is built so a team can start on Essential for the crawlability pass and move to Professional when the crawler-log half becomes a QBR question.

What we will not sell as technical GEO

A paid "AI schema pack" from a local SEO vendor we have not verified. Invented BrightLocal or GeoRanker AI SKUs. A promise that llms.txt is required. A robots.txt that blocks GPTBot "for training" and then a complaint that ChatGPT Search cannot cite the docs. The four are the four ways technical GEO gets sold as something it is not, and each one is a product that does not do what its sales copy claims.

Local businesses still need the local SEO playbook. GBP and NAP are not replaced by a Promptwatch schema suggestion. A schema suggestion is a markup idea, and GBP and NAP are the local infrastructure that a markup idea cannot substitute for.

FAQ

Should we block GPTBot and allow OAI-SearchBot only?

That is a legal and security choice. If Search is in scope, OAI-SearchBot needs to fetch. If you block both, we will not run a citation program and then blame the model. The block is a choice, and the choice has a consequence, and the consequence is no citations.

Does schema get us into AI Overviews?

Google says structured data should match the page. It does not say a new schema type is an Overviews lever. We add accurate markup. We measure Overviews in the same Promptwatch project, not in the schema validator. The validator tells you the markup is valid. The project tells you the Overview appeared. Those are different answers.

How we run it

  1. Export current robots.txt and CDN bot rules.
  2. Run Promptwatch crawlability. On Professional, Business, or an agency plan, open Agent Analytics errors on the money URLs.
  3. Fix blocks. Re-fetch. Only then edit schema and copy.
  4. Skip any llms.txt work that is being sold as a Google ranking task.
  5. hello@1001seomedia.com if you want the technical pass on the retainer. Send robots.txt. Do not send a promised Overviews lift.