HomeAI SearchAbout MeContact

Before rewriting for AI search, check whether your page can be read

By Sumith Parambat DamodaranPublished in TechnologyOctober 11, 20266 min read
Before rewriting for AI search, check whether your page can be read

A page can look complete in a browser while its initial response contains very little of the answer. Before commissioning another GEO article, I would check that the evidence already published can be retrieved. Otherwise the content team may be solving a delivery problem with more words.

The page that passed the wrong test

Imagine a fictional booking product, StudioLedger. Its rescheduling page has a clear heading, a polished animation and an accordion explaining payment restrictions. A marketer opens it, reads the answer and approves it. But suppose the initial HTML contains the heading and a loading shell; the restriction arrives through a later request.

The marketer tested a rendered experience. The missing test was whether the important fact was present in the response available to the intended consumer. That distinction matters before asking whether the wording is persuasive.

This is a hypothetical diagnosis, not a claim about StudioLedger or a production customer. The same investigation is useful for a documentation site, a product comparison or a pricing page.

Name the access you actually want

“Allow AI bots” is too vague for an implementation ticket. Search discovery, model training and a user asking an assistant to visit a specific URL are different purposes. The owner of the site should make those decisions explicitly.

OpenAI’s documentation separates OAI-SearchBot for search, GPTBot for potential training use and ChatGPT-User for certain user-triggered visits. Search and training controls are independent. The documentation also says robots.txt may not apply to user-triggered actions. Check the current official guidance before changing a production policy. OpenAI crawler documentation

A user-agent label in a log is also not proof of identity. If that distinction affects an access decision, verify against the provider’s current published information rather than trusting the string alone.

A three-layer audit

LayerQuestionEvidence to retainWhat it cannot establish
AccessCan the intended consumer reach the approved URL?Status, redirects, applicable policy and challenge responseThat the answer will use it
ContentDoes the retrieved response contain the relevant claim and conditions?Initial HTML and the corresponding rendered sectionThat every service processes it identically
SelectionDoes a sampled answer cite or correctly represent the page?Exact prompt, environment, answer and source URLThe private ranking or retrieval pipeline

The layers are a debugging sequence. They are not a ranking formula. A page can pass the first two and still be absent from an answer because another source is selected or no external retrieval occurs.

Run one useful inspection

Choose a public page that answers a real buyer question. Write down one fact that must survive: for example, that a rescheduled booking retains its payment only under specified provider and price conditions.

  1. Retrieve the page and record the final URL, HTTP status and any challenge or login screen.
  2. Inspect the initial HTML for the fact, its limitation and a link to the supporting policy. Compare this with what a browser shows after loading.
  3. Check that an ordinary internal link leads to the authoritative page. The link should describe the decision or evidence it helps with.
  4. Check the canonical URL and indexing controls separately. A crawl permission decision is not the same as an indexing decision.
  5. If logs are available, inspect a defined window and the relevant URL. Record the response status and verified agent category; do not infer a recommendation from a request.
  6. Repeat the same inspection after a correction. Keep the before and after artefacts together.

Do this for one important page before expanding to a site-wide backlog. The first pass should produce a reproducible finding, not an impressive total of bot names.

Write a ticket somebody can close

An unhelpful ticket says, “Make this page GEO ready.” A useful one says, “The rescheduling restriction is visible after hydration but absent from the initial HTML. Include the approved answer and conditions in the generated response, and keep the visible version consistent.”

Its acceptance criteria are concrete: the public URL returns the expected page; the initial HTML includes the exact approved restriction; the rendered page agrees; and the supporting policy link works. The content owner approves the wording, while engineering owns delivery.

Static generation and server rendering are possible approaches, not prescriptions for every site. My own blog generates article HTML at build time. That gives me an inspectable output; it does not prove that every AI service will select an article.

Keep logs in their lane

More bot requests can mean many things: discovery, repeat visits, errors, training collection or user-driven fetches. A crawl-to-click ratio does not describe the commercial value of a particular page. Join observations carefully with referral data and sampled answers, and retain the gaps rather than filling them with a story.

For the same reason, I would not use a successful “summarise this URL” request as the sole test for automatic search discovery. It checks a different path. The measurement guide explains how to keep coverage, citations, accuracy and outcomes separate.

What I would fix first

Start with a missing or contradictory fact on a page that matters to a customer. Make the answer available in the approved delivery path, keep the limitation beside it and make the page easy to find. Then repeat the observation that originally exposed the problem.

The deliverable is a better source page and a closed, evidenced ticket. Citation changes can be investigated afterwards without becoming the only reason the work has value.

Sources and attribution

  • Charlie Norledge and Petar Jovetic, Advanced GEO & AI Optimisation, supplied reference deck, pages 68–88, particularly the initial-HTML comparison and rendering approaches on pages 73–78. Reviewed as reference material.
  • OpenAI crawler documentation, reviewed 11 October 2026. The provider-specific distinctions above follow this source; the audit and example are my proposed application.
  • Google’s crawlable link guidance. Descriptive, ordinary links are useful navigation, without guaranteeing selection in generated answers.

Browse the AI search field guide for the complete reading paths.


Tags

AI SearchTechnical SEOCrawlingGEO
Previous ArticleEntities, schema and RAG: a practical evidence map
Sumith Parambat Damodaran

Sumith Parambat Damodaran

Product Manager

Topics

General
Product Management
Technology

Related Posts

Your product story has a maintenance problem
Sumith Parambat Damodaran
October 11, 2026 6 min

Quick Links

AI search field guideAbout MeContact

Social Media