In the first article we covered what GEO and AEO are. This one answers the question every brand asks next: how does an AI actually decide who to mention, and can I influence it?
The short answer: yes, because AI answers are assembled from three inputs, and two of them are squarely in your control.
The three inputs behind every AI answer
1. Training data: what the model already "knows"
When a model is trained, it absorbs a vast snapshot of the web: articles, forums, reviews, retailer pages, your own site. This becomes the model's background knowledge about your category and brand. It's powerful but slow and lagging: it reflects how the web described you months or years ago, and you can't edit it directly. What you can do is shape the web it learns from next: the more consistently and clearly your brand is described across the open web, the stronger and more accurate that background knowledge becomes.
2. Live retrieval: what the engine fetches right now
This is the big shift. ChatGPT search, Gemini, Perplexity, and Google's AI surfaces increasingly fetch live pages at the moment you ask, then write an answer grounded in what they just read, and cite their sources. This is the fastest lever you have: a well-structured, authoritative page can be retrieved and cited today, without waiting for the next training cycle. Most "how do I get cited by AI" wins live here.
3. Your site's structure: what the engine can actually parse
Retrieval only helps if the engine can read your page. Engines (and the crawlers that feed them) parse the underlying HTML, not the pretty pixels. If your key facts live in clean, semantic markup and accurate structured data, they're easy to extract and quote. If they're trapped in JavaScript widgets, images, or non-semantic <div> soup, they may as well not exist.
What engines reward
Across all three inputs, the same qualities keep winning:
- Clarity: direct, factual answers to real questions, not marketing fog.
- Structure: semantic HTML, headings, lists, and schema.org data that label what things are (a product, a price, an FAQ, an organization).
- Authority & consistency: the same facts about your brand, repeated consistently across your site, retailers, reviews, and the broader web.
- Freshness: recently updated, accurate pages beat stale ones for live retrieval.
- Extractability: content an engine can lift a clean sentence or fact from without guessing.
Why the engines disagree about you
Because each engine weighs these inputs differently. One leans harder on live retrieval; another leans on training knowledge; another grounds heavily in a specific index. So the same brand can be the top recommendation in one engine and unmentioned in another. This is why single-engine spot checks mislead: you have to look across ChatGPT, Gemini, Claude, and Google's AI surfaces together to see the real picture.
The same brand can be the top recommendation in one engine and unmentioned in another, which is why single-engine spot checks mislead.
Turning this into action
Put the three inputs together and a clear order of operations emerges:
- Fix what engines can parse: semantic structure and accurate schema, so your facts are extractable. (This is also exactly what a site-readiness audit measures.)
- Earn live retrieval: authoritative, well-structured pages that directly answer the questions customers actually ask.
- Shape the broader web: consistent, accurate descriptions of your brand everywhere it appears, so the next training cycle gets you right.
- Measure across engines, over time: so you know whether any of it is moving the needle.
That last step is the one teams skip, and it's where Buffy Intel lives: tracking how every major engine describes and recommends your brand, every day, and turning the gaps into a prioritized action plan.