Kick Ads
Back to Blog

Generative Engine Optimization (GEO): How to Get Cited in Google AI Overviews, ChatGPT & Perplexity (2026)

By Frankie Chan, Co-Founder · Updated 24 July 2026

Generative Engine Optimization (GEO): How to Get Cited in Google AI Overviews, ChatGPT & Perplexity (2026)

Generative Engine Optimization (GEO) is getting your content cited inside AI answers: the box at the top of Google, the reply ChatGPT hands back, the sourced summary Perplexity writes. Most of what is sold as GEO is repackaged SEO, and some of it (llms.txt files, "AI schema", content-chunking) does nothing at all, on Google's own word. It sits next to Answer Engine Optimization (AEO), the same idea aimed at being the direct answer. This guide separates the tactics with real evidence behind them from the ones a cottage industry is charging for, and backs it with first-party data: our own pages are already cited in AI answers, including a live Google AI Overview that cites us for a Chinese-language query, and our GA4 shows exactly which engines send the traffic back.

Quick answer: To get cited in AI Overviews, front-load a complete answer to the question in your first ~200 words, back it with original statistics, quotes and citations, and make the page unambiguous about who wrote it and what entity it belongs to. AI Overviews and AI Mode overwhelmingly pull from pages that already rank in the top of organic search, so strong SEO is the entry ticket. You do not need llms.txt, special AI schema, or content-chunking. Google's own position is that GEO is still SEO.

What is Generative Engine Optimization (GEO)?

Generative Engine Optimization is optimizing your content so generative AI systems retrieve it, trust it, and cite it in their answers. Answer Engine Optimization (AEO) is the closely related idea of being the direct answer to a question, whether that surface is an AI Overview, a featured snippet, or a voice assistant. In practice the two overlap almost completely, and Google's stated view in its AI features and your site guidance is blunt: there is no separate playbook. The things that earn AI citations are the same things that earn rankings. GEO is a lens on SEO for a world where the search result is increasingly a synthesized answer rather than ten blue links.

The distinction that matters is the surface each one optimizes for.

ApproachWhat it optimizes forThe surface
SEORanking a page in organic resultsThe classic list of blue links
AEOBeing the single direct answer to a questionFeatured snippets, voice, People Also Ask, AI Overviews
GEOBeing retrieved and cited inside an AI-generated answerGoogle AI Overviews & AI Mode, ChatGPT, Gemini, Perplexity

The overlap is the point. A page that is clear, well-sourced and authoritative tends to rank organically, win the snippet, and get pulled into an AI answer, because all three systems are reading for the same signals. GEO is not a replacement for SEO; it is what you do on top of solid SEO to raise the odds the AI picks you as a source.

How do AI Overviews and AI assistants choose which sources to cite?

AI answers are built in two steps: retrieval, then generation. The system first retrieves a set of candidate pages relevant to the query, then writes an answer and cites the pages it leaned on. The retrieval step is where SEO does its work, because these systems draw from an index that is close to, or the same as, the organic search index. Independent analyses put the overlap high: one study of Google's AI Mode found that around 88% of the links it cites already rank in the top-20 organic results, and other research shows roughly 97% of AI Overviews cite at least one page from the top 20. Rank is not the only factor, but if you are not ranking at all, you are rarely in the candidate pool to begin with.

Diagram: how a web page becomes a cited source in an AI answer — indexed, retrieved, cited, then referral traffic in GA4

Once you are in the pool, the generation step decides who gets cited. Two things move that decision. The first is entity clarity: the system needs to understand what your page is about, who wrote it, and which organization stands behind it, without guessing. Ambiguity gets you skipped in favour of a source the model is more confident about. The second is extractability: content the model can lift a clean, self-contained answer from is easier to cite than content where the answer is buried three scrolls down and tangled in caveats. So the job splits neatly: rank well enough to be retrieved, then write clearly enough to be the piece the model quotes.

How do you get your content cited in Google AI Overviews?

There is no single trick. There is a checklist, and the pages that get cited do most of it. Here is the one we run.

Front-load the answer

Answer the question fully in the first ~200 words. If someone asks "how do I appear in AI Overviews" and your page opens with three paragraphs of preamble before it gets to the point, the model has to work to find your answer, and it will often prefer a page that stated the answer up front. Lead with a direct, complete response, then expand. The Quick answer box at the top of this article is that principle in practice.

Add original statistics, quotes and citations

This is the GEO tactic with the strongest published evidence so far. The GEO study from Princeton, the Allen Institute for AI and IIT Delhi, which tested tactics across roughly 10,000 queries, found that adding relevant statistics lifted a source's visibility in generative answers by around 41%, and that adding quotations and citations produced gains in the 30-40% range. Crucially, combining these tactics beat any single one. Original data, a named expert quote, a citation to a primary source: each gives the model a concrete, attributable thing to pull, and concrete attributable things are what AI answers are made of.

Build entity clarity and E-E-A-T

Generative engines reward content that reads as genuine expertise. Google's guidance is blunt: "Don't just recycle what others on the internet have already said, or could easily be produced by a generative AI model." It wants unique, experience-led content that says something the ten pages before it did not. That is E-E-A-T (Experience, Expertise, Authoritativeness, Trustworthiness) doing its job. Publish first-party data, real cases, and a clear point of view, under a named author with genuine credentials, on a site whose entity is unambiguous. A page that recycles the consensus gives the model no reason to cite you over the source you recycled from.

Structure for extraction

Write so the answer is easy to lift. Question-shaped H2 and H3 headings that mirror how people ask, concise definitions in the first sentence under each heading, lists where a list fits, and an FAQ section for the long tail of related questions. This is why this article is built entirely on question headings. You are not writing for a keyword; you are writing the exact answer to a question a model might be asked to summarize.

Make sure AI crawlers can reach you

None of this matters if the crawlers are blocked, though the nuance matters. For AI Overviews and Gemini inside Search, Googlebot is the gatekeeper, the same crawler that indexes you for organic Search, so if you rank, you are already reachable. The AI-specific agents matter for the assistants: for ChatGPT, GPTBot governs whether OpenAI can train on your content while OAI-SearchBot governs whether ChatGPT search can cite you, so do not block that one; PerplexityBot covers Perplexity. One thing to get right: Google-Extended controls whether your content is used to train and ground Gemini, not whether you appear in Google Search or AI Overviews, so blocking it will not remove you from AI Overviews, though it does affect how the Gemini app uses your content. Audit your robots.txt so an over-eager copy-paste is not quietly blocking Googlebot, Bingbot, GPTBot or PerplexityBot.

Structured data — optional, not required

Schema helps, but it is not a magic key. Well-formed Organization, Article and Person schema, with author, dateModified and sameAs populated, reduces ambiguity about who and what your page is, which feeds the entity-clarity signal above. But there is no special AI-Overview schema, and Google does not require structured data to cite you. Treat schema as good hygiene that removes doubt, not as the thing that gets you in.

Keep it fresh

Recency is favoured, especially for anything time-sensitive. A populated dateModified and a genuine refresh, updated statistics, current references, a re-checked claim, signals the content is maintained. Stale stats and dead references cut the other way. Refresh your cornerstone pages on a real schedule rather than publishing once and walking away.

What you do NOT need (the GEO myths)

This is where GEO gets oversold, so here is the honest de-hype, sourced to Google. In its own AI-optimization guidance, Google states plainly that there is nothing new to do specifically for its AI features. Concretely, you do not need:

  • An llms.txt file. Google does not use it. It is a proposed standard some tools promote, but Google's crawlers do not read it to decide citations, and shipping one changes nothing about your AI Overview visibility.
  • Special AI-specific schema or files. There is no AI-Overview markup, no "generative" schema type, no hidden file that flags your page to the model. Anyone selling one is selling a myth.
  • Content-chunking for the AI. You do not need to pre-slice your content into model-sized fragments. Write for humans in a clear structure; the system handles retrieval-level chunking itself.

The reason this matters commercially: a whole cottage industry is charging for llms.txt installs and "AI schema" packages that Google says do nothing. Spend that effort on the checklist above, which is evidenced, instead of on files that are not.

How do you optimize for Perplexity, ChatGPT, Gemini and other AI engines?

The fundamentals travel across every engine, be clear, be well-sourced, be extractable, but each retrieves and cites differently, and the lever changes with it. We have ordered these by what our own GA4 shows actually sends us traffic: Gemini and Perplexity together drove about three-quarters of our AI-assistant referrals, so a Hong Kong site should weight effort toward them rather than default to ChatGPT.

Diagram comparing how Perplexity, Gemini, ChatGPT, Copilot, Claude and Grok retrieve and cite sources, and the optimization lever for each

Perplexity is the most citation-driven engine, and it retrieves in two ways. It runs its own crawler, PerplexityBot, which builds a standing index and is explicitly not a training crawler, and it fetches live pages in real time when a query needs fresh information. Every answer carries prominent inline citations, so being the clear primary source of a fact, stated in clean, extractable structure, is what gets you pulled in. Recency and a genuinely quotable claim matter more here than anywhere else.

Google AI Overviews and AI Mode (Gemini) are the biggest surface, and they are grounded in Google's core Search index. If you already rank organically you are most of the way there: seoClarity's analysis of 432,000 keywords found ~97% of AI Overviews cite at least one page from the top-20 organic results. The standalone Gemini app can also ground on Google Search. So the Google GEO checklist above is the lever, strong organic presence carries straight through.

ChatGPT Search uses distinct bots for distinct jobs, and getting this wrong is the common own-goal. Per OpenAI's own docs, GPTBot governs training, while OAI-SearchBot is what determines whether your pages appear in ChatGPT search answers; ChatGPT-User handles individual live fetches. Do not block OAI-SearchBot if citations matter. ChatGPT's browsing has historically leaned on Bing, so Bing indexability still helps, alongside the same clarity and sourcing signals.

Microsoft Copilot is built on Bing's index plus GPT models, so Bing is the whole game here. Bing indexability and Bing Webmaster Tools, which now ships an AI Performance report showing your Copilot citations, are the levers. IndexNow gets fresh URLs into Bing fast.

Claude now searches the web and returns inline citations on every plan, and it shows up in our referral data. Under the hood, Claude's web search reportedly runs on Brave Search rather than Google or Bing, so broad, clean web indexability, not any single search engine, is what surfaces you. The same clarity and sourcing signals apply.

Grok (xAI) retrieves from a general web crawl plus privileged, real-time access to X posts, which gives it citation behaviour unlike the others; presence and mentions on X carry weight here. DeepSeek we could not verify a current, authoritative live-web citation mechanism for, so we are not making a claim about it.

Across all of them the winning content is the same shape: a clear, well-structured page that states an attributable answer a model can quote with confidence. You are not building six different pages; you are building one good one and making sure each engine can reach it.

How do you know if it's working? Tracking AI citations in GA4

Getting cited is only half the win; you need to see the traffic. As of mid-2026, GA4 ships a native "AI Assistant" channel. Google's default definition names ChatGPT, Gemini and Copilot among others; in our own property the channel also picks up Perplexity and Claude, because those referrals arrive with an "ai-assistant" medium. It is still imperfect: some assistants strip the referrer, so a share of genuine AI visits lands in Direct and undercounts the true total. A custom channel group with a regex is worth adding so Perplexity, Claude and anything the default misses are all captured cleanly.

Here is our own first-party proof that this is not theoretical. Our YouTube ads complete guide is already being cited in a live Google AI Overview for a Chinese-language "youtube 廣告" query, our own content picked as a source inside the answer box, and it is not the only guide of ours that AI Overviews now cite by name.

Google AI Overview citing the KickAds YouTube ads guide as a source

And the referral traffic is showing up. Here is the actual split from our GA4 for 1-24 July 2026, every session where an AI assistant sent a visitor to our site:

AI engineUsersReferral sessions
Gemini1419
Perplexity116
ChatGPT57
Claude34
All AI assistants2346

GA4 Traffic acquisition showing the AI Assistant channel with referrals from Gemini and ChatGPT

Two things stand out, and both cut against common GEO advice. Gemini led on users and sessions, and ChatGPT, where most GEO guidance points, came third by users. Perplexity shows a gap worth watching: 16 sessions from a single user, so read engine data by both users and sessions before you conclude anything. The single most-cited landing page was our YouTube ads guide, the same page Google's AI Overview cites. The volumes are modest and early, but the pattern is real and measurable: apply the checklist, and these engines send traffic to the pages they cite. Treat a page appearing in this channel as the signal that your GEO work landed.

How we write a page to get cited

This is the workflow behind the pages above, the same one we run for clients. It is deliberately plain, because the tactics that work are the ordinary ones done with discipline.

  1. Start with the real question, not the keyword. We pull search data and read the AI answers already showing for a topic, then shape the page's headings as the exact questions people and models ask. The heading is the question; the first sentence under it is the answer.
  2. Front-load a complete answer. The first 200 words have to stand on their own as the answer, because that is the block a model lifts. Everything after is depth for the reader who stays.
  3. Put original data in, not just opinion. Every flagship page carries at least one number, test result or first-party stat that exists nowhere else. That is the single thing that makes a model cite you rather than the ten pages that only restate the consensus.
  4. Verify every claim against the primary source. Before a factual line ships, it is checked against the official documentation, not memory. On a page about being trusted by AI, a wrong fact is the fastest way to lose the trust.
  5. Let the structure carry the schema. A clean question-and-answer structure with a real FAQ produces valid Article and FAQPage schema on its own, which is what makes the page eligible to be pulled into an answer. You do not bolt schema on; you write in a shape that generates it.

None of this is exotic. It is SEO done with the extra discipline that AI citation rewards, and it is what produced the citations shown above.

Common GEO mistakes to avoid

Four mistakes undo the work, and the first is the one people reach for by instinct.

  • Keyword stuffing. The same GEO study found that stuffing keywords performs worse than the baseline in generative engines. These systems read for meaning, not density, and over-optimized text reads as low-quality. Write for a person; the density takes care of itself.
  • Chasing llms.txt and fake schema. Covered above: Google says these do nothing for AI Overviews. Time spent installing them is time not spent on statistics, sourcing and structure.
  • Faking off-site mentions. Genuine brand mentions and citations across the web build the entity authority that helps you get cited. Manufacturing them, mass-produced guest posts, planted mentions, generated citations, falls under Google's scaled-content-abuse policy and risks a penalty. Earn the mentions; do not fabricate them.
  • Blocking AI crawlers by accident. The most common own-goal: a robots.txt copied from another site, or a blanket disallow, that quietly bars GPTBot, Google-Extended or PerplexityBot. Audit your robots file and confirm the crawlers you want can actually reach you.

FAQ

Which AI engines actually send referral traffic? In our own GA4, over 1-24 July 2026, the AI Assistant channel logged 46 referral sessions across four engines: Gemini (19), Perplexity (16), ChatGPT (7) and Claude (4), and by users Gemini led with ChatGPT second. Your split will differ, but do not assume ChatGPT is where AI traffic comes from, and read the numbers by both users and sessions.

Is GEO the same as SEO? Largely, yes. Google's stated position is that there is no separate optimization for its AI features, the content and technical work that earns rankings is what earns AI citations. GEO adds emphasis: front-loading answers, adding original statistics and citations, and writing for extractability. But the foundation is solid SEO, because AI Overviews and AI Mode overwhelmingly cite pages that already rank organically.

Do I need special schema or an llms.txt file for AI Overviews? No. Google explicitly says you do not need llms.txt, AI-specific schema, or content-chunking to appear in AI Overviews. Well-formed Organization, Article and Person schema is useful hygiene that reduces ambiguity about your page, but it is optional and there is no dedicated AI-Overview markup. Any tool selling an "AI schema" package as the key to AI Overviews is selling something Google says does nothing.

How long until I get cited? There is no fixed timeline, and it depends on whether you already rank. Because AI answers draw almost entirely from pages in the top organic results, a page that already ranks well can be cited quickly once it is clear and well-sourced; a page with no organic presence has to earn rankings first, which is the slower path. Treat AI citations as a downstream effect of ranking plus clarity, not an overnight switch.

Does GEO work for Hong Kong and Chinese-language content? Yes. Our own live example is a Traditional-Chinese "youtube 廣告" query where our guide is cited in the AI Overview, and our GA4 shows AI-assistant referrals landing on both English and Chinese pages. The same principles apply: answer the question clearly in the first section, source it, and keep the entity and author unambiguous. If anything, well-structured Chinese content faces less competition in AI answers than the crowded English space.


Getting cited in AI answers starts with content that already earns rankings and states a clear, sourced answer. If you want the fundamentals underneath it, our SEM ultimate guide and YouTube ads complete guide are built on exactly the front-loaded, evidence-led structure this article describes, and both are already being read by the AI engines. Want a second pair of eyes on whether your content is set up to be cited? Our free audit is a fair place to start.

Frankie Chan

About the author

Frankie Chan · Co-Founder, Kick Ads

Frankie is an ex-Googler and paid media strategist. He has managed Google Ads and Meta Ads for ecommerce and lead generation businesses across Hong Kong and Malaysia since 2017, working closely on strategy, reporting and client growth planning.

Want an expert to review your campaigns?

Get a Paid Media Health Check from Kick Ads — no pitch, just honest findings.

Get a Paid Media Health Check