Euracle
Marketing

Generative Engine Optimization: The Evidence-Based Guide

Generative engine optimization is the practice of structuring content and building presence so that AI systems retrieve it, cite it, and name your brand when answering a question. It differs from search engine optimization in its unit of success: how often you are mentioned across many generated answers, rather than where you rank on one results page.

TL;DR

  • GEO was defined in a peer-reviewed paper, GEO: Generative Engine Optimization, presented at KDD 2024. It introduced GEO-bench, a benchmark of 10,000 queries, and tested nine content strategies.
  • The strongest methods were citing sources, adding quotations and adding statistics. Improving fluency produced meaningful gains with no new content. Keyword stuffing performed below the unmodified baseline.
  • The headline "up to 40%" is a maximum, not an average. The top three methods produced roughly 30% to 40% relative improvement on the paper's own visibility metric.
  • There is more research than the industry cites. A NeurIPS 2025 benchmark titled C-SEO Bench: Does Conversational SEO Work? exists and is essentially never mentioned in agency guides.
  • The research tests on-page changes. The strongest reported signal is off-page. Analyses find AI engines favour authoritative third-party sources over brand-owned content, which means the best-evidenced discipline has its largest gap exactly where most of the work is.

If the terminology is what you came for rather than the practice, our guide to AEO vs GEO vs SEO settles which word to use and where each came from.

What is generative engine optimization?

Generative engine optimization is the practice of structuring content and building presence so AI systems retrieve it, cite it, and name your brand when answering a question. The term was coined in an academic paper presented at KDD 2024, which defined a generative engine as a system combining a language model with a search engine.

Three plain-English terms that make the rest readable:

  • A generative engine is any system that writes an original answer by pulling from sources it retrieves. ChatGPT, Perplexity, Google's AI features and Gemini all qualify.
  • Retrieval is the step where the system goes and fetches material before writing. What gets retrieved decides what can be cited.
  • A citation is the system naming or linking your page inside its answer. It is the thing GEO is trying to win.

The important structural point, which the next sections build on: the unit is a passage, not a page, and the outcome is a mention, not a position.

How is GEO different from SEO?

In what counts as winning. SEO competes for a position in a list of links. GEO competes for a mention inside a written answer.

SEOGEO
What you winA ranked linkA citation or brand mention
How success is measuredPosition and clicksFrequency across many prompts
Is the result stable?Broadly, between updatesNo. The same prompt can produce different answers
What gets retrievedA pageA passage
Where the work happensMostly your own siteSubstantially off your own site

The tactics overlap heavily and the foundations are identical: be crawlable, answer clearly, be worth citing. The divergence is that last row, and it is the part most programmes underfund. Our guide to answer engine optimization covers how that split should shape a budget.

What does the research actually say works?

Four content methods, measured against a 10,000-query benchmark, and one that measurably backfires.

The founding paper introduced GEO-bench, a benchmark of 10,000 queries across 25 domains, then tested nine ways of modifying content and measured the effect on visibility inside generated answers using a metric accounting for both how much of your content appears and where it appears.

Up to 40%: the visibility improvement reported by the founding GEO paper. A maximum from the strongest methods, not an average (Aggarwal et al., KDD 2024)

Reported effect
Cite SourcesAmong the strongest, consistent across domains
Quotation AdditionAmong the strongest
Statistics AdditionAmong the strongest. Replacing vague claims with figures
Fluency OptimizationMeaningful gains from clearer writing alone, with no new content added
Keyword StuffingRoughly 8% below the unmodified baseline, and about 10% below in the paper's Perplexity validation

Two things follow that most guides do not draw out.

The winning methods are not new. Citing your sources, quoting authorities, using real numbers and writing clearly are what good writing already required. GEO did not invent a technique; it produced evidence for practices that were previously matters of taste.

Effectiveness varied by domain. The paper found the methods performed differently across subject areas, which is why a universal checklist is a poor substitute for testing on your own material.

What has not been measured?

Most of it, and being clear about that is what separates a guide from a pitch.

There is more research than the industry cites. Beyond the founding paper, the literature includes a NeurIPS 2025 benchmark titled C-SEO Bench: Does Conversational SEO Work? (Puerto, Gubri, Green, Oh and Yun, arXiv:2506.11097), a March 2026 preprint on diagnosing and repairing citation failures in generative engine optimization, and work on how language models internalise citation practices. A benchmark asking, in its title, whether these techniques work is not something a serious guide should omit.

The research tests on-page changes. The founding paper measured what happens when you modify a page. It did not measure what happens when a review site, a directory or a community discussion mentions you, which is the mechanism most practitioners now believe matters most.

And the strongest reported signal sits in that gap. Analyses combining the founding paper with 2025 work on citation bias find that AI engines favour authoritative third-party sources over brand-owned content. So the best-evidenced discipline in marketing has its largest evidential hole exactly where most of the work is.

What this means for you: treat the four on-page methods above as evidence-backed and everything else in this guide, including the off-page section, as reasoned practice supported by observation rather than controlled experiment. Any vendor presenting off-page GEO tactics as proven is overstating what exists.

Why is there no position one in GEO?

Because generative engines are non-deterministic, meaning the same question asked twice can produce different answers with different sources. There is no fixed slot to occupy.

That single property changes what you are optimising for. In search you pursue a position. In GEO you pursue a frequency: across a set of prompts a buyer might realistically type, how often does your brand appear?

Three practical consequences.

Measure across many prompts, not one. A single check tells you nothing. A fixed set of 20 to 50 prompts, run repeatedly, tells you something.

Expect movement without cause. Your citation can appear and disappear without you changing anything, because model updates, reasoning settings and retrieval choices all shift.

Compare against competitors, not against yourself. Because the absolute number is noisy, the useful figure is your share of mentions relative to the other brands appearing for the same prompts.

One check, one answer, is not a measurement. It is an anecdote, and it is how most teams currently form beliefs about their AI visibility.

How do you do GEO on your own site?

Six steps, ordered so the evidence-backed work comes first.

Step 1. Be retrievable at all

A page that cannot be crawled cannot be cited. Check for noindex and nosnippet, confirm your CDN is not blocking legitimate crawlers, and confirm pages you want cited are actually indexed. This is the precondition for everything below.

Step 2. Cite your sources

The strongest measured method. Attribute claims to named publishers with dates, inline, where the claim appears. It also makes the passage safer for a model to reuse, because the attribution travels with it.

Step 3. Add statistics

Replace vague claims with figures. "Improves efficiency" becomes "cut triage time from four minutes to one across 2,000 tickets a month". Measured as among the strongest methods, and it is the single easiest edit to make on existing pages.

Step 4. Add quotations

Quoting a named authority was measured among the strongest methods. One relevant, attributed quotation per substantial section is sufficient.

Step 5. Improve fluency

The paper found meaningful gains from clearer writing with no new content added. Short sentences, one idea per paragraph, answer before elaboration.

Step 6. Write passages that survive being quoted alone

Assume any 200 to 300 word block will be retrieved without its surroundings. Name the subject in the first sentence rather than saying "it". Put the number and the source inside the block. Do not rely on a previous paragraph for meaning.

What not to do: keyword stuffing, which measured below the unmodified baseline. And llms.txt, which Google states its systems do not use, covered in our guide to llms.txt.

How do you do GEO off your own site?

By being present on the sources these systems retrieve from, which are frequently not your website. This is the larger half of the work and the one with less evidence behind it, so it is presented as reasoned practice rather than proven method.

Find the sources first. Run your buyers' real questions through ChatGPT, Perplexity and Gemini and record which domains get cited. You are looking for the review sites, directories, comparison pages and communities that recur in your category. That list is your target list, and it will differ by industry.

Then audit your presence on each. Not your website. Your listing, your profile, your third-party coverage. Most on-page-only programmes stall precisely here.

Four categories, in rough order of how often they appear:

What to do
Review and directory sitesClaim the listing, complete every field, keep it current
Comparison and listicle pagesThird-party pages naming your category. Being absent from them is being absent from the answer
Communities and forumsParticipate genuinely. Manufactured presence is detectable and reputationally expensive
Your own earned coverageOriginal research, data and commentary that other people cite. The only category you fully control

Entity consistency matters across all four. Your company name, description, founding year and service list should match exactly everywhere. Inconsistency degrades a model's ability to resolve that all those mentions refer to one company.

The assistant-specific version of this work is set out in how to get cited by ChatGPT.

How do you measure generative engine optimization?

Three sources, in descending order of reliability, and one calculation.

Google's first-party data. Search Console's Generative AI performance report shows how content performs in Google's generative AI features. It is the only first-party AI visibility data any platform publishes.

Your own prompt monitoring. Fix 20 to 50 buyer questions, run them monthly across the assistants your buyers use, record every brand and domain cited. Run each at more than one reasoning setting where the product allows it, because source sets diverge between modes. This is a measurement you build, not a metric you buy, and it is auditable.

Referral traffic in GA4. Filter for assistant domains. Note that a large share of assistant traffic arrives without a referrer and lands as direct, so this understates the channel.

The calculation: share of mentions

Run your prompt set. For each prompt, record which brands appear. Then:

Share of mentions = (prompts where your brand appears) ÷ (total prompts in the set)

Worked example, illustrative:

Value
Prompts in set40
Prompts mentioning your brand6
Your share of mentions15%
Prompts mentioning your closest competitor14
Their share35%

That gap is the number to move, and it is meaningful in a way that raw citation counts are not. Re-run monthly and track the direction. A single reading is noise; three readings are a trend.

<a id="cost-and-time"></a>

What does GEO cost and how long does it take?

The on-page work is cheap and fast. The off-page work is neither.

On-page workOff-page workMonitoring
What it involvesApplying the four measured methods to existing pagesListings, third-party coverage, community presence, original researchRunning a fixed prompt set
Effort30 to 45 minutes per pageOngoing, distributed across the year2 to 4 hours monthly, less if automated
Time to any effectWeeks, since pages are already indexedMonthsImmediate, it is a measurement
Evidence behind itControlled experimentsObservation and inferencen/a
Where GEO is the weaker investment
Here, if your buyers do not use assistants. AI referrals average around 1% of sessions for most sites. If your category's buyers are not asking assistants about it, on-page hygiene is worth the hours and a funded off-page programme is not. Check your own prompt set before committing budget

The honest sequence: do the on-page work first because it is cheap, fast and evidence-backed. Build a prompt set and measure for two months. Only fund the off-page programme if the measurement shows your category is genuinely being asked about.

What do most teams get wrong?

Buying tactics with no evidence behind them. The four measured methods are public. Anything beyond them is reasoned practice, and a vendor should say which is which.

Quoting the 40% figure as an average. It is a maximum from a specific benchmark on a specific metric. Repeating it as a general expectation turns a research finding into a sales claim.

Checking once and forming a belief. Generative engines are non-deterministic. One prompt, one answer, is an anecdote.

Treating GEO as a one-time optimisation. Source sets shift with model updates. A brand dominating a prompt this quarter can be absent next quarter with no on-site change.

Funding the on-page half and calling it a programme. The on-page work is the part with evidence and the smaller part of the job.

Manufacturing third-party presence. Fake reviews and astroturfed community posts are detectable, and the reputational cost when found exceeds any citation gained.

Who pays for this: the marketing lead who funded it, because the channel that grows is brand and enquiry quality while the channel that shrinks is measurable sessions. Agree the success metric before the work starts.

How does Euracle run generative engine optimization?

Evidence first, then measurement, then the expensive half only if the measurement justifies it.

The Eureka Method, Euracle's discovery sprint, runs four phases.

Discover builds the prompt set from the client's real sales questions and establishes a baseline share of mentions before anything is changed, so there is something to measure against later.

Design separates the evidence-backed on-page work from the reasoned off-page work in writing, and says which is which.

Deploy ships the four measured methods across existing pages first, because that is the cheapest work with the strongest support.

Scale re-runs the prompt set monthly and only funds off-page investment where the baseline shows the category is genuinely being asked about.

The stack is deliberately ordinary: Ahrefs and SEMrush for search data, Google Search Console and GA4 for first-party measurement, and n8n with the Claude API to run the monthly prompt set automatically, so recurring measurement is a scheduled job rather than a recurring analyst cost.

Two structural commitments come from how Euracle is set up. Senior practitioners only: the people in the pitch do the work, which matters most in the judgement call about which tactics are evidenced and which are not. And one contract across six disciplines, so a finding that the client's category is not yet being asked about, and the budget belongs elsewhere, does not require a second vendor to act on.

If you want the baseline measured before anyone quotes you a programme, that is Euracle's AEO and GEO service. The gap between on-page and off-page effort is widest for B2B SaaS companies, where buyers research through comparison sites and communities the brand does not control.

The Google-specific half of this work sits in how to rank in AI Overviews, which covers a surface with published documentation and first-party reporting.

FAQ

FAQ

Generative engine optimization is the practice of structuring content and building presence so AI systems retrieve it, cite it and name your brand when answering a question. The term was coined in a paper presented at KDD 2024, which defined a generative engine as a system combining a language model with a search engine.

Four content methods have controlled evidence behind them: citing sources, adding quotations, adding statistics and improving fluency. The founding paper measured these on a 10,000-query benchmark. Most other GEO tactics circulating are reasoned practice rather than measured method, and a benchmark titled C-SEO Bench: Does Conversational SEO Work? exists and deserves reading.

SEO competes for a ranked position in a list of links. GEO competes for a mention inside a written answer, measured as frequency across many prompts rather than as a position. The foundations overlap almost entirely. The real divergence is that GEO depends substantially on third-party sources you do not own.

Fix a set of 20 to 50 buyer questions, run them monthly across the assistants your buyers use, and calculate share of mentions: prompts where your brand appears divided by total prompts. Add Google Search Console's Generative AI performance report for Google's surfaces, and GA4 referral data with the caveat that it understates the channel.

On-page changes to indexed pages can show effect within weeks. Off-page work runs on months, typically one to three for directory and review-site corrections and longer for genuine community presence. Judge a programme at six months, and expect the measurement to move before the traffic does.

The on-page work is roughly 30 to 45 minutes per page, so a 40-page retrofit is about 25 internal hours. Monitoring is two to four hours monthly, less if automated. Off-page work is the expensive part and is open-ended, which is why measuring first is worth doing before committing to it.

No. It measured roughly 8% below the unmodified baseline in the founding GEO paper, and about 10% below in the paper's Perplexity validation. Adding keywords to content performed worse than doing nothing, which is the opposite of what the same tactic once achieved in early search engines.

They overlap heavily and the boundary is contested. GEO covers how your brand is represented across generated answers generally. AEO focuses on being extracted as a direct answer. GEO is the only one of the two with a formal published definition, so it is the safer term to use precisely.

Conclusion

You can now separate what is known from what is asserted. Apply the four measured methods to your existing pages, because they are cheap, fast and the only part of this discipline with controlled experiments behind them. Build a prompt set of your buyers' real questions, calculate your share of mentions, and re-run it monthly, because one check is an anecdote. Then fund the off-page work only if the measurement shows your category is genuinely being asked about, and treat anyone presenting off-page tactics as proven with appropriate scepticism. If you want the baseline measured before anyone quotes you a programme, talk to Euracle about AEO and GEO.

Sources

  1. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., and Deshpande, A. GEO: Generative Engine Optimization. arXiv:2311.09735, November 2023; KDD '24, Barcelona, August 2024. DOI 10.1145/3637528.3671900. https://arxiv.org/abs/2311.09735
  2. Puerto, H., Gubri, M., Green, T., Oh, S. J., and Yun, S. C-SEO Bench: Does Conversational SEO Work? NeurIPS Datasets and Benchmarks 2025. arXiv:2506.11097.
  3. Diagnosing and Repairing Citation Failures in Generative Engine Optimization, preprint, 11 March 2026. arXiv:2603.09296.
  4. Google Search Central, Optimizing your website for generative AI features on Google Search, last updated 10 July 2026. https://developers.google.com/search/docs/fundamentals/ai-optimization-guide
  5. Google Search Central, AI Features and Your Website, on the Generative AI performance report. https://developers.google.com/search/docs/appearance/ai-features
  6. Method-level results from the founding paper, and the finding that AI engines favour authoritative third-party sources over brand-owned content.
Image of Devanshu

Written by

AI-Native Product Manager + GTM Engineer

Keep reading

More from the journal.

Marketing

How Long Does SEO Take? The Honest Answer Is Three Numbers

SEO takes 30 to 60 days to tell you whether it is working, 3 to 12 months to produce rankings, and longer than that to produce revenue. Those are three separate clocks and most answers to this question report only the middle one, which is why the answer never feels useful.

Image of Devanshu
Devanshu Takkar
AI-Native Product Manager + GTM Engineer
Sep 3, 2026
Read article
Marketing

AEO vs GEO vs SEO: Three Names, One Job

You do not need three programmes. SEO, AEO and GEO describe one discipline at three scopes: ranking a page, being quoted as an answer, and being represented across AI-generated responses. The work overlaps heavily. If you have been quoted three separate retainers, that is a pricing decision rather than a technical one.

Image of Devanshu
Devanshu Takkar
AI-Native Product Manager + GTM Engineer
Sep 2, 2026
Read article
Marketing

llms.txt: You Are Probably Being Sold the Wrong Use Case

llms.txt is a plain-text file placed at the root of a website that lists, in Markdown, the pages an AI system should read and what each one covers. It was proposed in September 2024 by Jeremy Howard of Answer.AI as a way to give language models a clean, low-noise map of a site's documentation.

Image of Devanshu
Devanshu Takkar
AI-Native Product Manager + GTM Engineer
Aug 28, 2026
Read article

The breakthrough, delivered

Your breakthrough is one conversation away.

Let's find your spark