# Does GEO Work? What 45 Studies Actually Prove
> Does GEO really work? A critical review of 45 studies: proven levers, weak claims, and a rigorous method for measuring visibility in AI search.
- Canonical HTML: https://verityscore.io/en/blog/does-geo-work-45-studies-evidence-2026/
- Markdown alternate: https://verityscore.io/en/blog/does-geo-work-45-studies-evidence-2026.md
- Language: en
- Content type: blog
- Published: 2026-08-24
- Updated: 2026-08-24
- Tags: geo, ai-search-optimization, chatgpt, google-ai-mode, ai-citations, geo-measurement, geo-research, aeo
Does GEO work, or has the industry simply given a new name to familiar content and search practices? It is a fair question. One set of studies reports visibility gains of **up to 40%**. More recent experiments find that most proposed tactics have little effect and sometimes make performance worse.

Those findings are less contradictory than they look. They measure different stages.

A page can be accessible to a crawler without being retrieved for an answer. It can be retrieved without being cited. It can be cited without receiving a click. A click can still produce no revenue. As long as all four outcomes are collapsed into the word “visibility,” almost any GEO promise can be made to sound true.

We reviewed the [critical survey of 45 studies published between 2023 and July 2026](https://arxiv.org/abs/2607.14035), then returned to the most important primary papers to check their methods, evidence, and publication status. This is what the research supports, and what it does not yet justify selling as certainty.

## The short answer: yes, but not as advertised

**Yes, GEO describes a real problem.** Generative engines select, order, and synthesize sources. Changes to a page can alter the probability that its information appears in an answer.

**No, there is no universal recipe for “ranking in ChatGPT.”** None of the 45 studies covered by the critical survey establishes a stable, longitudinal, cross-platform causal effect running all the way from a page edit to organic discovery, traffic, and revenue.

The strongest evidence concerns intermediate outcomes:

- how closely a source matches the question;
- whether it enters the retrieved document set;
- where it appears inside the model's context;
- whether important facts are explicit, current, and extractable;
- whether the measurement is repeated enough to separate signal from run-to-run variation.

The evidence becomes thinner as the outcome moves toward commercial impact: click, conversion, and incremental revenue.

That distinction separates what is **strictly demonstrated** — an intervention changed citation behavior under a defined protocol — from a **strategic interpretation** — the same intervention may improve acquisition if every other stage also works.

## GEO is not one ranking: it is a chain of probabilities

A traditional search result still looks like an ordered list. A generative answer has several systems between your page and the final sentence.

| Stage | The actual question | Useful measurement |
|---|---|---|
| 1. Access | Can the crawler read the page? | HTTP status, robots.txt, served HTML |
| 2. Discovery | Is the page known and eligible for indexation? | index coverage, links, sitemap |
| 3. Retrieval | Is it selected for this particular question? | inclusion in candidate documents |
| 4. Reranking | Where is it placed in the context? | source position and relevance score |
| 5. Citation | Is the URL shown as a source? | citation frequency and citation rank |
| 6. Absorption | Which facts enter the generated answer? | claims used and claim fidelity |
| 7. Click | Does the user open the source? | referral sessions and landing pages |
| 8. Conversion | Does the visit create value? | purchase, margin, lead, assisted revenue |

[Generative Engine Optimization](/en/blog/what-is-geo/) now refers to this entire chain. That is useful for naming the field, but dangerous when interpreting a study.

A citation gain at stage 5 does not prove a discovery gain at stage 3. It certainly does not prove a sales increase at stage 8.

## Where the “40% more GEO visibility” claim comes from

The industry's most repeated number comes from the [foundational GEO paper published at KDD 2024](https://arxiv.org/abs/2311.09735). The researchers created GEO-bench, a set of 10,000 queries, and tested changes including citations, statistics, direct quotations, style improvement, simplification, and keyword repetition.

For some query categories and metrics, the strongest methods produced visibility gains approaching 40%. The paper also ran a validation on Perplexity. Quotations, statistics, and external citations sometimes improved how prominently a document appeared in an answer. Keyword stuffing made one measure worse.

The finding matters: **how information is expressed can affect whether a model uses it**.

But the protocol has a major limit. The documents under evaluation were already present in the supplied set or generation context. The study did not test whether an unknown page would be discovered on the open web, retrieved against millions of alternatives, and then visited.

The accurate statement is:

> In a controlled setting where a source was already available to the engine, some transformations increased its measured visibility by up to roughly 40%.

“GEO increases your traffic by 40%” is not a conclusion of that study.

## What later experiments corrected

Research since 2024 has moved closer to real retrieval conditions. It makes three important corrections to the original narrative.

### 1. A tactic that works alone can lose its effect when competitors copy it

[C-SEO Bench, published in the NeurIPS 2025 Datasets and Benchmarks track](https://proceedings.neurips.cc/paper_files/paper/2025/hash/27aa3aeff0f8460a7b43d30fa6c5c032-Abstract-Datasets_and_Benchmarks_Track.html), evaluates 1,900 questions, 16,360 documents, six domains, and both factual answers and product recommendations.

Its results are far less dramatic. Most conversational optimization methods were largely ineffective and sometimes harmful. In the matrix reproduced by the critical survey, only **3 of 54 cases** showed a statistically significant positive result. Approaches that improved a source's position inside the model context worked better than cosmetic rewrites.

The benefit also declined as more competing actors adopted the same method. This is intuitive. If every page adds a statistic, an expert quote, and three question headings, those elements stop differentiating any one source. GEO is competitive, not absolute.

### 2. Optimizing for citation can hurt retrieval

[SAGEO Arena, accepted at KDD 2026](https://arxiv.org/abs/2602.12187), evaluates a more complete chain: retrieval, reranking, and generation over a large corpus with structured web information.

Its most useful finding for practitioners is uncomfortable. Several rewrites that appeared favorable at generation time degraded retrieval or reranking. A passage can become easier to quote after it is supplied while becoming less likely to enter the candidate set in the first place.

This explains part of the contradiction in public GEO testing. A consultant manually gives a page to an assistant and sees a better answer. The live engine must first find that page among every alternative. A change wins only if it survives both stages.

### 3. Relevance and context position outweigh cosmetic formatting

[What Gets Cited?, published at SIGIR 2026](https://arxiv.org/abs/2605.25517), runs 252,000 controlled trials across six models and 18 factors. Two documents are placed in context and the researchers observe which source receives the first citation.

The largest effects come from **topical relevance** and **list position**. Explicit prices and recent dates also help consistently. Completeness and some trust cues have smaller effects. Formatting-only changes contribute little.

The limitation remains visible: both sources have already been retrieved. The paper shows what separates available candidates, not how a page enters the initial selection.

## Evidence table: demonstrated, promising, or unproven

| Claim | Evidence as of August 2026 | Practical meaning |
|---|---|---|
| Exact relevance increases the chance of citation | **Strong but conditional** | Answer a narrow intent instead of merely lengthening a generic page |
| A source placed earlier in context is cited more often | **Strong but conditional** | Retrieval and upstream authority still matter |
| Explicit price, recency, and completeness can help | **Repeated in controlled settings** | State important facts instead of forcing the model to infer them |
| Citations, statistics, or quotations always improve visibility | **Varies by domain and engine** | Add only relevant and verifiable evidence |
| Clear structure improves extraction | **Plausible and supported across studies** | Organize information, but do not treat formatting as a hack |
| FAQ format triggers AI citations | **Not established as a universal effect** | Use FAQs for readers, not as a citation guarantee |
| Schema markup is sufficient for recommendation | **Not established** | Keep structured data, but expose the same facts visibly and consistently |
| llms.txt improves Google AI Mode visibility | **Contradicted by Google** | Google Search says it ignores the file |
| One run accurately represents AI visibility | **False** | Repeat prompts and measure a distribution |
| More AI referrals prove an optimization worked | **False without a comparison group** | Separate channel growth from incremental impact |
| GEO guarantees durable traffic or sales growth | **Not established** | Measure through conversion before making the claim |

The conclusion is not that nothing works. It is that **the strength of a commercial promise should match the strength of the evidence**.

## Why one ChatGPT screenshot proves nothing

Generative answers vary with phrasing, word order, model, date, location, and sometimes between identical runs.

The [Don't Measure Once study](https://arxiv.org/abs/2604.07585) tracked four engines, four verticals, and 45 days. It found that a single observation can give a misleading picture of visibility. Within the setting studied, the authors proposed roughly **seven to eight repetitions** to stabilize some estimates.

That number is not a universal law. The study covers a limited universe, including a Swiss market, and tests at most ten repetitions. The broader method is still persuasive: generative visibility is a distribution, not a fixed rank.

A credible measurement should preserve:

- the exact question and its paraphrases;
- the engine and, where visible, the model;
- the language, country, and browsing context;
- the date and time;
- the complete answer, not just the brand mention;
- every cited URL and its order;
- the claims attributed to each source.

[Share of AI Voice](/en/blog/share-of-ai-voice-guide-2026/) becomes meaningful only when it aggregates a stable query panel and repeated observations. A single score creates false precision.

## A citation does not mean the source was understood correctly

Citation counting has another flaw: it measures whether a URL appears, not whether the answer used it faithfully.

A [human evaluation published at Findings of EMNLP 2023](https://aclanthology.org/2023.findings-emnlp.467/) examined four answer engines available at the time. Only 51.5% of generated statements were fully supported by citations, while 74.5% of citations supported the statement associated with them.

Those systems are from 2023, so the percentages should not be treated as a benchmark for 2026 products. They establish an important measurement principle that still applies: **citation and faithfulness are different outcomes**.

For a brand, at least three cases matter:

1. the page is cited and the claim is faithful;
2. the page is cited but the engine distorts the price, availability, or return policy;
3. the brand is mentioned through another source without citing its own site.

The first creates an asset. The second can damage trust. The third shows why brand visibility extends beyond the brand domain: comparison sites, publishers, marketplaces, forums, and review platforms can all shape the answer.

## Traffic data proves the market exists, not that GEO caused the growth

It would be equally wrong to conclude that GEO has no commercial importance.

Google launched [AI Overviews and AI Mode in France on July 22, 2026](https://blog.google/intl/fr-fr/nouveautes-produits/explorez-obtenez-des-reponses/recherche-ia-apercus-mode/). Google says AI Mode questions are three times longer on average than traditional searches. Its query fan-out mechanism breaks complex requests into multiple parallel searches, so one purchase question can create several opportunities to select a source.

Commercial signals are rising outside Google as well:

- Adobe measured 393% year-over-year growth in AI-referred traffic to US retail sites in Q1 2026. In March, those visits converted 42% better than non-AI traffic and generated 37% more revenue per visit within its sample.
- Valiuz observed a twentyfold increase in generative-AI traffic across 13 large French retailers between May 2025 and May 2026.
- Ahrefs associated the presence of an AI Overview with a 58% lower average click-through rate for the first organic result across 300,000 keywords, using December 2025 data.

These studies do not answer “did a GEO intervention cause more sales?” Adobe and Valiuz describe the growth and quality of a channel. Ahrefs measures an association rather than a randomized experiment. They establish that the search surface is changing and that its visitors can have value. They do not validate an optimization checklist.

## The two studies closest to commercial impact

Two publications go beyond citation experiments, with different limitations.

### Pinterest: a positive production signal at large scale

A Pinterest team describes a 2026 production system that uses visual models to generate intent queries and collection pages connected across billions of images and tens of millions of collections. The paper reports a 20% organic traffic increase in the deployment studied and a large increase in generative-engine traffic compared with a nearest-neighbor retrieval baseline.

This may be the most substantial positive production case published so far. It is still one organization's industry preprint. The paper does not disclose every group size, assignment unit, and uncertainty interval needed for broad generalization. It also appears to combine traditional organic and generative acquisition.

The defensible interpretation is: **pages built around observed intent, proprietary data, and coherent internal linking can create material value at scale**. It does not follow that mass-generating pages will produce the same result on a site with no distinctive data.

### Glasp: why the control group changes the conclusion

A June 2026 natural experiment studies hundreds of thousands of YouTube-related question-and-answer pages on one domain. Treated pages recorded 5.7 times more ChatGPT referrals, compared with 3.5 times for untreated pages on the same site. An interrupted time-series analysis estimated a 1.82-fold effect with a 95% confidence interval of 1.31 to 2.54.

However, a placebo-in-time test returned p=.16, weakening the causal conclusion. The pre-treatment period was short and noisy.

The result is encouraging, not definitive. Its most important contribution is methodological. Without untreated pages, the entire 5.7-fold rise could have been attributed to the intervention. The comparison group shows that much of it came from ChatGPT's overall growth as a referral channel.

## What Google officially says, and what it does not say

Google's [official guide for performing well in AI search experiences](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) resolves several recurring debates:

- AI features in Search use Google's core ranking and quality systems;
- no special technical requirement exists beyond normal index and snippet eligibility;
- no special markup, AI file, or artificial content chunking is required;
- Google Search ignores `llms.txt`;
- structured data remains useful when it matches visible content, but it is not a generative shortcut;
- unique, helpful, non-commodity content remains the best investment.

For Google, this confirms that GEO does not replace [SEO](/en/kb/geo-vs-seo/). It does not disclose the exact behavior of ChatGPT, Claude, or Perplexity. Every engine has different sources, crawlers, retrieval systems, and interfaces.

The strategic conclusion is more precise than either “GEO is just SEO” or “SEO is dead”: **discovery remains the entry condition, while selection, absorption, and synthesis add new stages to optimize and measure**.

## For ecommerce, the best GEO lever is not a format

A shopping question forces the engine to compare facts: product, variant, price, stock, purpose, material, compatibility, delivery, returns, and customer proof. When a fact is missing, the engine must ignore it, find it somewhere else, or infer it.

Our [measurement of 475 French Shopify stores](/en/blog/geo-barometer-shopify/) shows where the real gap sits. The discovery layer is close to universal: 97.9% of stores already serve an `llms.txt`, largely generated by Shopify. On 457 product pages, however, only 10.9% exposed a GTIN, 6.1% attached an `AggregateRating`, and 4.6% published structured shipping or return information.

The file marketed as a new GEO tactic is already supplied by the platform. The commercial facts needed for comparison remain scarce.

An ecommerce store should prioritize one **consistent product source of truth**:

1. state price, availability, variants, and essential specifications in served HTML;
2. make those facts match the `Product` JSON-LD;
3. keep the same values in product feeds and Merchant Center where applicable;
4. expose shipping and return terms instead of hiding them behind a dynamic interface;
5. connect reviews to the right product and make their evidence accessible;
6. publish comparison and decision pages built on evidence competitors cannot simply paraphrase.

None of this guarantees a citation. It reduces ambiguity throughout the chain: entity understanding, retrieval, comparison, and answer fidelity. That is a mechanism, not a hack.

## A GEO measurement protocol derived from the evidence

A brand does not need a research laboratory to test GEO. It does need enough discipline to avoid mistaking noise for impact.

### 1. Build the panel around real decisions

Separate questions by intent: category discovery, comparison, objection, compatibility, price, use case, and final choice. Fifty questions close to purchase are more informative than 500 generic prompts containing the brand name.

### 2. Record a repeatable baseline

Run several paraphrases of each question and repeat the observations. Seven or eight runs are a benchmark from one study, not a universal standard. What matters is keeping the protocol stable before and after the change.

### 3. Change one family of factors at a time

If you rewrite the page, add schema markup, repair a feed, and launch a press campaign simultaneously, you cannot attribute the outcome. Work with page or product cohorts and preserve an unchanged comparison group where possible.

### 4. Measure each stage independently

| Level | Indicator |
|---|---|
| Retrieval | share of answers in which the domain appears among consulted sources |
| Citation | share of answers displaying the URL and average citation position |
| Presence | share of answers mentioning the brand or product, with or without a direct citation |
| Fidelity | share of price, use, and policy claims repeated accurately |
| Acquisition | assistant-attributed visits, landing pages, and engagement |
| Value | conversion, margin, new customers, and assisted revenue |

The [AI traffic measurement guide](/en/blog/measure-ai-traffic-ecommerce/) and [GEO ROI framework](/en/blog/measure-geo-roi-2026/) extend the method into analytics and commercial outcomes.

### 5. Compare distributions, not screenshots

Report the median, dispersion, and number of observations. A change from 20% to 28% presence across five responses is not equivalent to the same movement across 500 observations over six weeks.

### 6. Allow the right delay for each stage

A served-HTML fix can appear immediately in a direct fetch but take longer to enter an index. Referral visits may remain rare even after citations rise. Set the observation window in advance rather than selecting the date that tells the strongest story.

## Seven mistakes that make a GEO test unusable

1. **Testing only a branded query.** It mainly measures whether the engine already knows the entity.
2. **Using one answer.** The result may be run-level randomness.
3. **Counting only citations.** The source may be misrepresented or receive no click.
4. **Treating two engines as interchangeable.** Their indexes, models, and interfaces differ.
5. **Changing the entire site at once.** No effect can then be attributed.
6. **Ignoring channel-wide growth.** More ChatGPT referrals do not prove that your intervention worked.
7. **Turning correlation into a rule.** Frequently cited pages often have more mentions and proof, but those signals can be consequences of authority as much as causes.

A serious [GEO audit](/en/kb/geo-audit/) should name the stage observed, the source of the data, and the limit of the inference. Otherwise, it produces a legible score without a reliable decision.

## What to do now

The 2023–2026 literature supports neither cynicism nor magic.

Cynicism ignores a real change: generative engines select and synthesize sources for questions that increasingly resemble purchase decisions. Magical claims ignore the limits: effects depend on retrieval, engine, domain, competitors, and measurement design.

The most defensible strategy has four priorities:

1. **Be discoverable.** Fix access, indexation, internal linking, and actual crawler blocks.
2. **Be selectable.** Address a precise intent with unique, current, verifiable facts.
3. **Be interpretable.** Align visible content, structured data, product feeds, and external evidence.
4. **Be measurable.** Track retrieval, citation, fidelity, clicks, and conversion with repetition and a comparison group.

GEO therefore works less like a new tag to install and more like a discipline of consistency. The goal is not to write for an AI. It is to reduce, at every stage, the amount the system must guess about your brand and products.

## Method and limitations of this analysis

This synthesis begins with the critical survey of 45 studies published on July 15, 2026, then returns to the most important primary publications. Scientific status matters: the foundational GEO study appeared at KDD 2024, C-SEO Bench at NeurIPS 2025, What Gets Cited? at SIGIR 2026, and SAGEO Arena at KDD 2026. Other cited studies remain preprints and should be interpreted accordingly.

The Adobe, Valiuz, and Ahrefs figures describe market change, not the causal effectiveness of a GEO method. The Pinterest and Glasp cases move closer to real traffic but do not yet support a universal rule. This analysis covers evidence available through August 24, 2026. The field will change as engines, interfaces, and longitudinal research evolve.

## Key takeaways

- GEO has measurable effects, particularly after a source has been retrieved.
- “Up to 40%” means neither 40% more traffic nor 40% more sales.
- Relevance, retrieval, and context position outweigh formatting tricks.
- Citation-focused optimization can hurt retrieval when treated in isolation.
- One screenshot does not measure visibility; repeated observations are required.
- Being cited guarantees neither a faithful claim, a click, nor a conversion.
- For ecommerce, consistent product facts are a more defensible lever than a file or tag sold as magic.

The right test is not “does ChatGPT cite me today?” It is: **across a stable set of purchase decisions, is my brand retrieved, accurately cited, visited, and chosen more often than before, and more often than an unchanged comparison group?**
## FAQ

### Does GEO really work?

Yes, some interventions change the probability that a generative engine uses or cites a source, particularly after that source has already been retrieved. However, none of the 45 studies in the critical review establishes a universal method that sustainably increases organic discovery, traffic, and sales across platforms. Results depend on the engine, query, competitors, and stage being measured.

### Does the original GEO paper prove a 40% gain?

The paper reports visibility gains approaching 40% in parts of its protocol, but candidate documents were already supplied to the engine. It shows that rewriting can affect a response once a page is in context. It does not show that the page will be crawled, retrieved from the open web, visited, or converted more often.

### Which GEO factors have the strongest evidence in 2026?

The most repeatable results concern exact topical relevance, inclusion in the retrieved source set, position within the context, and explicit facts such as price and recency. Clear structure can make facts easier to extract, but neither FAQ formatting nor schema markup guarantees a citation on its own.

### Do I need an llms.txt file for Google AI Mode?

No. Google says its AI search features use the core requirements and systems of Google Search, with no special AI file or markup required, and explicitly says Google Search ignores llms.txt. Pages should be indexable, eligible for snippets, useful, and understandable. This official guidance applies to Google and does not necessarily describe every other assistant.

### How should a brand measure visibility in ChatGPT?

Use a stable set of purchase questions, multiple paraphrases, and repeated runs. Save the engine, model, language, market, date, full answer, and cited URLs. Measure brand presence, citations, claim fidelity, referral visits, and conversions separately, then compare distributions before and after instead of relying on one screenshot.

### Does GEO replace SEO?

No. GEO adds stages to the visibility journey but still depends heavily on discoverability, indexation, authority, and page quality. Google specifically says its AI features use the core systems of Search. GEO extends SEO by making a source easier to select, cite, interpret, and measure; it does not remove the need to be found first.

### What GEO actions should an ecommerce store prioritize?

Start by making price, availability, variants, product identifiers, shipping, returns, and customer proof explicit in served HTML and consistent with structured data and merchant feeds. Then create pages that answer real comparison and selection questions with original evidence. Finally, measure whether those pages are retrieved, cited, visited, and associated with conversions.

## Sources

- [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026), review of 45 studies](https://arxiv.org/abs/2607.14035) (academic)
- [GEO: Generative Engine Optimization, KDD 2024](https://arxiv.org/abs/2311.09735) (academic)
- [C-SEO Bench: Does Conversational SEO Work?, NeurIPS 2025 Datasets and Benchmarks](https://proceedings.neurips.cc/paper_files/paper/2025/hash/27aa3aeff0f8460a7b43d30fa6c5c032-Abstract-Datasets_and_Benchmarks_Track.html) (academic)
- [SAGEO Arena: A Systematic Benchmark for Search-Augmented Generative Engine Optimization, KDD 2026](https://arxiv.org/abs/2602.12187) (academic)
- [What Gets Cited? Controlled Experiments on Source Selection in LLM-Based Search, SIGIR 2026](https://arxiv.org/abs/2605.25517) (academic)
- [Don't Measure Once: Instability and Repetition in Generative Engine Visibility, 2026](https://arxiv.org/abs/2604.07585) (academic)
- [Measuring the Causal Impact of ChatGPT Referral Traffic on Web Engagement, 2026](https://arxiv.org/abs/2606.04362) (academic)
- [Generative Engine Optimization in Production: Retrieval-Driven Content Generation at Pinterest, 2026](https://arxiv.org/abs/2602.02961) (case_study)
- [Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering, Findings of EMNLP 2023](https://aclanthology.org/2023.findings-emnlp.467/) (academic)
- [Google Search Central: Top ways to ensure your content performs well in Google's AI experiences](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) (official)
- [Google France: AI Overviews and AI Mode launch in France, July 22, 2026](https://blog.google/intl/fr-fr/nouveautes-produits/explorez-obtenez-des-reponses/recherche-ia-apercus-mode/) (official)
- [Adobe Digital Insights: AI-sourced traffic insights, Q2 2026](https://business.adobe.com/resources/sdk/.2026-q2-ai-traffic-report/q2-2026-adi-ai-sourced-traffic-insights.pdf) (industry)
- [Valiuz: GEO barometer across 13 French retailers, May 2025 to May 2026](https://www.valiuz.com/fr/post/nouvelle-%C3%A9dition-de-notre-barom%C3%A8tre-sur-l-%C3%A9volution-du-geo-l-irr%C3%A9sistible-ascension-du-visiteur) (industry)
- [Ahrefs: AI Overviews Reduce Clicks by 58%, updated February 2026](https://ahrefs.com/blog/ai-overviews-reduce-clicks-update/) (industry)

