A survey of 45 studies puts the GEO evidence on far shakier ground than the sales decks suggest. Only two levers hold up: topical relevance, and where your content sits in the model’s context. The rest either fails to transfer between settings, or backfires outright. So if someone sells you GEO work on the promise of a 40% lift, you are buying a number from a lab setup that does not match how search really runs.
The 40% everyone quotes came from a fixed context
That number traces back to one paper. Aggarwal and co-authors at Princeton and IIT Delhi coined the term GEO and built GEO-bench, a set of 10,000 queries. Their 2024 KDD paper reports that GEO “can boost visibility by up to 40% in GE responses.” That result is real. But read the setup.
The test hands the model a fixed set of sources, then rewrites one of them. So the page has already won its slot before the test starts. In other words, the 40% measures how much better a cited page can look once it is cited. It says nothing about getting cited in the first place.
That gap matters more than it sounds. Most of the work of AI search happens upstream, in the crawl and the fetch and the ranking. A lift that only applies after all of that is a lift on the easy part.
What the GEO evidence actually supports
Olivier Martinez reviewed the field in a critical survey published in July 2026, covering 45 studies from November 2023 onward. His read is blunt. Topical relevance and context position are “the most reproducible levers.” Generic tricks “transfer poorly.” Rival sites erode whatever gains one site makes. And citation-oriented rewrites can hurt retrieval.
The survey’s summary line deserves quoting in full. No reviewed technique “shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior.” Not one, after three years and 45 studies.

Body-only rewrites made things worse, not better
The sharpest test comes from SAGEO Arena, a benchmark built by Sunghwan Kim and colleagues at Yonsei. Instead of a fixed context, it runs the full pipeline: retrieval, then reranking, then the answer. The setup holds 2,700 queries against a corpus of 171,003 documents.
Then they applied standard GEO tactics to page bodies and measured every stage. Top-20 retrieval fell from 0.58 to 0.53. Top-10 presence after reranking dropped from 1.00 to 0.84. Citation rate slipped from 0.50 to 0.47. Adding jargon cost 4.54 rank spots on average, because it pulled the page’s wording away from how people phrase queries.

Their verdict: “body-only optimization is insufficient for enhancing visibility throughout the pipeline.” Put plainly, tactics that win in a fixed context can bury you before the answer stage ever happens. You do not get a citation you were never retrieved for.
Why the GEO evidence keeps flipping on you
Part of the mess is testing. Martinez cites work by Schulte finding that 57.8% of repeated ChatGPT prompts never fired a web search at all. So more than half the runs answered from memory. He also cites Grossman and colleagues, who put URL-level overlap between platforms at Jaccard scores of 0.11 to 0.18.
Translated: two engines asked the same question cite almost wholly different pages. Meanwhile the same engine asked twice may not even search. Any audit built on one run per prompt reads noise and calls it signal.
This is why brand tracking dashboards swing so hard week to week. Often the tool did not catch a change in your ranking. It caught a coin flip.
Most GEO evidence you get sold sits at the bottom of the scale
Martinez grades claims on a five-tier scale. The top tier means a real field trial with logs and controls. One rung down means live engines tested with repeats, reworded prompts and multiple dates. At the weakest tier sit fixed-context or synthetic studies.
Nearly every case study in a vendor deck lands on those bottom rungs. So does the 40% result. That does not make either one worthless. Still, it does mean the number cannot carry the weight the market puts on it.
This gives you a cheap filter for pitches. Ask how many times they ran each prompt. Then ask whether they held a control page, and which dates the runs cover. Vendors with real GEO evidence answer in seconds, because they logged it. Vendors without it change the subject to case studies.
None of this means you should sit still. It means you should spend on the parts with support behind them, and treat the rest as a bet rather than a plan. Budget accordingly. A tactic at the weakest evidence tier is worth trying on ten pages, not rolling out across four hundred.
The second lever is position, not polish
Topical relevance gets all the attention. Context position, the other lever the GEO evidence supports, gets almost none. Yet it is the one that explains why the rest of the playbook keeps failing.
Context position means where your page sits in the stack of sources the model receives. A page handed over third gets read differently than the same page handed over fourteenth. You cannot set that slot by hand. Instead it falls out of retrieval and reranking, which is exactly the stage most GEO advice skips.
So the two supported levers are not really separate jobs. Relevance decides whether you enter the set. Rank inside that set decides how much weight you carry. Both are won upstream, before a single word of your page reaches the model. Polish the prose all you like, but it cannot move either one.
Optimize for retrieval first, the answer second
The practical read is not “GEO is fake.” It is that order matters, and most advice has the order backwards.
Retrieval is the gate. If your page does not surface in the top 20 for the way people phrase the question, nothing downstream helps. So write in the words of the query, not the words of your industry. Cover the topic fully instead of sprinkling citation bait. Drop the jargon your buyers never type. Then, and only then, worry about how the passage reads once a model holds it.
This is also why we keep telling clients that schema markup does not buy citations. Same failure mode: a tactic aimed at the answer stage, bolted onto a page that never clears retrieval.
How to test GEO evidence on your own site
You do not need a research lab. But you do need to stop drawing conclusions from single runs. Martinez’s protocol scales down fine:
- Run each prompt seven to eight times, not once.
- Write three to five reworded versions of every prompt you care about.
- Hold an untouched control page next to the page you change.
- Log the raw answer, the citations, the timestamp, the locale, and whether the engine actually searched.
- Repeat across several dates before you call anything a result.

That is more work than a dashboard screenshot. Yet it is the difference between knowing a change worked and hoping it did. Run it on five prompts that matter rather than fifty that do not.
Want help telling real GEO gains from noise
We build AI search programs on tests that survive repetition, not on tactics borrowed from a fixed-context paper. If you want a read on what is really moving for your brand, get in touch.