How to Tell a Real GEO Agency From a Repackaged SEO Shop
Every marketing agency added an AI page to its website in the last eighteen months. Some of them do this work properly. Most are selling the same content retainer with different vocabulary, and the difficult part for a buyer is that both sound identical on a sales call. There is a reliable way to tell them apart, and it comes down to how a vendor talks about measurement.
Start here: a screenshot proves nothing
The single most important fact in this category is that AI answers are unstable, and it is the thing most vendors would prefer you did not know. Researchers at the University of St Gallen published a study in April 2026 that isolated exactly how unstable. They ran identical prompts up to ten times on the same day, more than three thousand paired comparisons, and found that only about a third to two fifths of the cited sources overlapped between runs of the same question on the same day. Because nothing on the web changed in those minutes, the instability is coming from the model itself, not from the world. Independent work agrees from other angles. Ahrefs tracked more than 43,000 keywords and found a roughly 70 percent chance an AI Overview changes between one check and the next, with nearly half the cited sources swapping. SparkToro ran a study with 600 volunteers across almost 3,000 queries and found under a one in a hundred chance of getting an identical list of brands back when you ask the same question twice. Kevin Indig's analysis found only 2.2 percent of citations appear consistently after three runs.
That has a direct consequence for anyone selling you something. A screenshot of ChatGPT recommending your brand is one draw from a distribution. So is a screenshot of it recommending your competitor. Neither is evidence, and both are trivially easy to produce by asking repeatedly until you get the answer you wanted. The same study gives the honest threshold. You need roughly seven to eight runs of a question before a single day's reading is stable, and a rolling window of two to four weeks before a trend means anything.
The questions worth asking
How many times do you ask each question, and over what period? If the answer is once, the number they are about to show you has a margin of error wide enough to swallow any improvement they might claim. This is the question that separates the serious from the rest, and most vendors will not have a good answer. Can you show me the exact prompts? The questions are the measurement. If they are unwilling to show you the question set, you cannot judge whether they are measuring anything a buyer would actually type. We have caught this in our own library, where a question had been written that described our own service so precisely that no competitor could ever have matched it. It scored beautifully and meant nothing.
What happens when the number does not move? A vendor who has never had that conversation has not been doing this long enough. Who writes the content, and who checks it? AI generated recommendations are confidently wrong at a meaningful rate, and in this category the output goes onto a client's public website. Ask what the review step is and who performs it.
Answers that should end the conversation
Any guarantee of citations. Google's own documentation states that meeting every requirement still does not mean your content will be crawled, indexed or served. OpenAI states there is no way to guarantee top placement. A vendor guaranteeing what the platforms explicitly refuse to guarantee is telling you something about themselves.
Heavy acronym pressure. John Mueller of Google put it bluntly in August 2025: the higher the urgency and the stronger the push of new acronyms, the more likely they are just making spam and scamming.
Anything involving buying Reddit activity. Search Engine Journal documented vendors in June 2026 selling exactly this, aged accounts and paid upvotes aimed at manipulating what AI engines cite, following original reporting by 404 Media. It is the link farm of this era and it carries the same ending.
Programmatic comparison pages at scale. Lily Ray has been direct about where this crosses a line, describing the practice of generating a thousand near identical pages comparing yourself to every possible competitor as scaled content abuse under Google's existing spam policies.
What actually works, so you know what you are buying
The foundational research here is a paper from Princeton and Georgia Tech presented at KDD in 2024, which tested nine optimisation tactics across ten thousand queries. Adding quotations from credible sources, adding statistics, and citing sources properly all produced meaningful visibility gains, with the headline result reaching up to 40 percent. Keyword stuffing produced nothing, and in some measurements made things slightly worse.
The other finding worth knowing is where citations come from. A study by AirOps with the analyst Kevin Indig examined more than 21,000 brand mentions and found roughly 85 percent traced back to third party domains rather than the brand's own website, making a brand several times more likely to be mentioned through somebody else's page than its own.
That tells you what a competent vendor should be doing. Evidence dense content, real numbers, quotable claims, and a serious effort to get you onto the third party pages the engines retrieve from. If a proposal is entirely about publishing more blog posts on your own domain, it is aimed at the smaller half of the problem.
Sources
- Schulte, Bleeker and Kaufmann, Don't Measure Once: Measuring Visibility in AI Search, University of St Gallen, arXiv:2604.07585, April 2026
- Ahrefs, AI Overview Change, 11 November 2025
- SparkToro, AIs are highly inconsistent when recommending brands, 28 January 2026
- Kevin Indig, How to measure the impact of AI search the right way, Growth Unhinged, 15 July 2026
- Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024, arXiv:2311.09735
- AirOps with Kevin Indig, The Influence of Offsite Signals in AI Search, 17 October 2025
- Google Search Central, AI features and your website
- OpenAI Help Center, ChatGPT Search
- John Mueller via Search Engine Roundtable, 14 August 2025
- Search Engine Journal, Buying Reddit To Win AI Citations Is The New Link Farm, 29 June 2026
- Lily Ray via ppc.land, 13 May 2026
##FAQs
How long before I can judge whether it is working? Two to four weeks for a trustworthy baseline reading, and realistically a quarter before movement means anything. Should I just buy a monitoring tool and skip the agency? If you have someone with the time to act on what it tells you, yes. The tool reports the gap, it does not close it. Is any of the vendor research trustworthy? Some of it is good, and nearly all of it is published by companies selling the thing it supports. The St Gallen paper is the strongest independent source in the category because its authors have nothing to sell.
See where you stand.
Start with a free audit. We will show you exactly where you are cited across AI engines, and where you are not.
Get a free audit