Why your ChatGPT visibility keeps changing, and what actually gets you cited
Ask ChatGPT the same buyer question twice and you can get two different sets of recommended businesses. Not months apart, after you have changed your website. The same day, the same prompt, minutes apart.
Independent research published by Search Engine Land on 8 July 2026, by Chris Green and Suganthan Mohanadasan, put a number on it: across repeated runs of the same prompts, ChatGPT changed its primary search source 11.6% of the time. When that happened, the sources it cited to the user changed substantially with it.
If you have ever checked whether ChatGPT recommends your business, got an answer, and quietly assumed that answer was your position, this is the piece that explains why one check tells you almost nothing.

The same question, a different answer
The research repeated identical prompts and watched what came back. When ChatGPT stayed on the same backend search source between runs, the answers were reasonably stable. When the source switched, the overlap collapsed.
| What was measured | Same source | Source switched |
|---|---|---|
| URL overlap between repeat runs | 0.273 | 0.149 |
| Domain overlap between repeat runs | 0.265 | 0.155 |
Read the second column plainly. When the backend source changed, the two answers to the same question shared roughly 45% fewer of the same links and about 42% fewer of the same domains. A business cited in the first run could be absent from the second, replaced by a competitor, with nothing about the business, the question, or the wider web having changed in between.
This is the part most "AI visibility" checks never tell you: the output is noisy by design. A single snapshot is one roll of the dice reported as a verdict.
What is happening behind the answer
The citation cards you see in a ChatGPT answer are the end of a longer process. The research describes a retrieval layer that sits behind those cards and does the actual fetching, and it does not behave like a single search engine.
Two mechanisms matter for anyone trying to be recommended.
It fans out, and it goes looking for your competitors
A single buyer question does not become a single search. The research describes ChatGPT fanning out into many searches behind one prompt, including site: probes, pricing checks, and searches for competitors the user never named. So when a prospect asks "who is the best [your category] near me", the model is not just looking you up. It is assembling a shortlist, pricing you against it, and running its own searches for rivals you may not know you are being compared to.
That is the competitive set the wider internet assigns you, which is often not the one you think you are in. You do not get to nominate your competitors. The retrieval layer does.
It routes through several sources, and they disagree
The researchers found the system routing queries through internal source labels that never appear in the citations the user sees: Labrador, Bright, Oxylabs and SERP. In one dataset, Labrador handled 88.1% of primary sources, with the rest split across the others. They also observed a turn_use_case classification that causes some prompts to skip web search entirely.
We should be careful with these names. They are inferred from observed behaviour by outside researchers, not confirmed by OpenAI, and the exact plumbing will change. Treat the labels as a snapshot, not gospel. The finding that survives is the one that matters to you: ChatGPT is not one retriever with one view of the web. It is several, they weight the web differently, and which one answers a given prompt is not fully stable. That is the source of the 11.6%.
Why one check lies to you
Put the two findings together and the practical conclusion is uncomfortable for the "run a quick check" model of AI visibility.
If which backend answers is partly a coin toss, and different backends return substantially different sources, then a single query result is a sample of one from a distribution you cannot see. It might catch you on a good roll or a bad one. Either way it cannot tell you your actual standing, only that you appeared, or did not, on that one run.
This is why we measure the way we do, and why we are open about it. The free AI Referral Check tests a small set of buyer questions across ChatGPT, Claude and Gemini to give you a fast read. The paid Benchmark runs up to 100 real buyer phrasings across all four engines, including Perplexity, repeatedly, and names every competitor the engines refer buyers to instead of you. Repetition is not padding. It is the only way to turn a noisy signal into a rate you can trust. A mention rate or top-pick rate means something because it is measured across many prompts and many runs. A single "yes, ChatGPT mentioned me" means very little, because the next run might say no.
Honesty about this is the point, not a hedge. Anyone selling you a one-number AI score off a single check is selling you a coin toss with a confident voice.
What actually gets you cited across all of it
Here is the useful half. If you cannot control which backend answers, the winning move is to be the business that gets picked no matter which one does. You raise your floor across every pipeline at once.
The research lands on the same recommendation from the other direction. To be selected consistently across these systems, it points to plain HTML, crawlable facts, clear pricing and specifications, text-heavy pages, and strong third-party coverage. That is not a growth hack. It is the unglamorous foundation, and it maps almost exactly onto what we deploy in a Sprint.
Three things do the heavy lifting.
Be readable to a crawler. The retrieval layer fetches raw HTML. If your key facts, your services, your prices, your location, live only in JavaScript that renders in a browser but not in the raw page, large parts of your site are invisible to the fetch, whichever backend runs. Plain, server-rendered, text-heavy pages are legible to all of them.
State your facts where they can be lifted. Crawlable facts, clear pricing and specifications, direct answers to real buyer questions. Fan-out runs pricing checks and site: probes. Give those probes clean answers on your own pages rather than making the model infer them, or find them on someone else's.
Build third-party signal. This is the one owners underrate. For "best of" and "who should I use" questions, the engines lean heavily on what the rest of the web says about you: reviews, named expert mentions, consistent profiles across directories. Your own website arguing that you are the best counts for little across every backend. Independent corroboration counts for a lot. The study's emphasis on strong third-party coverage is the same finding in different words.
None of this makes the noise disappear. It shifts the odds on every roll. The more legible and corroborated you are, the more backends can find and trust you, and the less it matters which one happens to answer.
Readiness is the setup, the referral is the result
One caution, because it is the whole of our positioning. Everything above is readiness: the inputs you control. Clean HTML, stated facts, third-party signal. Getting those right is necessary. It is not the same as winning.
Readiness is the setup. The referral rate is the live result. A perfectly marked-up site still loses if the engines keep naming a competitor, and the only way to know which is true for you is to measure the output, repeatedly, across the engines your buyers actually use. Do the foundational work because it moves the outcome. Then check the outcome, more than once, because now you know the outcome is noisy.
Curious where you actually stand across ChatGPT, Claude and Gemini? Run the free AI Referral Check and see your score in about a minute: https://invisiblecompetitor.com/?scan=1