Across two Gemini runs, citation volume by model, high-coverage domains, source mix and Google Top 10 overlap remained in similar ranges. Exact page-level citations had lower repeatability.
This difference defines the baseline for this research series. Sleepwalker should report patterns that repeat and show normal run-to-run variation, or the noise floor, alongside URL-level changes. A cross-model comparison without a repeat of the same model can mistake normal model variation for a model difference.
Two repeat runs across three Gemini Flash models
We used 50 English-language consumer-electronics prompts, split evenly across
mobile phones, wearables, TVs, gaming consoles and headphones. None named Best
Buy, which was the target entity. Each prompt ran on gemini-3.5-flash,
gemini-3.6-flash and gemini-3.5-flash-lite in the United States locale.
Run A and Run B used the same prompts, models, target and locale. Together they produced 300 probes, 2,437 citations and zero execution errors. Run B began 19 to 21 hours after Run A, depending on the model. A Google organic Top 15 snapshot provided the comparison baseline.
| Model | Citations, Run A / B | Average citations per answer, A / B | Zero-citation prompts, A / B |
|---|---|---|---|
gemini-3.5-flash | 757 / 714 | 15.1 / 14.3 | 4 / 5 |
gemini-3.6-flash | 254 / 229 | 5.1 / 4.6 | 19 / 21 |
gemini-3.5-flash-lite | 238 / 245 | 4.8 / 4.9 | 12 / 12 |
The signals that held across both runs
The model-level citation pattern held. gemini-3.5-flash returned about 15
citations per answer in both runs. gemini-3.6-flash returned no citations on
19 and 21 prompts. gemini-3.5-flash-lite returned no citations on exactly 12
prompts in both runs.
High-coverage domains also held. reddit.com appeared on 29 of 50 prompts for
gemini-3.5-flash in both runs. rtings.com appeared on 15 and 15 prompts
for gemini-3.5-flash, and on 13 and 13 prompts for
gemini-3.5-flash-lite.
At the Top 10 depth, model-domain overlap changed by 0.1, 6.2 and 4.6 percentage points across the three models:
| Model | Google Top 10 domain overlap, Run A / B |
|---|---|
gemini-3.5-flash | 41.7% / 41.6% |
gemini-3.6-flash | 59.9% / 66.1% |
gemini-3.5-flash-lite | 67.1% / 62.5% |
Each model had no more than one answer with citations but no shared domain in the Google Top 10 in either run. Matched domains ranked 6.3 to 7.1 on average, and only about 15% to 18% were Google’s No. 1 result. Gemini pulled domains from across the first page instead of concentrating on the top organic result.
The source mix held too. YouTube accounted for 27% to 34% of Gemini citations across the six model-run combinations, against 10.8% of the Top 10 organic baseline. Retailers accounted for 0.8% to 2.0% of Gemini citations, against 7.0% of the baseline. The models were consistently reweighting the available source pool.
Exact pages need a repeat run for the same model
We compared each model’s citation set in Run A with its own citation set in Run B. Exact URL self-overlap ranged from 13.3% to 32.2%. Domain self-overlap was higher, from 30.9% to 46.1%.
For every prompt with at least one citation in either answer, we divided the number of shared distinct URLs or domains by the number of distinct URLs or domains across both answers. The table shows the average score across those prompts. A whole-run return rate answers a different question because it counts unique items across all 50 prompts.
| Model | URL self-overlap | Domain self-overlap |
|---|---|---|
gemini-3.5-flash | 13.3% | 30.9% |
gemini-3.6-flash | 22.4% | 34.6% |
gemini-3.5-flash-lite | 32.2% | 46.1% |
gemini-3.5-flash had an average exact URL overlap of 13.3% between runs. This
makes a single URL-level snapshot a poor basis for a claim about change over
time or a difference between two models.
One mid-tier domain moved substantially. techradar.com fell from 12 to four
prompts for gemini-3.6-flash. This example sets a limit on the measurement.
Give the stable high-coverage patterns more weight.
Average cross-model overlap replicated across runs
The three model pairs shared 14.4% of exact citation URLs on average in Run A and 12.9% in Run B. Domain overlap averaged 24.6% in Run A and 23.9% in Run B. No pair produced an identical citation set for any comparable prompt across the two runs.
Pair ordering did not repeat. In Run A, gemini-3.5-flash and
gemini-3.6-flash had the lowest URL overlap at 9.4%. In Run B, their overlap
was 10.9%, while gemini-3.5-flash and gemini-3.5-flash-lite had the lowest
value at 10.7%.
The models drew on different citation sets in both runs. The evidence does not support a ranking of which model pair is furthest apart from one 50-prompt run.
Best Buy ranked often and appeared in few citation sets
In Run A, bestbuy.com appeared in the Top 15 organic results for 30 prompts.
It was cited by the three models on five, one and two prompts, respectively.
That is an average of 2.7 prompts per model. The count is too small to treat as
a brand benchmark, but it is enough to show why rank and citation visibility
need separate measurement.
Report the noise floor with every result
Measure AI visibility per platform, model and prompt cluster. Then repeat the measurement before assigning meaning to page-level movement.
One run can support aggregate signals: average citation volume, zero-citation rate, source-type mix and domains that appear on at least 15 of 50 prompts. It cannot reliably support a story about a specific URL, a mid-tier domain or the ordering of two close model-comparison figures.
For URL-level reporting, include the same-model self-overlap beside the cross-model overlap. Without it, a dashboard can turn normal output variation into a model-release narrative.
Limitations
- This is a 50-prompt, one-target study in English for the United States.
- The two runs show short-term variation over an interval of about 19 to 21 hours. They do not measure multi-week drift.
- Changes in the live search index may be present in the self-overlap figure.
- Two runs do not provide confidence intervals. URL-level confidence requires at least three to five repeat runs.
- Source-type classification is manual. The long-tail bucket accounts for 21% to 28% of citations.
- Brand-level counts in this study range from two to seven mentions per model. They establish direction, not a reliable brand baseline.
- This design does not isolate the difference between commercial and speculative queries. The earlier sports comparison used different prompts, model generations, dates and a different target.