The 3 GPT-5.6 runs averaged 22.0% domain overlap with Google’s matching organic Top 10. The GPT-5.5 run reached 30.0%. We measured the difference with 50 US consumer-electronics prompts on July 10 and 11, 2026, using Sleepwalker through an MCP connection in Claude.
The same runs found 14.4% mean exact citation URL overlap between GPT-5.5 and GPT-5.6.
OpenAI launched GPT-5.6 Sol, Terra and Luna on July 9, 2026. The 3 GPT-5.6 runs also shared little with each other. Their pairwise average was 13.4%, with comparisons ranging from 8.8% to 17.9%.
These 4 runs produced materially different citation sets while the platform, prompts and market stayed fixed. The result supports tracking AI visibility by platform and exact model.
It does not mean a brand will lose 86% of its citations after a model switch. Each model ran once. The study cannot separate model effects from normal answer variation. It shows how different these four citation sets were. Repeat runs are needed before we attribute the difference to a release.
What we measured
We ran gpt-5.6-terra on July 10, 2026. We ran gpt-5.5, gpt-5.6-sol
and gpt-5.6-luna on July 11. All 200 prompts completed without errors and
produced 651 citations.
The 50 US English prompts covered mobile phones, wearables, TVs, gaming consoles and headphones, with 10 prompts in each category. None named Best Buy, the target entity. Every model received the same prompts in the same order.
The study measured OpenAI API responses, not the consumer ChatGPT product.
Google Search used hl=en and gl=us. The search set excluded Shopping, AI
Overviews and video carousels. We collected the matching organic results in the
July 10 to 11 study window.
For each model pair, we compared the distinct citation URLs within each prompt.
We divided the shared set by the combined set, then averaged the 50 prompt
scores. URL matching removed www., fragments, trailing slashes and common
tracking fields. This is a per-prompt Jaccard score.
For Google, we calculated the share of citation domains found in the matching organic Top 10, then averaged the answer scores. The Top 10 is a comparison window, not evidence of which search results a model used.
Source types were classified by hand. Manufacturer pages are product or shop pages on a brand-owned domain. Support, news and community subdomains stayed outside that group. Retailers are multi-brand shops. Review and editorial sources form a separate group.
All four models returned about three citations per answer
| Model | Citations | Citations per answer |
|---|---|---|
gpt-5.5 |
171 | 3.42 |
gpt-5.6-sol |
161 | 3.22 |
gpt-5.6-terra |
161 | 3.22 |
gpt-5.6-luna |
158 | 3.16 |
The highest and lowest totals differed by 13 citations across 50 answers. The models differed far more in which pages they cited than in how many citations they returned.
Exact cited pages changed substantially by model
The GPT-5.5 run shared 15.0%, 17.1% and 11.2% of exact citation URLs with the Sol, Terra and Luna runs. That averages 14.4%.
| Model pair | Exact URL overlap |
|---|---|
| GPT-5.5 and Sol | 15.0% |
| GPT-5.5 and Terra | 17.1% |
| GPT-5.5 and Luna | 11.2% |
| Sol and Terra | 17.9% |
| Sol and Luna | 13.6% |
| Terra and Luna | 8.8% |
The GPT-5.6 models did not form one consistent citation profile. Sol and Terra agreed more with each other than either did with GPT-5.5. Terra and Luna had the lowest agreement in the study. One combined “GPT-5.6” result would hide that difference.
The GPT-5.5 run had the highest Google Top 10 overlap
| Model | Cited domains in matching Google Top 10 |
|---|---|
gpt-5.5 |
30.0% |
gpt-5.6-sol |
21.7% |
gpt-5.6-terra |
20.8% |
gpt-5.6-luna |
23.5% |
The GPT-5.6 average was 22.0%. A Top 10 comparison is generous here. Each answer contained about three citations. Even at that depth, most cited domains did not appear in the matching Google results.
Reddit makes the difference concrete. It ranked first organically for 22 of 50 queries and received no citations in any of the 200 OpenAI answers. Ranking first in Google was not enough to enter this citation set.
rtings.com appeared in 26% of answers
rtings.com received 78 of the 651 citations. It appeared in 52 of 200 answers. That is about 1 in 4 answers. Its answer coverage ranged from 11 to 16 prompts per model.
Its TV coverage was stronger. In seven of the 10 TV prompts, at least one model used rtings.com as the only cited domain. In those answers, rtings.com made up the entire cited source set.
The GPT-5.6 runs cited more manufacturer pages
Every GPT-5.6 run had a higher share of manufacturer product-page citations and a lower share of review or editorial citations than the GPT-5.5 run. Each table cell shows the citation count and its share of that model’s total.
| Model | Manufacturer pages | Review and editorial | Retailers |
|---|---|---|---|
| GPT-5.5 | 45 (26.3%) | 77 (45.0%) | 2 (1.2%) |
| Sol | 78 (48.4%) | 38 (23.6%) | 2 (1.2%) |
| Terra | 70 (43.5%) | 42 (26.1%) | 2 (1.2%) |
| Luna | 48 (30.4%) | 60 (38.0%) | 6 (3.8%) |
| All 4 runs | 241 (37.0%) | 217 (33.3%) | 12 (1.8%) |
The remaining citations fell into support, news, community and other source categories.
Samsung and Nintendo were clear examples. The main samsung.com domain
received 5 citations from GPT-5.5, compared with 12 from Sol, 8 from Terra and
12 from Luna. The main nintendo.com domain moved from 1 citation in GPT-5.5
to 5, 7 and 3 in the 3 GPT-5.6 runs.
Retailers remained nearly absent. bestbuy.com received 10 citations across
the 4 runs. Amazon did not appear among the 20 most-cited domains. Across this
sample, multi-brand retailers received 12 citations, compared with 241 for
manufacturer product pages.
The GPT-5.5 run used a more personal recommendation voice
A literal phrase check gives the tone observation a fixed rule. An answer counted when it contained “I’d” or “my pick.” The GPT-5.5 run used one of those phrases in 31 of 50 answers. Sol did so in 6, Terra in 11 and Luna in 18.
This does not score answer quality or count every possible first-person phrase. It shows a clear wording difference under one narrow rule. The 3 GPT-5.6 runs more often stated a choice without presenting it as the model’s personal pick.
AI visibility is a platform-model matrix
Citation volume stayed close across these 4 runs while the exact cited pages differed. A useful AI visibility test therefore records the platform, exact model, prompt set, market and run date.
A separate Gemini citation reproducibility test reached the same measurement lesson: broad source patterns can persist while exact cited URLs move substantially. Repeat the same model to establish its normal variation before treating a cross-model difference as a release effect.
Sleepwalker supports model-specific AI Visibility research through MCP, the API and the CLI. It uses pay-as-you-go credits, so the same prompt matrix can be tested against the platforms and exact models relevant to a brand.
Limits of this study
- The sample contains 50 US English consumer-electronics prompts for 1 target entity. It does not represent other markets, industries or prompt types.
- Each OpenAI model ran once. The study describes a dated snapshot, not a long-term trend or a causal model effect.
- The study measured OpenAI API responses, not the consumer ChatGPT product.
- Source-type classification was manual and depends on the stated categories.
- The results do not explain why a source was selected or predict future model behavior.