The short answer
- AI models justify their rankings with a small set of sources. In every category we measured, the five most-cited sources carry 35% to 60% of all references.
- The dominant source changes from one kind of market to the next. It's G2 for CRM software, Skytrax for airlines, Hodinkee for watches, Apple Podcasts for podcasts, SWE-bench for AI coding tools.
- In most categories, that source is the shared frame everyone is judged against. The models cite it at a similar rate for the winner and for the brands in 2nd to 5th place. Benchmarks are the exception: SWE-bench is cited mostly for the models that top it.
- A brand's own website describes the brand. It doesn't rank it. Citations to a brand's own site are almost always used to explain that one brand. The category winner's own site accounts for 15% or less of the links in its category.
- Where no neutral scoreboard exists, each firm is described by its own research. In AI consulting, the five most-cited sources are all consulting firms' own websites.
What this measuresThe bias built into the models by their training. Every answer here came from
the model's own knowledge, with no web search.
Scope
What the model believes before it searches
Most AI assistants now search the web before they answer, and that search can bring in sources and brands the model would never have named on its own. That live layer is not part of this study.
What we measure is the layer underneath it: what each model already believes about a category before it looks anything up, and the starting point any search adds to.
The question
Where does a confident, ranked answer come from?
Ask ChatGPT, Claude or Gemini which CRM to buy and you get a confident, ranked answer. Brands increasingly want to know how that answer is made. Is it the brand's website, its ads, its reviews? Or something else entirely?
To find out, we didn't ask one model once. We ran the same buyer question through a panel of models, again and again, in ten very different categories, from CRM software and no-code tools to business-class airlines, luxury watches, UK neobanks and MBA programmes. Each time, the model had to rank its top five and name the sources behind each pick, answering from its training with no web search.
- ranked answers
- 1,533
- AI models · 8 labs
- 14
- categories
- 10
- source references
- 20,253
The result is the largest cross-category look we know of at what AI models recall when they recommend a brand. Six patterns show up across the categories, and they all point the same way.
Finding 1
A few sources do most of the work
Each answer named two to three sources per brand, about 13 references per answer. You might expect those references to be spread thinly across the web. They aren't.
In every category, the five most-cited sources account for between 35% and 60% of all references. In CRM software it's 60%, in business-class airlines 58%, in EVs and podcasts 57%. The long tail exists, but the models return to the same handful of places again and again.
35–60%
of all references go to the five most-cited sources, in every category
Top-five share, highest four categories
G2
Capterra
Everything else · 52%
31%17%
CRM software: sources behind 134 answers from 14 models. G2 alone is 31% of 1,875 references, Capterra another 17%.
CRM is the extreme case. Nearly half of every justification the models gave (48%) points to just two software review directories. But the shape repeats everywhere: one or two dominant sources, a supporting cast, then a long tail of occasional mentions.
Finding 2
Every category has a different scoreboard
This is the pattern that matters most, and the one generic "AI SEO" advice misses. The dominant source changes from one kind of market to the next. There is no universal list of sites that AI models trust. Only CRM and no-code, both software, share one.
Share of references to each category's most-cited source
G2CRM software31%
G2No-code platforms31%
Apple PodcastsBusiness podcasts27%
Skytrax World Airline AwardsBusiness-class airlines22%
HodinkeeWatch brands21%
TrustpilotUK neobanks21%
Financial TimesMBA programmes17%
Car and DriverEV brands16%
SWE-benchAI models for coding11%
Firms' own sites (BCG X)AI consulting9%
Look at what's on that list. It isn't one kind of website. It's a review directory in software, an awards body in aviation, a league table in business schools, an enthusiast magazine in watches, a streaming platform's charts in podcasts, and a technical benchmark in AI coding. We found six distinct kinds of scoreboard:
- 01Review directoriesG2, Capterra, TrustpilotProducts that look similar on paper
- 02Awards and rankingsSkytrax, FT, US News, Poets&QuantsAn established annual league table
- 03Specialist pressHodinkee, Car and Driver, EdmundsWhere expert taste carries authority
- 04Platform chartsApple PodcastsThe distributor's own listings
- 05Technical benchmarksSWE-bench, Aider, LMArenaWhere performance can be scored
- 06NoneFirms' own researchNo neutral scoreboard
The consequence is simple. A brand can't know which site shapes its AI ranking by reasoning from another category. A tactic that works for a SaaS company (collect G2 reviews) is irrelevant for an airline, whose standing is shaped by Skytrax and points-and-miles blogs like The Points Guy and One Mile at a Time.
Business-class airlines
Skytrax22%
The Points Guy16%
One Mile at a Time11%
Finding 3
The scoreboard is the same for everyone
You might assume the winning brand wins because it has a special relationship with the dominant source, as if Hodinkee simply "likes" Omega more. The data says otherwise. In most categories, the models cite the dominant source at a similar rate whether they're explaining the #1 brand or the #5.
UK neobanksTrustpilot
CRM softwareG2
AirlinesSkytrax
No-codeG2
WatchesHodinkee
#1#2#3#4#5
Share of the sources behind each of the top five brands going to the dominant source.
That's the key to how these rankings work. The dominant source isn't a vote for one brand. It's the frame the model compares every brand within. Starling Bank, Revolut, Tide, Monzo and Wise are all measured against Trustpilot. HubSpot, Salesforce, Zoho and Pipedrive are all measured against G2. Each brand is explained by how it looks on that shared scoreboard.
Benchmarks work differently. In AI coding, SWE-bench is 30% of the sources behind Anthropic, 16% behind OpenAI, 9% behind Google and 1% behind Meta. A benchmark isn't cited as a neutral frame. It's cited as evidence for whoever leads it.
Brands don't compete on their own websites. They compete on someone else's
page, and in almost every category it's a different page.
Each brand has its own shop window, its website. But when a model explains its ranking, it points to one shared board that every brand is listed on.
Finding 4
Your website describes you. It doesn't rank you.
That doesn't mean brands' own sites are ignored. They show up often, but for a specific job. When a model cites a brand's own website, it's almost always to describe that same brand: its pricing page, its features, its product lineup.
Across the ten categories, citations to a brand's own site were almost always attached to that brand's own entry, and for most brands every one of them was. Models use hubspot.com to explain HubSpot, and G2 to compare HubSpot with everyone else. In HubSpot's case, its own site makes up a quarter of the links behind its #1 ranking, but G2 makes up 39%, and it's G2 the model uses for Salesforce, Zoho and Pipedrive too.
Category winner's own site, share of all linked references
Qatar AirwaysAirlines2.1%
TeslaEV brands3.9%
Stanford GSBMBA programmes4.0%
OmegaWatches4.7%
HubSpotCRM software5.5%
AirtableNo-code6.8%
AnthropicAI coding7.4%
Starling BankUK neobanks9.2%
McKinsey & CompanyAI consulting14.6%
Even the winners' own sites are a minor share of what the models cite in their category. The one clear exception is McKinsey at 14.6%, in the one category with no outside scoreboard.
Finding 5
No scoreboard? Self-published research wins.
AI consulting is the only category in our study without a dominant outside scoreboard. There's no G2 for transformation partners, no Skytrax for strategy firms; the closest, Gartner, is sixth with 8% of links. So what do the models cite instead?
The firms' own publications. The five most-cited sources in the category are all consulting firms' own websites (McKinsey, BCG, Accenture, Deloitte and IBM), together about 62% of all linked references. Two-thirds of the links behind McKinsey's #1 ranking point to mckinsey.com, much of it its "State of AI" research.
62%
of linked references in AI consulting go to five firms' own sites
- McKinsey
- BCG
- Accenture
- Deloitte
- IBM
Each firm's site is cited only for that firm: mckinsey.com explains McKinsey, bcg.com explains BCG. Where there's no neutral referee, the models describe every firm through its own publications. Those sites are the most-cited sources in the category. McKinsey and BCG are level at the top (137 and 136 links), and McKinsey is also ranked first. We didn't measure how much each firm publishes, so this shows what the models cite, not why.
Finding 6
Models from different labs read the same scoreboards
Our panel mixed models from Anthropic, OpenAI, Google, xAI, Mistral and three Chinese labs: DeepSeek, Moonshot (Kimi) and Alibaba (Qwen). They are trained by different teams on different data mixes. Yet they converge on the same scoreboards.
The dominant source was cited by nearly every model on the panel: 14 of 14 in CRM and MBA, 12 of 12 in airlines, watches and EVs, 11 of 12 in no-code. Even in the most fragmented categories, at least 9 models cited it. US and Chinese models named the same top source in 6 of the 10 categories.
CRM software14 of 14
MBA programmes14 of 14
Airlines12 of 12
Watches12 of 12
EV brands12 of 12
No-code11 of 12
Whatever differences exist between models, and there are some, they were trained on the same public web. A brand's standing on a category's scoreboard doesn't depend much on which assistant the buyer uses.
What it means
Find your scoreboard
Taken together, the findings answer the question in the title. In nine of the ten categories, AI models justify their rankings by recalling how each brand stands on a small number of third-party scoreboards, and those scoreboards are specific to each category. A brand's own site helps the model describe the brand, but the comparison happens elsewhere. The exception is AI consulting, which has no outside scoreboard, so the models describe each firm through its own site. This is the bias the models carry from training; a live web search can add to it, but it starts from here.
For a brand, three things follow:
01
Generic "AI optimisation" is aimed at the wrong target. Your own pages are where the model learns what you are. They aren't where it learned how you compare.
02
Your scoreboard has to be measured, not guessed. Nobody would guess that a streaming platform's charts shape podcast rankings, or that points-and-miles blogs rival an awards body for airlines. The only reliable way to find your category's scoreboard is to see what the models actually cite.
03
If your category has no scoreboard, your own site is the evidence. In consulting, the firms' own sites are the most-cited sources, and McKinsey, level with BCG as the most-cited, is ranked first.
What's the scoreboard in your category? Every BrandRanking.AI study shows which sources AI models cite in a category, which brands they rank, and where models disagree. Browse the category studies or request a study for your own market.
How we measured
- Question. One natural buyer question per category, e.g. "What CRM software would you recommend?" or "Which airline would you recommend for business-class travel?"
- Models. 14 models from 8 labs: claude-opus-5, claude-haiku-4.5, gpt-6-astra, gpt-5.6-luna, gemini-3.8-flash, gemini-3.5-flash-lite, grok-4.6, deepseek-v4.1-flash, kimi-k3, qwen3.8-max and mistral-medium-3.5 (the core panel, about 10 runs each per category), plus earlier runs from gpt-5.6-sol, gpt-5.6-terra and gemini-3.7-flash in some categories.
- Answers. Each run asks for a ranked top five with reasoning and sources for each pick. 1,533 successful runs across 10 categories; watch brands and AI coding were sampled more deeply.
- Sources. Sources are recalled by the models from their training, not retrieved by live search, so they show what a model associates with a brand, not pages it read while answering. About a quarter are named without a link. Shares in Findings 1 and 2 are over all references, grouped by site, with a site's subdomains counted together (uk.trustpilot.com with trustpilot.com, rankings.ft.com with ft.com). Other shares are calculated over references that include a link.
- Scope. This measures the bias built into the models by their training: what they recommend when answering from their own knowledge. Assistants that search the web before answering (ChatGPT search, Perplexity, Google AI Overviews) add a live layer on top that this study does not measure.
BrandRanking.AI © 2026 · Study data: brandranking.ai/categories