The short answer
- Asking an AI once can easily give you the wrong winner. Omega won our watch study overall. Yet when we looked at the answers one at a time, only 39% of them put Omega first. Across all 11 categories we studied, a single answer picked the overall winner about 6 times in 10.
- Consistency belongs to the model, not the lab. mistral-medium-3.5 said Rolex 50 times out of 50. deepseek-v4.1-flash gave seven different #1s. Two models from the same lab can rank brands and cite sources very differently, so a strong position with one doesn't carry over to the next.
- No single answer matches the ranking, but pooled answers settle fast. The watch answers held 245 different top fives, and none had the final order. Pooled, 10 answers per model put the top five in the right order 90% of the time. What keeps growing is the list of brands named, from 12 to 28.
- Ten runs find a model's favourite. For 8 of the 11 models, 10 runs name the model's own #1 correctly at least 95% of the time. Getting its top three in order takes about 50. deepseek-v4.1-flash never settles, because its top two brands are genuinely tied.
What this measuresHow many answers it takes before an AI ranking of brands stops moving, and
which models make it wobble.
The experiment
The same question, 550 times
Ask an AI assistant which watch to wear on a first date and it gives you one confident answer. Ask again and you may get a different one. Ask a different assistant and you probably will.
So how many answers do you need before you know where a brand stands? Most of our category studies ask each model about 10 times. For watches we went five times further: the same question, 50 times, to each of 11 models from 8 labs.
- answers
- 550
- AI models · 8 labs
- 11
- runs per model
- 50
- watch brands named
- 28
Each answer is a ranked top five. Across all 550, Omega was the #1 pick 39% of the time, Cartier 26%, Seiko 14% and Rolex 13%. Every other brand shared the last 8%.
Omega
Cartier
Seiko
Everything else · 21%
39%26%14%
Watch brands: share of 550 answers that put each brand #1. Rolex took another 13%.
Finding 1
Every model has a personality
Line up each model's 50 answers side by side, one square per answer coloured by its #1 pick, and the personalities jump out.
Watch brands: each model's #1 pick, 50 answers in order
mistral-medium-3.5Rolex 50 of 50
gemini-3.8-flashOmega 46 of 50
gpt-6-astraCartier 36, NOMOS 13
qwen3.8-max8 different #1s
deepseek-v4.1-flash7 different #1s
- Omega
- Cartier
- Rolex
- Seiko
- NOMOS
- Tudor
- Grand Seiko
- Jaeger-LeCoultre
Grey squares are answers that came back empty.
The metronome. mistral-medium-3.5 put Rolex first in all 50 answers, and gave the exact same top five in 47 of them. The panel as a whole ranks Rolex fourth.
The fence-sitter. gpt-6-astra, OpenAI's newest model on the panel, couldn't choose between Cartier and NOMOS Glashütte. For its first 14 answers it swapped between the two almost every time. Over all 50, Cartier won 36 to 13.
The dice. deepseek-v4.1-flash named seven different #1s, and its favourite, Omega, won only 14 times. Its first ten answers made it look like a Tudor and Grand Seiko fan. Forty answers later, it looked like an Omega fan.
The shuffler. qwen3.8-max almost never repeats itself. Take any two of its answers: out of 1,176 possible pairs, only 6 gave the same top five in the same order. For mistral-medium-3.5, it was 1,082 out of 1,225.
Ask deepseek-v4.1-flash ten times and you'd call it a Tudor fan. Ask it fifty
times and it's an Omega fan.
Finding 2
Consistency belongs to the model, not the lab
Is it the lab that makes a model steady or erratic? We scored each model's consistency across all 11 categories: how often its answer named its own usual #1.
How often each model names its own usual #1, 11 categories
Google
gemini-3.8-flash96%
gemini-3.5-flash-lite90%
Anthropic
claude-opus-595%
claude-haiku-4.587%
OpenAI
gpt-6-astra93%
gpt-5.6-luna85%
One model per lab
mistral-medium-3.596%
grok-4.687%
qwen3.8-max87%
kimi-k383%
deepseek-v4.1-flash69%
50%100%
The first 10 answers per model in each category. Each lab's two models sit 6 to 8 points apart.
Most consistent: mistral-medium-3.5 and gemini-3.8-flash, which stick with their usual #1 96% of the time. Least consistent: deepseek-v4.1-flash, at 69%. It is the only model that never gave the same #1 in all 10 answers, in any category.
Every lab with two models on the panel has a steadier one and a looser one, 6 to 8 points apart. That is a wider gap than the one between the labs' averages. Knowing the lab doesn't tell you how steady the model is.
It doesn't tell you what the model thinks either. A strong position with one model doesn't carry over to another model from the same lab. OpenAI's two models agree that Cartier is #1 and disagree on much of the rest:
Watch brands: same lab, different opinions
Brandgpt-5.6-lunagpt-6-astra
Cartier9594
Omega8052
NOMOS Glashütte1585
Grand Seiko284
Each brand's score out of 100 with each model, across 50 answers.
Anthropic's two models don't even share a #1: claude-opus-5 picks Seiko and claude-haiku-4.5 picks Omega. Cartier scores 46 with Opus and 2 with Haiku. Google's two Geminis both pick Omega, but Cartier scores 82 with one and 34 with the other.
The sources differ too. gpt-5.6-luna cites Hodinkee in every answer; gpt-6-astra, OpenAI's newer model, cites it in only a third and leans on brands' own websites instead. So what got you cited by one model may not work for its successor.
There is a second kind of wobble, further down the list. Ask a model to name its top 5 brands, 10 times over. A perfectly steady model ends up with the same 5 brands. A restless one ends up with many more.
claude-haiku-4.5 usually agrees with itself about its #1. Below that, it roams: asked for 5 brands 10 times, it ends up with about 10 different brands, twice as many as it was asked for. In watches it named 18 different brands in 50 answers, plus three generic answers like "fashion watch brands". mistral-medium-3.5 named 7.
Asked for 5 brands, 10 times: how many different brands each model ends up with
claude-haiku-4.510.4
gpt-5.6-luna8.8
qwen3.8-max8.6
grok-4.68.0
deepseek-v4.1-flash7.9
kimi-k37.7
gemini-3.5-flash-lite7.5
gemini-3.8-flash7.1
gpt-6-astra6.6
claude-opus-56.4
mistral-medium-3.56.0
5 is the minimum: the same 5 brands every time. Average of 11 categories.
So is wobble a bell curve, with most models somewhere in the middle? No. Take every model in every category, 121 pairs in all. In 57% of them, the model gave the same #1 all 10 times. The wobble sits in a thin tail, and one model is in it more than any other: deepseek-v4.1-flash accounts for 4 of the 16 pairs at 6 in 10 or below.
Answers naming the model's usual #1, 121 model and category pairs
45678910
Answers out of 10
Finding 3
Answers scatter, the ranking settles
With models this different, no single answer is a good guide to the ranking. Of the 550 watch answers, 546 came back with a top five. The other four were empty: three from deepseek-v4.1-flash and one from qwen3.8-max, where the model spent its whole reply reasoning and never wrote the list. The 546 held 245 different top-five lists. Even the first answer from each of the 11 models was a different list. Not one answer listed the final top five in order: Omega, Cartier, Seiko, Rolex, Tudor. Only 69 even named the same five brands.
0 of 546
answers gave the panel's final top five in order
How can that be? The final ranking isn't an answer any model gave. It's an average of all of them. Each brand scores points for where it lands in each answer, and the points are added up. It works like "the average family has 2.4 children": the average is accurate, but no single family matches it.
Answers that match the final order, from the top down
Omega at #1213
Omega, Cartier at #1–2131
Omega, Cartier, Seiko at #1–331
Omega, Cartier, Seiko, Rolex at #1–428
All five in order0
Out of 546 answers with a top five.
Tudor is why the last step fails. It came 5th overall, but hardly any answer put it 5th. It appears in 170 answers and sits at #5 in only 8 of them. Usually it's #3 or #4, mostly from kimi-k3, deepseek-v4.1-flash and gemini-3.8-flash. When Tudor is in an answer, Rolex is usually pushed to #5: the most common list after mistral's is Omega, Cartier, Tudor, Seiko, Rolex. And when an answer has Rolex at #4, #5 goes to Casio, Hublot, Richard Mille or others, never Tudor. A third of answers rank Tudor fairly high and the rest leave it out. That averages to 5th.
The ranking comes from pooling the answers, and pooling is where more runs pay off. We redrew each model's answers at random, 1,000 times at each size, and checked how often the pooled top five came out in the final order.
Chance the pooled top five comes out in the final order
12351020
Answers per model, 11 models
The #1 settles almost at once. The rest takes longer. With one answer per model, Omega came out on top in 98% of redraws, but the whole top five was in the right order only 37% of the time. At 10 answers per model it's 90%, and at 20 it's 99%. The erratic models cancel out because they wobble in different directions, but it takes a few rounds for that to happen.
The second is the long tail. Every new batch of answers turned up brands nobody had named before: Breitling in run 7, Patek Philippe in run 15, Sinn in run 39, Oris in run 45. Each was named once, and three of the four came from claude-haiku-4.5. By 50 runs, the models had named 28 brands between them, more than twice as many as after one.
Watch brands named at least once, as runs are added
13102050
Runs per model
Each model's answers in the order they were collected. Generic answers, such as 'fast-fashion brands', aren't counted.
Finding 4
Ten runs find a model's favourite
People don't act on an average of 11 models. They ask one assistant. So the question that matters is how many times you need to ask one model before you know what it thinks.
We ran each model 50 times on watches. From those runs we worked out how often 10 runs get the right answer, and extrapolated what 100 runs would give.
95%
the chance that 10 runs name a model's own #1 correctly, for 8 of the 11 models
Watch brands: what 10, 50 and 100 runs per model get right
Runs per modelOwn #1 right 95%+ of the timeOwn top 3 in the right order± margin on its scores
108 of 11 models1 of 11±17 (5 to 33)
509 of 116 of 11±7 (2 to 15)
100, projected10 of 118 of 11±5 (2 to 10)
Margins are for the typical gap between two brands in a model's own top five, on the 0 to 100 score. Every figure comes from simulated batches drawn from each model's 50 real answers.
Ten runs tell you what most models prefer. Fifty tell you their top three. Model by model, it looks very different:
How often a model's own #1 comes out right
Model10 runs50 runs100 runs, projected
mistral-medium-3.5100%100%100%
claude-opus-5100%100%100%
claude-haiku-4.5100%100%100%
gemini-3.8-flash100%100%100%
gemini-3.5-flash-lite100%100%100%
kimi-k3100%100%100%
gpt-6-astra97%100%100%
gpt-5.6-luna97%100%100%
grok-4.688%97%100%
qwen3.8-max70%91%97%
deepseek-v4.1-flash50%66%72%
Share of 4,000 simulated batches, drawn from each model's 50 real answers, that named the same #1 as all 50. We did not run any model 100 times.
- mistral-medium-3.5 is settled at 10 runs. Its #1, top three and top five are all fixed.
- Most models get their #1 right at 10 runs. gpt-6-astra, gpt-5.6-luna, kimi-k3, both Geminis and both Claudes name their own #1 correctly 97% to 100% of the time. Their top three need about 50 runs.
- grok-4.6 needs about 50 runs. At 10, its #1 is right 88% of the time, because it rates Cartier and Omega closely. At 50, it's 97%.
- qwen3.8-max needs 50 to 100 runs. At 10, its #1 is right only 70% of the time, then 91% at 50 and 97% at 100.
- deepseek-v4.1-flash never settles. It's right 50% of the time at 10 runs, 66% at 50 and 72% at 100. Its top two are genuinely tied, Omega on 67 and Cartier on 65, so more runs only confirm the tie.
A tie is the honest answer in that last case. When a model scores two brands within a few points of each other, no number of runs will split them. Report it as a tie, not a winner.
What it means
Know which models you're asking
Your position isn't one number. It's a separate answer from every model your buyers use.
01
Ten runs give you a solid read on each model's #1. If your brand sits between #2 and #5, more runs, up to 50 per model, give you a much clearer picture of where it stands.
02
Models from the same lab answer differently. A strong position with gpt-5.6-luna doesn't carry over to gpt-6-astra, and the sources that got you cited by one may not work for the next. Check each model on its own.
Where does your brand stand across the models? Every BrandRanking.AI study asks a panel of models the same question about 10 times each and shows where they agree and where they don't. Browse the category studies or request a study for your own market.
How we measured
- Data. 550 answers to "Which watch brand would you wear on a first date?": 50 from each of 11 models, collected between 9 and 26 September 2026. Four came back empty. Five answers from a retired model, gemini-3.7-flash, are left out. The comparisons across categories use the 11 category studies from Note 01 plus AI brand visibility tools, with the first 10 answers per model. Every answer is a ranked top five from the model's training alone, with no web search.
- Rankings. A brand gets 100 points for a #1, 80 for a #2 and so on down to 20 for a #5, and 0 when it's left out. The points are averaged for each model, then across models with every model counted equally. First places are the share of answers that put a brand #1.
- Consistency. A model's usual #1 is the brand it put first most often in a category. Consistency is the share of its answers that named that brand.
- Redraws. For the run tests, we redrew each model's answers at random 1,000 times at each size.
- Runs per model. For each model, we redrew its 50 watch answers at random, with replacement, 4,000 times at 10, 50 and 100 answers, and compared each redraw with the ranking from all 50. No model was run 100 times: the 100-answer figures are this resampling projected past the data, and hold only if the model keeps answering as it did. The margin is 1.96 times the spread of the gap between two brands in a single answer, divided by the square root of the number of runs. Only watches has 50 runs per model, so these figures are for watches only.
BrandRanking.AI © 2026 · Study data: brandranking.ai/categories