# How consistent are AI models when ranking brands?

_We asked 11 AI models which watch brand they would wear on a first date, 50 times each. Some gave the same answer every time. Others seemed to roll dice. The 546 answers held 245 different top-five lists, and not one matched the final ranking. Pooled, they settle fast: 10 answers per model get the top five right nine times in ten. Asking more changes the long tail, not the top of the list._

Field Notes 02 · 28 Sept 2026 · 7 min read · https://www.brandranking.ai/field-notes/how-consistent-are-ai-models-when-ranking-brands

---

## The short answer

1. **Asking an AI once can easily give you the wrong winner.** Omega won our watch study overall. Yet when we looked at the answers one at a time, only 39% of them put Omega first. Across all 11 categories we studied, a single answer picked the overall winner about 6 times in 10.
2. **Consistency belongs to the model, not the lab.** mistral-medium-3.5 said Rolex 50 times out of 50. deepseek-v4.1-flash gave seven different #1s. Two models from the same lab can rank brands and cite sources very differently, so a strong position with one doesn't carry over to the next.
3. **No single answer matches the ranking, but pooled answers settle fast.** The watch answers held 245 different top fives, and none had the final order. Pooled, 10 answers per model put the top five in the right order 90% of the time. What keeps growing is the list of brands named, from 12 to 28.
4. **Ten runs find a model's favourite.** For 8 of the 11 models, 10 runs name the model's own #1 correctly at least 95% of the time. Getting its top three in order takes about 50. deepseek-v4.1-flash never settles, because its top two brands are genuinely tied.

> **What this measures**
>
> How many answers it takes before an AI ranking of brands stops moving, and
> which models make it wobble.

_The experiment_

## The same question, 550 times

Ask an AI assistant which watch to wear on a first date and it gives you one confident answer. Ask again and you may get a different one. Ask a different assistant and you probably will.

So how many answers do you need before you know where a brand stands? Most of our category studies ask each model about 10 times. For watches we went five times further: the same question, 50 times, to each of 11 models from 8 labs.

- **550** answers
- **11** AI models · 8 labs
- **50** runs per model
- **28** watch brands named

Each answer is a ranked top five. Across all 550, Omega was the #1 pick 39% of the time, Cartier 26%, Seiko 14% and Rolex 13%. Every other brand shared the last 8%.

| Source | Share |
| - | - |
| Omega | 39% |
| Cartier | 26% |
| Seiko | 14% |
| Everything else | 21% |

_Watch brands: share of 550 answers that put each brand #1. Rolex took another 13%._

---

_Finding 1_

## Every model has a personality

Line up each model's 50 answers side by side, one square per answer coloured by its #1 pick, and the personalities jump out.

**Watch brands: each model's #1 pick, 50 answers in order**

| Name | Context | Answers | First choices |
| - | - | - | - |
| mistral-medium-3.5 | Rolex 50 of 50 | 50 | Rolex 50 |
| gemini-3.8-flash | Omega 46 of 50 | 50 | Omega 46, Cartier 4 |
| gpt-6-astra | Cartier 36, NOMOS 13 | 50 | Cartier 36, NOMOS 13, Seiko 1 |
| qwen3.8-max | 8 different #1s | 50 | Seiko 23, Cartier 10, NOMOS 7, Omega 4, Grand Seiko 3, Jaeger-LeCoultre 1, No answer 1, Tudor 1 |
| deepseek-v4.1-flash | 7 different #1s | 50 | Omega 14, Cartier 10, Tudor 9, Seiko 6, Grand Seiko 5, No answer 3, NOMOS 3 |

_Grey squares are answers that came back empty._

**The metronome.** mistral-medium-3.5 put Rolex first in all 50 answers, and gave the exact same top five in 47 of them. The panel as a whole ranks Rolex fourth.

**The fence-sitter.** gpt-6-astra, OpenAI's newest model on the panel, couldn't choose between Cartier and NOMOS Glashütte. For its first 14 answers it swapped between the two almost every time. Over all 50, Cartier won 36 to 13.

**The dice.** deepseek-v4.1-flash named seven different #1s, and its favourite, Omega, won only 14 times. Its first ten answers made it look like a Tudor and Grand Seiko fan. Forty answers later, it looked like an Omega fan.

**The shuffler.** qwen3.8-max almost never repeats itself. Take any two of its answers: out of 1,176 possible pairs, only 6 gave the same top five in the same order. For mistral-medium-3.5, it was 1,082 out of 1,225.

> Ask deepseek-v4.1-flash ten times and you'd call it a Tudor fan. Ask it fifty
> times and it's an Omega fan.

---

_Finding 2_

## Consistency belongs to the model, not the lab

Is it the lab that makes a model steady or erratic? We scored each model's consistency across all 11 categories: how often its answer named its own usual #1.

**How often each model names its own usual #1, 11 categories**

| Name | Group | Value |
| - | - | - |
| gemini-3.8-flash | Google | 96% |
| gemini-3.5-flash-lite | Google | 90% |

| Name | Group | Value |
| - | - | - |
| claude-opus-5 | Anthropic | 95% |
| claude-haiku-4.5 | Anthropic | 87% |

| Name | Group | Value |
| - | - | - |
| gpt-6-astra | OpenAI | 93% |
| gpt-5.6-luna | OpenAI | 85% |

| Name | Group | Value |
| - | - | - |
| mistral-medium-3.5 | One model per lab | 96% |
| grok-4.6 | One model per lab | 87% |
| qwen3.8-max | One model per lab | 87% |
| kimi-k3 | One model per lab | 83% |
| deepseek-v4.1-flash | One model per lab | 69% |

_The first 10 answers per model in each category. Each lab's two models sit 6 to 8 points apart._

**Most consistent: mistral-medium-3.5 and gemini-3.8-flash**, which stick with their usual #1 96% of the time. **Least consistent: deepseek-v4.1-flash**, at 69%. It is the only model that never gave the same #1 in all 10 answers, in any category.

Every lab with two models on the panel has a steadier one and a looser one, 6 to 8 points apart. That is a wider gap than the one between the labs' averages. Knowing the lab doesn't tell you how steady the model is.

It doesn't tell you what the model thinks either. **A strong position with one model doesn't carry over to another model from the same lab.** OpenAI's two models agree that Cartier is #1 and disagree on much of the rest:

**Watch brands: same lab, different opinions**

| Brand | gpt-5.6-luna | gpt-6-astra |
| - | - | - |
| Cartier | 95 | 94 |
| Omega | 80 | 52 |
| NOMOS Glashütte | 15 | 85 |
| Grand Seiko | 28 | 4 |

_Each brand's score out of 100 with each model, across 50 answers._

Anthropic's two models don't even share a #1: claude-opus-5 picks Seiko and claude-haiku-4.5 picks Omega. Cartier scores 46 with Opus and 2 with Haiku. Google's two Geminis both pick Omega, but Cartier scores 82 with one and 34 with the other.

The sources differ too. gpt-5.6-luna cites Hodinkee in every answer; gpt-6-astra, OpenAI's newer model, cites it in only a third and leans on brands' own websites instead. So what got you cited by one model may not work for its successor.

There is a second kind of wobble, further down the list. Ask a model to name its top 5 brands, 10 times over. A perfectly steady model ends up with the same 5 brands. A restless one ends up with many more.

claude-haiku-4.5 usually agrees with itself about its #1. Below that, it roams: asked for 5 brands 10 times, it ends up with about 10 different brands, twice as many as it was asked for. In watches it named 18 different brands in 50 answers, plus three generic answers like "fashion watch brands". mistral-medium-3.5 named 7.

**Asked for 5 brands, 10 times: how many different brands each model ends up with**

| Name | Share |
| - | - |
| claude-haiku-4.5 | 10.4 |
| gpt-5.6-luna | 8.8 |
| qwen3.8-max | 8.6 |
| grok-4.6 | 8.0 |
| deepseek-v4.1-flash | 7.9 |
| kimi-k3 | 7.7 |
| gemini-3.5-flash-lite | 7.5 |
| gemini-3.8-flash | 7.1 |
| gpt-6-astra | 6.6 |
| claude-opus-5 | 6.4 |
| mistral-medium-3.5 | 6.0 |

_5 is the minimum: the same 5 brands every time. Average of 11 categories._

So is wobble a bell curve, with most models somewhere in the middle? No. Take every model in every category, 121 pairs in all. In 57% of them, the model gave the same #1 all 10 times. The wobble sits in a thin tail, and one model is in it more than any other: deepseek-v4.1-flash accounts for 4 of the 16 pairs at 6 in 10 or below.

**Answers naming the model's usual #1, 121 model and category pairs**

| Answers out of 10 | Value |
| - | - |
| 4 | 1 |
| 5 | 7 |
| 6 | 8 |
| 7 | 10 |
| 8 | 15 |
| 9 | 11 |
| 10 | 69 |

---

_Finding 3_

## Answers scatter, the ranking settles

With models this different, no single answer is a good guide to the ranking. Of the 550 watch answers, 546 came back with a top five. The other four were empty: three from deepseek-v4.1-flash and one from qwen3.8-max, where the model spent its whole reply reasoning and never wrote the list. The 546 held 245 different top-five lists. Even the first answer from each of the 11 models was a different list. **Not one answer listed the final top five in order**: Omega, Cartier, Seiko, Rolex, Tudor. Only 69 even named the same five brands.

**0 of 546** answers gave the panel's final top five in order

How can that be? **The final ranking isn't an answer any model gave. It's an average of all of them.** Each brand scores points for where it lands in each answer, and the points are added up. It works like "the average family has 2.4 children": the average is accurate, but no single family matches it.

**Answers that match the final order, from the top down**

| Name | Share |
| - | - |
| Omega at #1 | 213 |
| Omega, Cartier at #1–2 | 131 |
| Omega, Cartier, Seiko at #1–3 | 31 |
| Omega, Cartier, Seiko, Rolex at #1–4 | 28 |
| All five in order | 0 |

_Out of 546 answers with a top five._

Tudor is why the last step fails. It came 5th overall, but hardly any answer put it 5th. It appears in 170 answers and sits at #5 in only 8 of them. Usually it's #3 or #4, mostly from kimi-k3, deepseek-v4.1-flash and gemini-3.8-flash. When Tudor is in an answer, Rolex is usually pushed to #5: the most common list after mistral's is Omega, Cartier, Tudor, Seiko, Rolex. And when an answer has Rolex at #4, #5 goes to Casio, Hublot, Richard Mille or others, never Tudor. A third of answers rank Tudor fairly high and the rest leave it out. That averages to 5th.

The ranking comes from pooling the answers, and pooling is where more runs pay off. We redrew each model's answers at random, 1,000 times at each size, and checked how often the pooled top five came out in the final order.

**Chance the pooled top five comes out in the final order**

| Answers per model, 11 models | Value |
| - | - |
| 1 | 37% |
| 2 | 56% |
| 3 | 62% |
| 5 | 79% |
| 10 | 90% |
| 20 | 99% |

**The #1 settles almost at once. The rest takes longer.** With one answer per model, Omega came out on top in 98% of redraws, but the whole top five was in the right order only 37% of the time. At 10 answers per model it's 90%, and at 20 it's 99%. The erratic models cancel out because they wobble in different directions, but it takes a few rounds for that to happen.

The second is the long tail. Every new batch of answers turned up brands nobody had named before: Breitling in run 7, Patek Philippe in run 15, Sinn in run 39, Oris in run 45. Each was named once, and three of the four came from claude-haiku-4.5. By 50 runs, the models had named 28 brands between them, more than twice as many as after one.

**Watch brands named at least once, as runs are added**

| Runs per model | Value |
| - | - |
| 1 | 12 |
| 3 | 19 |
| 10 | 23 |
| 20 | 26 |
| 50 | 28 |

_Each model's answers in the order they were collected. Generic answers, such as 'fast-fashion brands', aren't counted._

[Open the watch study](https://www.brandranking.ai/category-ranking/watch-brands)

---

_Finding 4_

## Ten runs find a model's favourite

People don't act on an average of 11 models. They ask one assistant. So the question that matters is how many times you need to ask one model before you know what it thinks.

We ran each model 50 times on watches. From those runs we worked out how often 10 runs get the right answer, and extrapolated what 100 runs would give.

**95%** the chance that 10 runs name a model's own #1 correctly, for 8 of the 11 models

**Watch brands: what 10, 50 and 100 runs per model get right**

| Runs per model | Own #1 right 95%+ of the time | Own top 3 in the right order | ± margin on its scores |
| - | - | - | - |
| 10 | 8 of 11 models | 1 of 11 | ±17 (5 to 33) |
| 50 | 9 of 11 | 6 of 11 | ±7 (2 to 15) |
| 100, projected | 10 of 11 | 8 of 11 | ±5 (2 to 10) |

_Margins are for the typical gap between two brands in a model's own top five, on the 0 to 100 score. Every figure comes from simulated batches drawn from each model's 50 real answers._

**Ten runs tell you what most models prefer. Fifty tell you their top three.** Model by model, it looks very different:

**How often a model's own #1 comes out right**

| Model | 10 runs | 50 runs | 100 runs, projected |
| - | - | - | - |
| mistral-medium-3.5 | 100% | 100% | 100% |
| claude-opus-5 | 100% | 100% | 100% |
| claude-haiku-4.5 | 100% | 100% | 100% |
| gemini-3.8-flash | 100% | 100% | 100% |
| gemini-3.5-flash-lite | 100% | 100% | 100% |
| kimi-k3 | 100% | 100% | 100% |
| gpt-6-astra | 97% | 100% | 100% |
| gpt-5.6-luna | 97% | 100% | 100% |
| grok-4.6 | 88% | 97% | 100% |
| qwen3.8-max | 70% | 91% | 97% |
| deepseek-v4.1-flash | 50% | 66% | 72% |

_Share of 4,000 simulated batches, drawn from each model's 50 real answers, that named the same #1 as all 50. We did not run any model 100 times._

- **mistral-medium-3.5 is settled at 10 runs.** Its #1, top three and top five are all fixed.
- **Most models get their #1 right at 10 runs.** gpt-6-astra, gpt-5.6-luna, kimi-k3, both Geminis and both Claudes name their own #1 correctly 97% to 100% of the time. Their top three need about 50 runs.
- **grok-4.6 needs about 50 runs.** At 10, its #1 is right 88% of the time, because it rates Cartier and Omega closely. At 50, it's 97%.
- **qwen3.8-max needs 50 to 100 runs.** At 10, its #1 is right only 70% of the time, then 91% at 50 and 97% at 100.
- **deepseek-v4.1-flash never settles.** It's right 50% of the time at 10 runs, 66% at 50 and 72% at 100. Its top two are genuinely tied, Omega on 67 and Cartier on 65, so more runs only confirm the tie.

A tie is the honest answer in that last case. When a model scores two brands within a few points of each other, no number of runs will split them. Report it as a tie, not a winner.

---

_What it means_

## Know which models you're asking

Your position isn't one number. It's a separate answer from every model your buyers use.

_01_

**Ten runs give you a solid read on each model's #1.** If your brand sits between #2 and #5, more runs, up to 50 per model, give you a much clearer picture of where it stands.

_02_

**Models from the same lab answer differently.** A strong position with gpt-5.6-luna doesn't carry over to gpt-6-astra, and the sources that got you cited by one may not work for the next. Check each model on its own.

**Where does your brand stand across the models?** Every BrandRanking.AI study asks a panel of models the same question about 10 times each and shows where they agree and where they don't. Browse the [category studies](https://www.brandranking.ai/categories) or request a study for your own market.

- [Browse categories](https://www.brandranking.ai/categories)
- [Request a study](https://www.brandranking.ai/#request-category-study)

---

## How we measured

- **Data.** 550 answers to "Which watch brand would you wear on a first date?": 50 from each of 11 models, collected between 9 and 26 September 2026. Four came back empty. Five answers from a retired model, gemini-3.7-flash, are left out. The comparisons across categories use the 11 category studies from [Note 01](https://www.brandranking.ai/field-notes/how-do-ai-models-rank-brands) plus AI brand visibility tools, with the first 10 answers per model. Every answer is a ranked top five from the model's training alone, with no web search.
- **Rankings.** A brand gets 100 points for a #1, 80 for a #2 and so on down to 20 for a #5, and 0 when it's left out. The points are averaged for each model, then across models with every model counted equally. First places are the share of answers that put a brand #1.
- **Consistency.** A model's usual #1 is the brand it put first most often in a category. Consistency is the share of its answers that named that brand.
- **Redraws.** For the run tests, we redrew each model's answers at random 1,000 times at each size.
- **Runs per model.** For each model, we redrew its 50 watch answers at random, with replacement, 4,000 times at 10, 50 and 100 answers, and compared each redraw with the ranking from all 50. No model was run 100 times: the 100-answer figures are this resampling projected past the data, and hold only if the model keeps answering as it did. The margin is 1.96 times the spread of the gap between two brands in a single answer, divided by the square root of the number of runs. Only watches has 50 runs per model, so these figures are for watches only.

_BrandRanking.AI © 2026 · Study data: [brandranking.ai/categories](https://www.brandranking.ai/categories)_
