Your Brand has an AI reputation and it's shaping purchase decisions right now.

Swipe up to continue

Category · CRM Software

What CRM software would you recommend?

9 AI models · 8 answers per model · ranked by score

ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0

  1. #1HubSpot
    98
    brand score
    94first choice
    12345
  2. #2Salesforce
    82
    brand score
    56first choice
  3. #3Zoho
    54
    brand score
    31first choice
  4. #4Pipedrive
    39
    brand score
    25first choice
  5. #5Microsoft Dynamics 365
    21
    brand score
    16first choice
Show 4 more
  1. #6Freshworks
    4
    brand score
    4first choice
  2. #7Freshsales
    2
    brand score
    2first choice
  3. #8monday.com
    0
    brand score
    0first choice
  4. #9Zendesk
    0
    brand score
    0first choice

Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.

Swipe left for the full report

Category · CRM Software

Model by model

Median rank each model gave each brand.

Scroll the table sideways to see every model →

claude-opus-5deepseek-v4.1-flashgemini-3.7-flashgpt-5.6-lunagpt-5.6-solgpt-5.6-terragpt-6-astrakimi-k3mistral-medium-3.5
HubSpot1.01.01.02.01.01.01.01.01.0
Salesforce2.02.02.01.02.02.02.02.02.0
Zoho4.03.53.04.03.03.03.03.03.0
Pipedrive3.04.04.05.04.54.04.04.04.0
Microsoft Dynamics 3655.05.03.04.05.05.05.0
Freshworks5.05.05.05.0
Freshsales5.0
monday.com5.0
Zendesk5.0
1st
2
3
4
5th

A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.

Category · CRM Software

Where the models disagree

0 of 9 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.

  1. #5Microsoft Dynamics 365

    gemini-3.7-flash 0 gpt-5.6-luna 60 · 2 of 9 never named it

  2. #4Pipedrive

    gpt-5.6-luna 20 claude-opus-5 53

  3. #3Zoho

    gpt-5.6-luna 40 mistral-medium-3.5 60

  4. #6Freshworks

    claude-opus-5 0 gemini-3.7-flash 18 · 5 of 9 never named it

  5. #1HubSpot

    gpt-5.6-luna 88 mistral-medium-3.5 100

  6. #2Salesforce

    claude-opus-5 80 gpt-5.6-luna 93

A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.

Category · CRM Software

Named, and named first

Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.

the default answera narrow favouritelisted, rarely ledthe long tail12345678
  1. 1HubSpot100/89
  2. 2Salesforce100/11
  3. 3Zoho100/0
  4. 4Pipedrive100/0
  5. 5Microsoft Dynamics 36568/0
  6. 6Freshworks18/0
  7. 7Freshsales11/0
  8. 8monday.com1/0

The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.

Category · CRM Software

What the models actually said

Models
9
Answers
72

HubSpot leads with a score of 98 of 100, ranked by 9 of 9 models.

Across the nine models this category resolves into a stable pecking order — HubSpot, then Salesforce, then Zoho, then Pipedrive — with the disagreement concentrated at the two ends rather than the middle.

The top of the ranking is where the models are most alike. HubSpot is the clear consensus pick, ranked first by seven of nine, and the only real dissent is gpt-5.6-luna, which parks it at a median rank of 2.

HubSpot

per-model scores 88100 of 100 · mean 98 across 9 models

That same model is the mirror image on Salesforce: gpt-5.6-luna is alone in leading with it, scoring it 92.5 where six other models sit flat at 80. So the HubSpot-versus-Salesforce question is effectively decided by one model's ordering, not by any broad divide — the two brands trade first and second place almost entirely inside gpt-5.6-luna's answers.

HubSpotSalesforce
  • 100claude-opus-580
  • 100gemini-3.7-flash80
  • 100gpt-5.6-sol80
  • 100gpt-5.6-terra80
  • 100gpt-6-astra80
  • 100mistral-medium-3.580
  • 98deepseek-v4.1-flash83
  • 95kimi-k385
  • 88gpt-5.6-luna93

HubSpot ahead on 8 of 9 models

The genuinely wide splits appear lower down. Microsoft Dynamics 365 carries the largest disagreement in the field: gpt-5.6-luna rates it at borda 60 and a median rank of 3, while claude-opus-5, gpt-6-astra and kimi-k3 land it at rank 5, and two models leave it off some runs entirely. The divide is about weighting, not fact — models agree it is enterprise-grade but disagree on how heavily licensing and implementation complexity should count against a general recommendation.

gpt-5.6-luna
60
rates it highest
gemini-3.7-flash
0
never named it

Microsoft Dynamics 365 · 60 points apart on a 0–100 scale

Below Dynamics, the tail is defined by coverage rather than conviction. Freshworks appears for only four models, and its spread is driven almost entirely by gemini-3.7-flash naming it in most runs while three others cite it rarely. Freshsales, monday.com and Zendesk are each the property of a single model — mistral-medium-3.5, gpt-5.6-terra and gemini-3.7-flash respectively — so their low, "agreed" scores mostly reflect shared silence rather than a shared judgment.

The recalled sources point the same way for everyone and do little to explain the splits. G2www.g2.com/categories/crm×116 dominates the entire category, with Capterra and PCMag close behind, and the brand-specific pages track wherever a model went. Because the heaviest-cited sources are common across models, they are unlikely to be what separates gpt-5.6-luna from the rest; that divergence looks like a difference in how the model weighs enterprise capability against ease of adoption — a hypothesis the reasonings support but the source data does not confirm.

The panel agrees on the roster and its rough order; the real disagreement is one model's willingness to reward raw enterprise capability over turnkey ease.
Category · CRM Software

What shaped the answers

54 sources across 992 references, grouped by site from 231 recalled names. The top 5 carry 62% of them.

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · CRM Software
#1

HubSpot

also named HubSpot CRM

98
brand score
94first choice
12345

Summary

HubSpot is the panel's clearest consensus pick, carrying a Borda score of 97.78 and ranked by all 9 of 9 models. The agreement is tight: per-model scores span 87.5 to 100 with a standard deviation of 3.99, which the report labels "models agree." Seven of the nine models placed it at a median rank of 1, and only gpt-5.6-luna diverges meaningfully, settling at a median rank of 2 and giving it the lowest score in the set.

HubSpot

per-model scores 88100 of 100 · mean 98 across 9 models

The reasoning behind the ranking is remarkably uniform. Models repeatedly cite the same three attributes — an approachable, intuitive interface, a genuinely useful free tier, and integrated marketing, sales, and service tools — while consistently attaching the same caveat that costs rise steeply as contacts, seats, and advanced features are added. Several models frame it explicitly as the safe default recommendation when requirements are unknown. The recalled sources cluster under user-review and directory themes, most prominently G2www.g2.com/categories/crm×116 and the vendor's own HubSpot CRMwww.hubspot.com/products/crm×19, supplemented by Capterra, PCMag, and Gartner references.

HubSpot offers the best balance of usability, a genuinely useful free tier, and a unified suite spanning marketing, sales and service, which makes it the safest general recommendation for small and mid-sized teams.

claude-opus-5

HubSpot is the strongest general recommendation because it combines an approachable interface, a useful free tier, and well-integrated sales, marketing, and service tools.

gpt-5.6-sol

Per model summary

  • claude-opus-5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    8/8

    Consistently frames HubSpot as the safest general recommendation for small and mid-sized teams, citing its usability, useful free tier, unified marketing/sales/service suite, and best-in-class onboarding, while noting costs escalate as contacts and add-ons grow.

  • deepseek-v4.1-flash
    brand score
    98
    first choice
    94
    median rank
    1.0
    present in
    8/8

    Recommends HubSpot first for most small and mid-sized teams because of its intuitive interface, useful free tier, fast setup, and integrated sales/marketing/service tools, with the caveat that costs and customization limits appear as usage scales.

  • gemini-3.7-flash
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    8/8

    Emphasizes HubSpot's exceptionally intuitive interface, robust free tier, and seamless all-in-one ecosystem across marketing, sales, and service, presenting it as the best-balanced, scalable choice despite expensive enterprise tiers.

  • gpt-5.6-sol
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    8/8

    Repeatedly calls HubSpot the strongest general recommendation for its approachable interface, useful free tier, and integrated marketing/sales/service ecosystem, tempered by substantially rising costs as features and contacts grow.

  • gpt-5.6-terra
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    8/8

    Describes HubSpot as the strongest all-around recommendation for small and midsize teams due to its approachable CRM, useful free tier, and natural integration with marketing and service workflows, with the trade-off of rising costs from added hubs, seats, and automation.

  • gpt-6-astra
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    8/8

    Treats HubSpot as its default recommendation when company size and requirements are unknown, citing its approachable CRM connected to sales, marketing, and service tools and easy start, while warning costs rise substantially with advanced features and expanded usage.

  • kimi-k3
    brand score
    95
    first choice
    88
    median rank
    1.0
    present in
    8/8

    Recommends HubSpot as the best overall balance of usability and capability for most businesses, stressing its intuitive interface, genuinely useful free tier, fast adoption, and integrated hubs, while noting costs climb steeply at higher tiers.

  • mistral-medium-3.5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    8/8

    Consistently describes HubSpot as a highly intuitive, user-friendly CRM with a robust free tier and seamless integration of marketing, sales, and service tools, ideal for small to mid-sized and scaling businesses.

  • gpt-5.6-luna
    brand score
    88
    first choice
    69
    median rank
    2.0
    present in
    8/8

    Positions HubSpot as an approachable top pick for small and midsize businesses, highlighting its free tier and unified marketing, sales, service, and content tools, while noting costs rise as advanced features and contacts are added.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-55 sites · 24 references
G2×8 · 33%

g2.com

deepseek-v4.1-flash7 sites · 24 references
G2×8 · 33%

g2.com

gemini-3.7-flash8 sites · 24 references
gpt-5.6-sol5 sites · 24 references
gpt-5.6-terra4 sites · 24 references
HubSpot CRM×9 · 38%

hubspot.com

gpt-6-astra2 sites · 12 references
HubSpot CRM×9 · 75%

hubspot.com

kimi-k36 sites · 24 references
G2×8 · 33%

g2.com

mistral-medium-3.54 sites · 24 references
gpt-5.6-luna3 sites · 20 references

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · CRM Software
#2

Salesforce

82
brand score
56first choice
12345

Summary

Salesforce scores strongly across the panel, drawing a Borda of 82.22 of 100 and recognition from all 9 of the 9 models. The spread is narrow — per-model borda runs from 80 to 92.5 with a standard deviation of 3.99, which the report labels as agreement. That consensus is thematic as well as numerical: nearly every model frames Salesforce as the enterprise standard, praised for unmatched customization, its AppExchange ecosystem, and depth of automation and reporting, while placing it just behind HubSpot for general buyers because of implementation complexity, administrative overhead, and total cost of ownership. The recalled evidence sits under those themes consistently, with G2×1 appearing across most models, reinforced by Gartner's Magic Quadrant citations and Salesforce's own product pages.

Salesforce

per-model scores 8093 of 100 · mean 82 across 9 models

The main point of divergence is placement at the top rather than presence: gpt-5.6-luna ranks it first in 5 of 8 runs for a borda of 92.5, whereas claude-opus-5, gemini-3.7-flash, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra, and mistral-medium-3.5 each land at 80, never ranking it first. The first-choice score of 55.56 of 100 reflects that split between models that lead with Salesforce and those that reserve first place for a more turnkey option.

Salesforce is the strongest overall recommendation for organizations that need a highly capable, scalable CRM with extensive customization, automation, analytics, and integrations.

gpt-5.6-luna

It ranks lower than HubSpot here only because of higher cost, implementation complexity and admin overhead for smaller teams.

claude-opus-5

Per model summary

  • gpt-5.6-luna
    brand score
    93
    first choice
    81
    median rank
    1.0
    present in
    8/8

    Frames Salesforce as the strongest overall choice for scalable, highly customizable CRM with broad automation, analytics, and integrations. Its breadth and ecosystem suit larger organizations, but implementation complexity and cost can be significant for smaller teams.

  • claude-opus-5
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    8/8

    Consistently frames Salesforce as the market leader with unmatched customization, AppExchange ecosystem, and enterprise-grade reporting, ideal for large or complex organizations. It ranks below HubSpot mainly due to cost, implementation complexity, and admin overhead that make it overkill for smaller teams.

  • deepseek-v4.1-flash
    brand score
    83
    first choice
    56
    median rank
    2.0
    present in
    8/8

    Repeatedly describes Salesforce as the most powerful and customizable CRM with an unmatched ecosystem for complex enterprise sales. It ranks it lower because of the steep learning curve, high cost, and the need for dedicated admins, making it less turnkey for smaller teams.

  • gemini-3.7-flash
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    8/8

    Consistently calls Salesforce the industry standard/benchmark for enterprise customization, analytics, and AppExchange integrations. It places it just behind HubSpot due to steep learning curve, complex setup, administrative overhead, and high total cost of ownership.

  • gpt-5.6-sol
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    8/8

    Uniformly emphasizes Salesforce's exceptional customization, automation, reporting, and ecosystem for complex or enterprise deployments. It ranks below HubSpot for general audiences because implementation, administration, and total cost can be substantial.

  • gpt-5.6-terra
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    8/8

    Consistently positions Salesforce as excellent for larger or complex organizations needing deep customization, integrations, governance, and a broad partner ecosystem. Its power comes with higher implementation effort, administration, and total cost than simpler options.

  • gpt-6-astra
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    8/8

    Repeatedly recommends Salesforce for organizations needing extensive customization, sophisticated sales workflows, and a large integration ecosystem. It ranks it below HubSpot for general buyers because implementation, administration, and total cost can be excessive for simpler needs.

  • kimi-k3
    brand score
    85
    first choice
    63
    median rank
    2.0
    present in
    8/8

    Consistently describes Salesforce as the industry standard and most powerful, customizable CRM with an unmatched AppExchange ecosystem. It ranks it second because licensing cost, admin overhead, and implementation complexity make it overkill for many smaller teams.

  • mistral-medium-3.5
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    8/8

    Uniformly frames Salesforce as the enterprise-level industry leader with unparalleled customization, scalability, and ecosystem. It notes a steeper learning curve and complexity that can make it overkill for smaller businesses.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gpt-5.6-luna3 sites · 20 references
claude-opus-57 sites · 24 references
deepseek-v4.1-flash6 sites · 24 references
G2×8 · 33%

g2.com

gemini-3.7-flash7 sites · 23 references
gpt-5.6-sol7 sites · 24 references
gpt-5.6-terra5 sites · 24 references
G2 CRM Software Category×8 · 33%

g2.com

gpt-6-astra3 sites · 12 references
Salesforce CRM×8 · 67%

salesforce.com

  • ×6/crm
  • ×1Salesforce /
  • ×1Salesforce Sales Cloud overview /sales
kimi-k39 sites · 24 references
G2×7 · 29%

g2.com

mistral-medium-3.55 sites · 24 references
G2×8 · 33%

g2.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · CRM Software
#3

Zoho

also named Zoho CRM

54
brand score
31first choice
12345

Summary

Zoho occupies a firmly mid-tier position in this study, ranked by all 9 of the 9 models and earning a Borda score of

Stat unavailable

with no model placing it first. The spread of per-model scores is narrow — from 40 to 60, a standard deviation of 6.64 — which the report labels as agreement. That consensus rests on a shared reading: models consistently frame Zoho as the strongest value-for-money option, feature-rich and tightly bound to the wider Zoho suite, well-suited to budget-conscious SMBs, while noting an interface and configuration experience that feels less polished than higher-ranked alternatives. The clustering of median ranks around 3 and 4 reflects that this trade-off caps its standing rather than sinking it.

Zoho

per-model scores 4060 of 100 · mean 54 across 9 models

The recalled sources sit largely under review-aggregator and vendor themes, with G2×1 and Capterra×4 recurring across nearly every model alongside Zoho's own pages, and PCMag×4 supporting the value-and-features framing. The value case and the polish caveat surface together throughout the reasonings.

Zoho CRM delivers the strongest feature-per-dollar value, with automation, analytics and a huge suite of integrated business apps at a fraction of competitors' prices.

claude-opus-5

Zoho provides broad CRM functionality and strong value, especially for small and midsize businesses that may also use its wider business-software suite.

gpt-5.6-sol

Per model summary

  • gemini-3.7-flash
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    8/8

    Emphasizes an exceptionally feature-rich, cost-effective CRM with native integration into the broader Zoho ecosystem, offset by an interface that can feel cluttered, fragmented, or disjointed.

  • gpt-5.6-sol
    brand score
    58
    first choice
    32
    median rank
    3.0
    present in
    8/8

    Consistently characterizes Zoho as strong value with broad functionality tied to its wider business-software suite for SMBs, while noting the interface and configuration feel less polished than top-ranked options.

  • gpt-5.6-terra
    brand score
    58
    first choice
    32
    median rank
    3.0
    present in
    8/8

    Highlights Zoho's strong value and broad CRM capability, especially for cost-conscious teams using its wider suite, while noting the interface and setup feel less polished than leading alternatives.

  • gpt-6-astra
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    8/8

    Ranks Zoho third as an affordable, value-oriented choice for budget-conscious SMBs with a broad business-software ecosystem, noting that configuring across its many options and applications takes more effort than simpler alternatives.

  • kimi-k3
    brand score
    58
    first choice
    32
    median rank
    3.0
    present in
    8/8

    Consistently frames Zoho as the best value with a broad feature set (AI, automation, multichannel) at a fraction of leaders' prices, ranked mid-pack because its interface, polish, and ecosystem trail HubSpot and Salesforce.

  • mistral-medium-3.5
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    8/8

    Repeatedly describes Zoho as a cost-effective, feature-rich option with automation and AI, well-suited to small businesses and startups, though possibly lacking advanced features for larger enterprises.

  • deepseek-v4.1-flash
    brand score
    50
    first choice
    29
    median rank
    3.5
    present in
    8/8

    Repeatedly positions Zoho as excellent value with a broad, customizable feature set for budget-conscious SMBs, while noting its interface can feel less polished or cluttered and the wider ecosystem sprawling.

  • claude-opus-5
    brand score
    48
    first choice
    28
    median rank
    4.0
    present in
    8/8

    Consistently frames Zoho as the best feature-per-dollar value, integrated with the broad Zoho One suite and suited to budget-conscious SMBs and international teams, while noting a less polished UI and inconsistent support relative to HubSpot and Salesforce.

  • gpt-5.6-luna
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    8/8

    Describes Zoho as a broad, cost-effective, value-oriented CRM with a large application ecosystem, ranked below leaders because its interface and administration experience feel less polished or streamlined.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gemini-3.7-flash4 sites · 23 references
G2×8 · 35%

g2.com

gpt-5.6-sol6 sites · 24 references
gpt-5.6-terra4 sites · 24 references
Zoho CRM×9 · 38%

zoho.com

gpt-6-astra2 sites · 12 references
Zoho CRM×11 · 92%

zoho.com

kimi-k37 sites · 24 references
PCMag×7 · 29%

pcmag.com

mistral-medium-3.55 sites · 24 references
deepseek-v4.1-flash5 sites · 24 references
G2×8 · 33%

g2.com

claude-opus-57 sites · 24 references
G2×8 · 33%

g2.com

gpt-5.6-luna3 sites · 20 references
G2×8 · 40%

g2.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · CRM Software
#4

Pipedrive

39
brand score
25first choice
12345

Summary

Pipedrive scores a Borda average of 38.61 of 100 and is ranked by all 9 of the 9 models, a consistent mid-table placement rather than a top pick — none of the 9 models ranked it first. With per-model borda spanning 20 to 52.5 and a standard deviation of 8.51, the report labels this "broad agreement," and the spread reflects fine gradations rather than genuine dispute: claude-opus-5 places it highest at 52.5, while gpt-5.6-luna anchors the bottom at 20.

Pipedrive

per-model scores 2053 of 100 · mean 39 across 9 models

The models converge on a single reading, with the theme steady across all nine: Pipedrive is a sales-first, visually-driven pipeline CRM that small teams adopt quickly with low overhead, held back only by comparatively thin marketing, service, reporting, and enterprise capabilities. That shared framing — strong at its narrow purpose, weaker as an all-in-one platform — explains why it clusters at rank 4 for most models while gpt-6-astra notes it "could be my first choice for a small, sales-focused team." The recalled evidence sits mainly on review aggregators and the vendor's own pages, with G2×1 and Capterra×4 recurring most heavily alongside frequent citations of Pipedrive's own site.

Pipedrive is a sales-first, pipeline-centric CRM that small sales teams can adopt in days, with clear visual deal stages and low pricing.

claude-opus-5

Pipedrive excels in simplicity and ease of use, focusing on sales pipeline management.

mistral-medium-3.5

Per model summary

  • claude-opus-5
    brand score
    53
    first choice
    30
    median rank
    3.0
    present in
    8/8

    Pipedrive is consistently framed as a sales-first, visual pipeline CRM that small teams adopt quickly with low overhead and affordable pricing, ranked lower because its marketing, service, and reporting capabilities are comparatively thin.

  • deepseek-v4.1-flash
    brand score
    40
    first choice
    26
    median rank
    4.0
    present in
    8/8

    Pipedrive is portrayed as a focused, easy-to-use pipeline CRM ideal for small sales teams, ranked lower because it lacks the broader marketing, service, and customization depth of all-in-one platforms.

  • gemini-3.7-flash
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    8/8

    Pipedrive is described as purpose-built around visual pipeline management for deal-driven sales teams with minimal onboarding, ranked lower for lacking native marketing and customer support features found in full-suite platforms.

  • gpt-5.6-terra
    brand score
    43
    first choice
    26
    median rank
    4.0
    present in
    8/8

    Pipedrive is framed as a strong fit for sales-led teams wanting a clean visual pipeline and quick adoption without enterprise complexity, ranked lower for limited native marketing, service, and cross-functional capabilities.

  • gpt-6-astra
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    8/8

    Pipedrive is portrayed as excellent for sales teams prioritizing pipeline visibility and deal tracking, ranked fourth for general use because it is less comprehensive for marketing and customer-service needs, though it could top a small sales team's list.

  • kimi-k3
    brand score
    43
    first choice
    26
    median rank
    4.0
    present in
    8/8

    Pipedrive is described as a sales-first CRM built around a clean visual pipeline that reps enjoy using, ranked lower because its marketing, service, and reporting features are thin and often rely on add-ons.

  • mistral-medium-3.5
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    8/8

    Pipedrive is consistently characterized by simplicity and ease of use for sales-focused teams, with the caveat that it offers fewer marketing and service features than competitors like HubSpot or Salesforce.

  • gpt-5.6-sol
    brand score
    30
    first choice
    23
    median rank
    4.5
    present in
    8/8

    Pipedrive is consistently described as easy to adopt and effective for visual, sales-focused pipeline management, ranked lower because its marketing, customer-service, and enterprise capabilities are narrower than broader platforms.

  • gpt-5.6-luna
    brand score
    20
    first choice
    20
    median rank
    5.0
    present in
    8/8

    Pipedrive is presented as an intuitive, sales-focused CRM strong on visual pipeline and deal management, ranked lower because it is less comprehensive for marketing, service, analytics, and enterprise needs.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-55 sites · 24 references
G2×8 · 33%

g2.com

deepseek-v4.1-flash5 sites · 24 references
gemini-3.7-flash6 sites · 23 references
gpt-5.6-terra3 sites · 24 references
Pipedrive×10 · 42%

pipedrive.com

gpt-6-astra2 sites · 12 references
Pipedrive×9 · 75%

pipedrive.com

kimi-k37 sites · 22 references
G2×8 · 36%

g2.com

mistral-medium-3.54 sites · 24 references
Capterra×8 · 33%

capterra.com

gpt-5.6-sol4 sites · 24 references
Pipedrive×9 · 38%

pipedrive.com

gpt-5.6-luna3 sites · 20 references
G2×8 · 40%

g2.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · CRM Software
#5

Microsoft Dynamics 365

also named Microsoft, Microsoft Dynamics, Microsoft Dynamics 365 Sales

21
brand score
16first choice
12345

Summary

Microsoft Dynamics 365 scores modestly across the panel, with a Borda average of 20.56 of 100 and a first-choice score of 15.86, ranked by 7 of the 9 models but placed first by none. The models converge on a shared narrative more than the numbers alone suggest: Dynamics 365 is consistently framed as a capable, enterprise-grade CRM whose natural home is organizations already committed to Microsoft 365, Teams, Azure and the Power Platform, but one that ranks lower as a general recommendation because of complex licensing, partner-led implementation, and a steeper adoption curve. Where the models differ is in how much weight that complexity carries: gpt-5.6-luna places it at a median rank 3 with a borda of 60, while claude-opus-5, gpt-6-astra and kimi-k3 all settle at a median rank 5, several of them explicit that the placement reflects ease of adoption rather than capability.

Microsoft Dynamics 365
12345

That gap defines the spread the report labels "broad agreement" despite a standard deviation of 16.78. The evidence behind these judgments leans on vendor and review sources — Microsoft Dynamics 365 Saleswww.microsoft.com/en-us/dynamics-365/products/sales×16 recurs across models alongside G2 review pages and Gartner's Magic Quadrant for Sales Force Automation Platforms — supporting a consistent picture of a powerful platform whose value is conditional on ecosystem fit.

Microsoft Dynamics 365

per-model scores 060 of 100 · mean 21 across 9 models

Dynamics 365 is widely seen as strong for Microsoft-centric enterprises but heavy for a general buyer, which caps its ranking without contesting its capability.

It ranks fifth as a default recommendation because configuration and licensing can be complex, though it could rank much higher for a Microsoft-centric enterprise.

gpt-6-astra

I rank it last here mainly because it's the least broadly applicable of the five, not because of capability.

claude-opus-5

Per model summary

  • gpt-5.6-luna
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    8/8

    Consistently describes it as a strong, flexible enterprise CRM for companies already invested in Microsoft 365, Teams, Azure and Power Platform, but notes its configuration, licensing, and expertise requirements make it more complex than simpler options.

  • gpt-5.6-sol
    brand score
    28
    first choice
    19
    median rank
    4.0
    present in
    6/8

    Repeatedly frames it as a compelling choice for organizations deeply invested in the Microsoft ecosystem, but ranks it lower (fourth or fifth) as a general recommendation due to licensing and implementation complexity, especially for smaller teams.

  • claude-opus-5
    brand score
    20
    first choice
    20
    median rank
    5.0
    present in
    8/8

    Consistently frames Dynamics 365 as a powerful enterprise option best suited for organizations already invested in Microsoft 365, Teams, Azure and Power Platform, but ranks it last due to complex licensing, partner-led implementation, and lower ease of adoption—not capability.

  • deepseek-v4.1-flash
    brand score
    25
    first choice
    18
    median rank
    5.0
    present in
    6/8

    Repeatedly positions it as a robust CRM ideal for organizations already invested in the Microsoft ecosystem, while ranking it low generally because of complex, costly licensing and implementation that feels heavier than simpler alternatives.

  • gpt-5.6-terra
    brand score
    13
    first choice
    13
    median rank
    5.0
    present in
    5/8

    Consistently presents it as a sensible option for larger organizations invested in Microsoft 365, Teams, Power Platform and Azure, but ranks it lower for general use because implementation and administration are complex and often require specialist or partner support.

  • gpt-6-astra
    brand score
    20
    first choice
    20
    median rank
    5.0
    present in
    8/8

    Repeatedly ranks it fifth as a general-purpose default due to configuration and licensing complexity, while emphasizing it could rank much higher for Microsoft-centric enterprises where its ecosystem fit is a strong advantage.

  • kimi-k3
    brand score
    20
    first choice
    20
    median rank
    5.0
    present in
    8/8

    Consistently describes it as a powerful, enterprise-grade CRM with deep Microsoft ecosystem integration, but ranks it last because of confusing licensing, high implementation cost, steep learning curve, and limited value outside Microsoft-centric organizations.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gpt-5.6-luna4 sites · 20 references
gpt-5.6-sol8 sites · 18 references
claude-opus-56 sites · 24 references
deepseek-v4.1-flash8 sites · 18 references
G2×5 · 28%

g2.com

gpt-5.6-terra5 sites · 15 references
Microsoft Dynamics 365 Sales×7 · 47%

microsoft.com

gpt-6-astra2 sites · 11 references
Microsoft Dynamics 365 Sales×9 · 82%

microsoft.com

kimi-k311 sites · 23 references
Gartner Magic Quadrant for Sales Force Automation×4 · 17%

no URL recalled

  • ×4named without a URL

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · CRM Software

How this was measured

The question

  • Unaided brand recommendation question: “What CRM software would you recommend?
  • The prompt asks each model to return exactly 5 brands, ranked 1 to 5, and for each one a reason for the recommendation and the sources that informed it.

Sampling

  • Models are not deterministic, so the question is asked over and over — 8 answers per model
  • Spellings of the same brand are normalized to the most commonly used form before counting

Measures of Position: Median and Quartiles

  • Median value indicates that in 50% of answers the brand held this position or higher.
  • Quartiles help visualize the spread of rankings per brand. Q1 indicates that 25% of answers had this rank or higher. Q3 means that 75% of answers ranked the brand as X or better.
  • Medians and quartiles are computed only from answers where the brand was present.

Coverage

  • Count of answers in which the given brand was named, as a percentage.

Brand score - Normalized Borda score

  • The brand score is calculated using a normalized Borda score. Borda count is a voting method: each ballot awards points by position instead of naming one winner. Our ranking responses from the LLM are always fixed to 5 answers, so we assign 100 points for rank 1, then 80, 60, 40, 20 — and 0 if the brand is missing.
Brand score=100ni=1n6ri5\text{Brand\ score} = \frac{100}{n} \sum_{i=1}^{n} \frac{6 - r_i}{5}

where

ri={1,,5if the brand appears6if the brand is not mentionedr_i = \begin{cases} 1,\dots,5 & \text{if the brand appears} \\ 6 & \text{if the brand is not mentioned} \end{cases}
  • This way we can calculate a common score for all brands mentioned across model responses. The score reflects both how high the brand was ranked and how frequently it was mentioned.

First choice score - adjusted MRR - top-rank indicator

  • The first choice score is based on Mean Reciprocal Rank (a common search-engine measure), but adapted to measure the position of a specific brand across repeated ranked responses.
  • Each occurrence receives a reciprocal-position score: 100 for 1st, 50 for 2nd, 33.3 for 3rd, 25 for 4th, 20 for 5th, and 0 when the brand is not mentioned. The scores are averaged across all responses.
First choice score=100ni=1nsi\text{First\ choice\ score} = \frac{100}{n} \sum_{i=1}^{n} s_i

where

si={1/riif the brand is present0if the brand is absents_i = \begin{cases} 1/r_i & \text{if the brand is present} \\ 0 & \text{if the brand is absent} \end{cases}
  • No rank below 1st place gets more than 50, so the score is heavily driven by first places.
Category · Consulting companies for AI transformation

Which consulting firm should a Fortune 500 company hire for an AI transformation?

6 AI models · 5 answers per model · ranked by score

ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0

  1. #1Accenture
    91
    brand score
    82first choice
    12345
  2. #2McKinsey & Company
    79
    brand score
    60first choice
  3. #3Boston Consulting Group
    53
    brand score
    32first choice
  4. #4Deloitte
    51
    brand score
    32first choice
  5. #5IBM Consulting
    15
    brand score
    13first choice
Show 1 more
  1. #6Bain & Company
    11
    brand score
    8first choice

Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.

Swipe left for the full report

Category · Consulting companies for AI transformation

Model by model

Median rank each model gave each brand.

Scroll the table sideways to see every model →

claude-opus-5gemini-3.7-flashgpt-5.6-lunagpt-5.6-solgpt-5.6-terrakimi-k3
Accenture2.03.01.01.01.01.0
McKinsey & Company1.01.02.03.02.02.0
Boston Consulting Group4.02.04.04.04.03.0
Deloitte3.05.03.02.03.04.0
IBM Consulting5.05.05.05.05.05.0
Bain & Company4.05.05.04.0
1st
2
3
4
5th

A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.

Category · Consulting companies for AI transformation

Where the models disagree

0 of 6 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.

  1. #4Deloitte

    gemini-3.7-flash 28 gpt-5.6-sol 72

  2. #2McKinsey & Company

    gpt-5.6-sol 52 gemini-3.7-flash 100

  3. #1Accenture

    gemini-3.7-flash 60 gpt-5.6-terra 100

  4. #3Boston Consulting Group

    gpt-5.6-terra 40 gemini-3.7-flash 80

  5. #6Bain & Company

    claude-opus-5 0 gemini-3.7-flash 28 · 2 of 6 never named it

  6. #5IBM Consulting

    gemini-3.7-flash 4 gpt-5.6-sol 28

A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.

Category · Consulting companies for AI transformation

Named, and named first

Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.

the default answera narrow favouritelisted, rarely ledthe long tail123456
  1. 1Accenture100/70
  2. 2McKinsey & Company100/30
  3. 3Boston Consulting Group100/0
  4. 4Deloitte100/0
  5. 5IBM Consulting63/0
  6. 6Bain & Company37/0

The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.

Category · Consulting companies for AI transformation

What the models actually said

Models
6
Answers
30

Accenture leads with a score of 91 of 100, ranked by 6 of 6 models.

Across the six models, the top of the field is settled and the disagreements concentrate in the middle. Accenture and McKinsey trade the top two spots, while Boston Consulting Group, Deloitte, IBM Consulting, and Bain sort into a more contested lower tier where individual models diverge noticeably.

Where the models agree

The clearest consensus is on the extremes. Accenture holds a category-wide median of 1 and IBM Consulting a rock-solid median of 5 (Q1 5, Q3 5) with every model landing on exactly 5. IBM's uniformity extends to reasoning: all six credit engineering depth, watsonx, and regulated-environment fit, and five of the six cite vendor bias toward IBM's own stack as the reason it sits mid-tier. There is little to separate the models on either brand.

McKinsey is also broadly agreed on in substance — C-suite strategy credibility and QuantumBlack depth recur everywhere — even though its placement varies by a rank or two.

The main splits

The sharpest divergences are best read model by model:

  • gemini-3.7-flash is the standout dissenter. It ranks Accenture lowest of any model (median 3 vs. 1 for four peers) and BCG highest (median 2), while pushing Deloitte to the bottom (median 5). Its framing of Accenture as an "integration and infrastructure play rather than a strategy leader" is consistent with this ordering, elevating the strategy-first houses over the large integrator.
  • gpt-5.6-sol is the mirror image on Deloitte, ranking it highest of all models (median 2) but placing McKinsey lowest (median 3, Q1 3, Q3 4). It reads Deloitte's regulated-enterprise breadth favorably while adding the most caveats to both Accenture (scope, cost, complexity) and McKinsey (delivery bench).
  • claude-opus-5 ranks McKinsey top (median 1) and Accenture second, the reverse of the four gpt/kimi models, framing Accenture as execution-led rather than a strategy leader.
  • kimi-k3 is the most skeptical of Deloitte among the mid-rankers (median 4), recasting its breadth as "generalist and integration-led."

BCG shows the widest disagreement of any brand: gemini places it 2nd, kimi 3rd, and the remaining four settle at 4. Deloitte spans an even wider model range (2 through 5) despite a tidy 3.5 category median — a case where the aggregate obscures real disagreement.

Plausible source patterns behind the splits

The source data is suggestive rather than conclusive, and these links are hypotheses.

The most legible pattern is on Accenture: models ranking it 1st lean on first-party material (Accenture service pages, Technology Vision, the $3B investment newsroom item), whereas gemini-3.7-flash — the lowest ranker — is described as drawing almost entirely on external outlets (Gartner, IDC MarketScape, Bloomberg, FT, WSJ) with no first-party Accenture citations. It is plausible that a reliance on outside analyst and press framing produced its more measured, integrator-oriented view, but the data shows correlation, not cause.

A similar hypothesis fits Deloitte: gemini is again noted as leaning more on Gartner, IDC MarketScape, Bloomberg, FT, and Fortune, coinciding with its lowest placement — consistent with a pattern of external sourcing tracking cooler positioning, though the sample is too small to confirm.

For McKinsey and BCG, the recalled sources are heavily first-party (State of AI, QuantumBlack; BCG X, AI Radar) across models regardless of rank, so the sourcing does not obviously explain why claude and gemini rate them higher than the gpt models. Here the placement differences appear to come from how each model weighs strategy versus at-scale delivery rather than from distinct source pools.

Caveat

Bain and IBM rest on fewer ranked answers (11 and 19) than the leaders (30), so their model-level medians are less stable. All sources are recalled by the models from memory rather than verified citations, which limits how far any source-to-view link can be pushed.

Category · Consulting companies for AI transformation

What shaped the answers

91 sources across 428 references, grouped by site from 256 recalled names. The top 5 carry 44% of them.

McKinsey State of AI×49 · 11%

mckinsey.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Consulting companies for AI transformation
#1

Accenture

91
brand score
82first choice
12345

Summary

Accenture lands at the top of the field, with a median rank of 1 (Q1 1, Q3 2) across 30 ranked answers from 6 models. Agreement is strong at the high end: four models — gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, and kimi-k3 — place it at a median rank of 1 with no spread (Q1 1, Q3 1), while claude-opus-5 sits slightly lower at median 2 (Q1 1, Q3 2) and gemini-3.7-flash is the clear outlier at median 3 (Q1 3, Q3 3). The recurring theme uniting the top rankings is scale of delivery and execution: strategy-through-implementation breadth, cloud and integration depth, a multi-billion-dollar AI investment, and hyperscaler alliances. Several models frame Accenture as the strongest default for moving beyond pilots to enterprise-wide deployment, though most also qualify that its strength skews toward execution rather than top-table strategy — a point noted explicitly by kimi-k3, claude-opus-5, and gemini-3.7-flash, with gpt-5.6-sol adding caveats around scope, cost, and complexity.

The sources behind these themes divide along two lines. Models placing Accenture at rank 1 lean heavily on first-party material — Accenture's own AI and data-AI service pages, Technology Vision 2024, the Reinvention in the Age of Generative AI insight, annual reports, and the $3 billion AI investment newsroom announcement — supplemented by third-party analysts such as Gartner, Everest Group, IDC, Forrester, and the Stanford AI Index. claude-opus-5 grounds its execution-focused view in earnings-release and investor material on generative AI bookings alongside NVIDIA/Microsoft alliance announcements, Gartner's Magic Quadrant, and HFS/Everest Group assessments. By contrast, gemini-3.7-flash, the lowest ranker, draws almost entirely on external outlets — Gartner, IDC MarketScape, Bloomberg, Financial Times, Forrester, and the Wall Street Journal — with no first-party Accenture citations, which aligns with its more measured positioning of the firm as an integration and infrastructure play rather than a strategy leader.

Per model summary

  • gpt-5.6-luna
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Accenture is presented as the strongest overall choice, combining strategy, technology implementation, and managed services at scale, particularly for moving beyond pilots to enterprise-wide deployment.

  • gpt-5.6-sol
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Accenture is framed as the strongest default choice given its broad combination of strategy, integration, cloud, and delivery scale, while noting clients must control scope, cost, and complexity.

  • gpt-5.6-terra
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Accenture is consistently the strongest all-around choice, pairing board-level strategy with large-scale delivery to move from pilots to enterprise-wide operating-model and workforce change.

  • kimi-k3
    brand score
    96
    first choice
    90
    median rank
    1.0
    present in
    5/5

    Accenture is highlighted for unmatched implementation and delivery scale, multi-billion-dollar AI investment, and key platform alliances, though its brand skews toward execution rather than top-table strategy.

  • claude-opus-5
    brand score
    88
    first choice
    70
    median rank
    2.0
    present in
    5/5

    Accenture is characterized as having the largest scaled AI delivery capability — multi-billion-dollar gen-AI bookings, tens of thousands of AI-trained practitioners, and deep hyperscaler partnerships — making it the strongest choice for end-to-end execution rather than strategy alone.

  • gemini-3.7-flash
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    5/5

    Accenture is described as offering unmatched global scale and systems integration for massive IT and infrastructure overhauls, backed by multi-billion-dollar AI investment, though it leans more toward implementation than pure top-down strategy.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gpt-5.6-luna6 sites · 14 references
Accenture Technology Vision×7 · 50%

accenture.com

gpt-5.6-sol8 sites · 15 references
Accenture Annual Report×8 · 53%

accenture.com

gpt-5.6-terra5 sites · 12 references
Accenture — AI services×7 · 58%

accenture.com

kimi-k38 sites · 15 references
Accenture×4 · 27%

accenture.com

claude-opus-57 sites · 15 references
Accenture + NVIDIA / Microsoft alliance announcements×5 · 33%

newsroom.accenture.com

gemini-3.7-flash7 sites · 15 references
Gartner×4 · 27%

gartner.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Consulting companies for AI transformation
#2

McKinsey & Company

also named McKinsey

79
brand score
60first choice
12345

Summary

McKinsey & Company holds a strong overall position, with a median rank of 2 (Q1 1, Q3 2) across 30 ranked answers from 6 models, and appears under both "McKinsey & Company" and "McKinsey." The models converge on a consistent underlying rationale: McKinsey is valued for C-suite strategy credibility, use-case prioritization, and operating-model redesign, with its QuantumBlack unit repeatedly cited as adding technical and data-science depth. This theme is supported by heavy recall of McKinsey's own research and capability pages — particularly the State of AI report and QuantumBlack insights — alongside third-party analyst and business-press sources such as Gartner, Forrester, Forbes, Harvard Business Review, IDC MarketScape, and coverage of McKinsey's internal Lilli tool (Reuters, CNBC).

Agreement on the reasoning is high, but placement varies more than the median suggests. Gemini-3.7-flash and claude-opus-5 rank it strongest (both median 1, with gemini at Q1 1/Q3 1 and no noted drawbacks), while gpt-5.6-luna, gpt-5.6-terra, and kimi-k3 land at median 2, and gpt-5.6-sol places it lowest at median 3 (Q1 3, Q3 4). The recurring caveat behind the lower placements is consistent across those models: McKinsey is seen as ranking below Accenture or larger systems integrators for hands-on, at-scale implementation and ongoing operations, with premium pricing and a thinner delivery bench sometimes requiring additional implementation partners.

Per model summary

  • claude-opus-5
    brand score
    92
    first choice
    80
    median rank
    1.0
    present in
    5/5

    McKinsey's QuantumBlack unit is framed as combining board-level strategy credibility with a large bench of data scientists, backed by influential research (State of AI), making it best for enterprise-wide value-case definition and operating-model redesign — though at premium cost and often needing implementation partners at scale.

  • gemini-3.7-flash
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    McKinsey is consistently presented as pairing world-class C-suite strategy with deep technical execution via QuantumBlack, positioning it as the premier end-to-end choice for large-scale Fortune 500 transformation and change management, with no noted drawbacks.

  • gpt-5.6-luna
    brand score
    76
    first choice
    47
    median rank
    2.0
    present in
    5/5

    McKinsey is described as strong at AI strategy, use-case prioritization, and operating-model change, but consistently ranked below Accenture because large-scale technical implementation and ongoing operations may require systems integrators or other partners.

  • gpt-5.6-terra
    brand score
    72
    first choice
    43
    median rank
    2.0
    present in
    5/5

    McKinsey is framed as excellent for defining the AI agenda, prioritizing use cases, and operating-model/executive alignment, with QuantumBlack adding capability, while consistently cautioning clients to validate hands-on engineering and delivery capacity for large implementations.

  • kimi-k3
    brand score
    84
    first choice
    60
    median rank
    2.0
    present in
    5/5

    McKinsey is characterized by unmatched C-suite credibility, strategy rigor, and QuantumBlack's technical depth plus influential State of AI research, but ranked second to Accenture due to premium pricing and a thinner hands-on, at-scale implementation bench.

  • gpt-5.6-sol
    brand score
    52
    first choice
    32
    median rank
    3.0
    present in
    5/5

    McKinsey is repeatedly cited for executive alignment, operating-model redesign, and tying AI to measurable business value with QuantumBlack support, but ranked below larger systems integrators for implementation-heavy, at-scale programs that may need additional partners.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-58 sites · 15 references
McKinsey State of AI report×8 · 53%

mckinsey.com

gemini-3.7-flash5 sites · 15 references
Forbes×5 · 33%

forbes.com

gpt-5.6-luna3 sites · 14 references
gpt-5.6-terra2 sites · 12 references
McKinsey — The State of AI×11 · 92%

mckinsey.com

kimi-k36 sites · 15 references
McKinsey & Company×8 · 53%

mckinsey.com

gpt-5.6-sol6 sites · 15 references

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Consulting companies for AI transformation
#3

Boston Consulting Group

also named BCG

53
brand score
32first choice
12345

Summary

Boston Consulting Group lands mid-pack across the six models, with a median rank of 3 (Q1 3, Q3 4) over 30 ranked answers. There is moderate agreement on its placement but a visible spread in enthusiasm: gemini-3.7-flash ranks it highest at a median of 2 (Q1 2, Q3 2), kimi-k3 places it at 3 (Q1 3, Q3 3), and the remaining four models — claude-opus-5, gpt-5.6-luna, gpt-5.6-sol, and gpt-5.6-terra — each land at a median of 4, with gpt-5.6-terra the most consistent at that level (Q1 4, Q3 4). The firm is named as both Boston Consulting Group and BCG.

The recurring theme is a strategy-plus-build profile: models consistently credit BCG for elite corporate strategy, operating-model redesign, and value/ROI framing, paired with technical delivery through BCG X, and they cite the firm's research output, including the 10-20-70 framing and OpenAI/Anthropic partnerships. The countervailing theme, and the main reason it ranks behind McKinsey, Accenture, and Deloitte in several answers, is a perceived thinner global implementation, systems-integration, and managed-services footprint for full Fortune 500-scale rollouts. Sources cluster around BCG's own materials — BCG X, BCG's AI capabilities pages, the AI Radar, and the BCG Henderson Institute — supplemented by MIT Sloan Management Review (including joint MIT Sloan–BCG research), Harvard Business Review, and the Harvard Business School–BCG field experiment on GenAI productivity. Analyst and press validation appears through Gartner, IDC MarketScape, the Forrester Wave (AI Consultancies / AI Services), Forbes, the World Economic Forum, and Financial Times and Reuters coverage of BCG's AI revenue share.

Per model summary

  • gemini-3.7-flash
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    5/5

    BCG pairs elite corporate strategy with technical build and engineering through BCG X, emphasizing proprietary domain-specific AI solutions, business value/ROI, and responsible AI governance for large enterprises.

  • kimi-k3
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    5/5

    BCG blends strategy with BCG X build capability and highly cited AI research (e.g., 10-20-70 framing), but ranks behind McKinsey and Accenture because its large-scale global delivery bench is thinner and less proven at full enterprise-wide deployment.

  • claude-opus-5
    brand score
    44
    first choice
    27
    median rank
    4.0
    present in
    5/5

    BCG X provides genuine build capability alongside top-tier strategy, with roughly a fifth of revenue from AI work and notable OpenAI/Anthropic partnerships, but its delivery footprint is smaller than Accenture's or Deloitte's for large-scale integration and long-term run services.

  • gpt-5.6-luna
    brand score
    48
    first choice
    28
    median rank
    4.0
    present in
    5/5

    BCG excels at business strategy, innovation, and business-model redesign around AI, but ranks below implementation-led firms because a Fortune 500 transformation may need more systems integration, engineering capacity, and managed services.

  • gpt-5.6-sol
    brand score
    48
    first choice
    28
    median rank
    4.0
    present in
    5/5

    BCG combines senior-level strategy and operating-model redesign with product and technical build capabilities, strong for high-value differentiated AI use cases, though its global implementation and managed-services scale trails Accenture's and Deloitte's.

  • gpt-5.6-terra
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    5/5

    BCG is compelling for value-led AI strategy, portfolio prioritization, and business-model innovation supported by BCG X, but ranks below the top firms because Fortune 500-scale rollouts often need larger global implementation and managed-delivery depth.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gemini-3.7-flash10 sites · 15 references
Gartner×3 · 20%

gartner.com

kimi-k37 sites · 15 references
Boston Consulting Group×7 · 47%

bcg.com

claude-opus-58 sites · 15 references
BCG X×8 · 53%

bcg.com

gpt-5.6-luna4 sites · 14 references
BCG Artificial Intelligence×10 · 71%

bcg.com

gpt-5.6-sol7 sites · 15 references
BCG Artificial Intelligence×8 · 53%

bcg.com

gpt-5.6-terra1 site · 12 references
BCG — Artificial Intelligence×12 · 100%

bcg.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Consulting companies for AI transformation
#4

Deloitte

51
brand score
32first choice
12345

Summary

Deloitte lands in the middle of the pack overall, with a median rank of 3.5 (Q1 3, Q3 4) across 30 ranked answers from 6 models. Model-level placement spans a moderate range: gpt-5.6-sol ranks it highest at a median of 2 (Q1 2, Q3 2), while claude-opus-5, gpt-5.6-luna, and gpt-5.6-terra cluster at a median of 3, and kimi-k3 and gemini-3.7-flash place it lower at medians of 4 and 5 respectively. The models converge on a shared characterization of breadth even as they diverge on how favorably to weigh it.

The dominant theme is Deloitte's end-to-end coverage tying AI to risk, regulatory, cyber, finance, tax, and workforce processes, making it a strong fit for regulated Fortune 500 enterprises—a framing echoed by gpt-5.6-sol, claude-opus-5, gpt-5.6-luna, and gpt-5.6-terra, several of which add the recurring caveat that delivery quality and AI depth vary by practice, member firm, and geography. The lower-ranking models reframe that same breadth as a limitation: kimi-k3 sees the AI work as more generalist and integration-led and less distinctive than MBB strategy houses or Accenture, while gemini-3.7-flash positions Deloitte as an execution and compliance partner rather than a frontier strategy or deployment leader. Sources underpinning these themes lean heavily on Deloitte's own materials—the State of Generative AI in the Enterprise reports, the AI Institute, the Trustworthy AI framework, Tech Trends, and various AI and Data service pages—supplemented by third-party analyst references including Gartner (Magic Quadrant for Data and Analytics Service Providers and the Market Guide for AI Consulting and System Integration Services), the IDC MarketScape, Forrester Wave, and Everest Group PEAK Matrix, with gemini-3.7-flash notably drawing more on Gartner and IDC MarketScape plus outlets such as Bloomberg, the Financial Times, and Fortune.

Per model summary

  • gpt-5.6-sol
    brand score
    72
    first choice
    45
    median rank
    2.0
    present in
    5/5

    Deloitte is emphasized for integrating AI with enterprise processes—risk, cyber, finance, tax, regulatory and workforce transformation—making it strong for complex regulated companies, though delivery quality and strategic distinctiveness vary by team and region.

  • claude-opus-5
    brand score
    52
    first choice
    30
    median rank
    3.0
    present in
    5/5

    Deloitte is consistently framed as a broad, end-to-end choice combining strategy, technology, risk and regulatory/AI-governance capability (Deloitte AI Institute, Trustworthy AI) well-suited to regulated Fortune 500 industries, with the recurring caveat that delivery quality varies by practice and geography.

  • gpt-5.6-luna
    brand score
    56
    first choice
    33
    median rank
    3.0
    present in
    5/5

    Deloitte is portrayed as offering broad multidisciplinary coverage (strategy, risk, cybersecurity, technology, compliance) valuable for highly regulated Fortune 500 firms, with the consistent caveat that experience and AI depth vary by member firm, geography and delivery team.

  • gpt-5.6-terra
    brand score
    68
    first choice
    40
    median rank
    3.0
    present in
    5/5

    Deloitte is consistently positioned as a strong fit where AI must be tightly integrated with risk, regulatory, cyber, finance, tax and workforce processes across complex or regulated enterprises, with a recurring recommendation to vet the specific proposed team since experience varies by practice and geography.

  • kimi-k3
    brand score
    32
    first choice
    23
    median rank
    4.0
    present in
    5/5

    Deloitte is credited with enormous breadth, its AI Institute, generative AI research and strength in regulated sectors tying AI to risk/tax/audit, but repeatedly ranked lower because its AI work is perceived as less distinctive and more generalist/integration-led than the MBB strategy houses or Accenture.

  • gemini-3.7-flash
    brand score
    28
    first choice
    22
    median rank
    5.0
    present in
    5/5

    Deloitte is described as offering comprehensive breadth in AI governance, regulatory compliance and risk via its Trustworthy AI framework and AI Institute/Academy, ideal for heavily regulated enterprises, but ranked fifth because it is seen as an execution/compliance partner rather than a frontier strategy or deployment leader.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gpt-5.6-sol8 sites · 15 references
Deloitte AI Institute×3 · 20%

www2.deloitte.com

claude-opus-57 sites · 15 references
Deloitte AI Institute / State of Generative AI in the Enterprise×7 · 47%

www2.deloitte.com

gpt-5.6-luna3 sites · 14 references
gpt-5.6-terra3 sites · 12 references
kimi-k38 sites · 15 references
Deloitte×4 · 27%

deloitte.com

gemini-3.7-flash8 sites · 15 references
Gartner×5 · 33%

gartner.com

  • ×4/
  • ×1named without a URL

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Consulting companies for AI transformation
#5

IBM Consulting

also named IBM

15
brand score
13first choice
12345

Summary

IBM Consulting occupies a consistent mid-tier position across the models, with a median rank of 5 (Q1 5, Q3 5) over 19 ranked answers from 6 models. The agreement here is unusually tight: every one of the six models—claude-opus-5, gemini-3.7-flash, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, and kimi-k3—landed on a median rank of 5, with no dispersion between the first and third quartiles at either the aggregate or individual level. The firm is referenced under two names, IBM Consulting and IBM.

The thematic consensus is equally strong. Models uniformly credit IBM with technical and engineering depth, hybrid-cloud capability, watsonx assets, data governance, and suitability for regulated and legacy-heavy environments—claude-opus-5 notes competitive pricing, while gpt-5.6-sol and gpt-5.6-terra emphasize fit where IBM or Red Hat is already strategic. The recurring reason for the lower placement is perceived vendor bias toward IBM's own stack and weaker vendor-neutrality, cited by claude-opus-5, gemini-3.7-flash, gpt-5.6-luna, gpt-5.6-sol, and gpt-5.6-terra; claude-opus-5 additionally flags weaker C-suite/board-level strategy influence, and kimi-k3 points to less breadth and large-scale transformation track record than bigger firms. These themes draw on a mix of IBM's own materials—IBM watsonx, IBM Consulting AI Services, the IBM Institute for Business Value, and IBM Annual Report and earnings commentary—alongside third-party analyst sources including Gartner, Forrester, IDC (including the IDC MarketScape), Everest Group PEAK Matrix assessments, and HFS Research.

Per model summary

  • claude-opus-5
    brand score
    24
    first choice
    21
    median rank
    5.0
    present in
    5/5

    Consistently frames IBM Consulting as strong on engineering depth, hybrid-cloud, and watsonx assets with competitive pricing, but ranks it lower due to perceived bias toward IBM's own stack and weaker C-suite/board-level strategy influence.

  • gemini-3.7-flash
    brand score
    4
    first choice
    4
    median rank
    5.0
    present in
    1/5

    Emphasizes deep hybrid-cloud and technical engineering strength for legacy modernization in regulated environments, while flagging potential vendor bias toward IBM's own platform ecosystem.

  • gpt-5.6-luna
    brand score
    12
    first choice
    12
    median rank
    5.0
    present in
    3/5

    Presents IBM as credible for hybrid cloud, data governance, and regulated enterprise environments, but ranks it lower because its fit is strongest for clients already aligned with IBM's ecosystem rather than those seeking vendor-neutral strategy.

  • gpt-5.6-sol
    brand score
    28
    first choice
    23
    median rank
    5.0
    present in
    5/5

    Positions IBM as strong for technically complex, hybrid-cloud, regulated, and legacy-heavy transformations, ranking it lower because its advantages depend on architecture aligning with IBM/Red Hat and it is seen as less vendor-neutral.

  • gpt-5.6-terra
    brand score
    12
    first choice
    12
    median rank
    5.0
    present in
    3/5

    Describes IBM as credible for hybrid-cloud, data-platform, and legacy modernization needs, especially where IBM is already strategic, but ranks it fifth for being more platform-centric and less vendor-neutral.

  • kimi-k3
    brand score
    8
    first choice
    8
    median rank
    5.0
    present in
    2/5

    Recognizes IBM's watsonx heritage and legacy-modernization strength, but ranks it fifth because its brand is tied to IBM's own stack and it lacks the breadth and large-scale transformation track record of bigger firms.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-57 sites · 15 references
IBM earnings commentary on generative AI book of business×9 · 60%

ibm.com

gemini-3.7-flash3 sites · 3 references
Forrester×1 · 33%

forrester.com

gpt-5.6-luna2 sites · 8 references
IBM Consulting AI Services×7 · 88%

ibm.com

gpt-5.6-sol7 sites · 15 references
IBM Consulting Artificial Intelligence×9 · 60%

ibm.com

gpt-5.6-terra2 sites · 7 references
IBM Consulting — Artificial intelligence×6 · 86%

ibm.com

kimi-k32 sites · 6 references
IBM Consulting×5 · 83%

ibm.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Consulting companies for AI transformation

How this was measured

The question

  • Unaided brand recommendation question: “Which consulting firm should a Fortune 500 company hire for an AI transformation?
  • The prompt asks each model to return exactly 5 brands, ranked 1 to 5, and for each one a reason for the recommendation and the sources that informed it.

Sampling

  • Models are not deterministic, so the question is asked over and over — 5 answers per model
  • Spellings of the same brand are normalized to the most commonly used form before counting

Measures of Position: Median and Quartiles

  • Median value indicates that in 50% of answers the brand held this position or higher.
  • Quartiles help visualize the spread of rankings per brand. Q1 indicates that 25% of answers had this rank or higher. Q3 means that 75% of answers ranked the brand as X or better.
  • Medians and quartiles are computed only from answers where the brand was present.

Coverage

  • Count of answers in which the given brand was named, as a percentage.

Brand score - Normalized Borda score

  • The brand score is calculated using a normalized Borda score. Borda count is a voting method: each ballot awards points by position instead of naming one winner. Our ranking responses from the LLM are always fixed to 5 answers, so we assign 100 points for rank 1, then 80, 60, 40, 20 — and 0 if the brand is missing.
Brand score=100ni=1n6ri5\text{Brand\ score} = \frac{100}{n} \sum_{i=1}^{n} \frac{6 - r_i}{5}

where

ri={1,,5if the brand appears6if the brand is not mentionedr_i = \begin{cases} 1,\dots,5 & \text{if the brand appears} \\ 6 & \text{if the brand is not mentioned} \end{cases}
  • This way we can calculate a common score for all brands mentioned across model responses. The score reflects both how high the brand was ranked and how frequently it was mentioned.

First choice score - adjusted MRR - top-rank indicator

  • The first choice score is based on Mean Reciprocal Rank (a common search-engine measure), but adapted to measure the position of a specific brand across repeated ranked responses.
  • Each occurrence receives a reciprocal-position score: 100 for 1st, 50 for 2nd, 33.3 for 3rd, 25 for 4th, 20 for 5th, and 0 when the brand is not mentioned. The scores are averaged across all responses.
First choice score=100ni=1nsi\text{First\ choice\ score} = \frac{100}{n} \sum_{i=1}^{n} s_i

where

si={1/riif the brand is present0if the brand is absents_i = \begin{cases} 1/r_i & \text{if the brand is present} \\ 0 & \text{if the brand is absent} \end{cases}
  • No rank below 1st place gets more than 50, so the score is heavily driven by first places.
Category · MBA​ Programs ROI

Which MBA program gives you the best return on investment?

6 AI models · 4 answers per model · ranked by score

ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0

  1. #1Stanford Graduate School of Business
    58
    brand score
    57first choice
    12345
    • gpt-5.6-sol 4.0
  2. #2INSEAD
    42
    brand score
    31first choice
  3. #3Harvard Business School
    34
    brand score
    21first choice
  4. #4BYU Marriott School of Business
    32
    brand score
    29first choice
  5. #5The Wharton School
    29
    brand score
    18first choice
    • gpt-5.6-sol 3.0
Show 19 more
  1. #6Texas McCombs
    24
    brand score
    15first choice
  2. #7Chicago Booth
    23
    brand score
    17first choice
    • gpt-5.6-terra 2.0
  3. #8Georgia Tech Scheller
    9
    brand score
    7first choice
  4. #9Darden
    9
    brand score
    5first choice
  5. #10Indiana Kelley
    8
    brand score
    5first choice
  6. #11Gies College of Business
    4
    brand score
    4first choice
  7. #12Massachusetts Institute of Technology
    4
    brand score
    3first choice
  8. #13Rice
    3
    brand score
    2first choice
  9. #14Ross
    3
    brand score
    2first choice
  10. #15Questrom School of Business
    3
    brand score
    2first choice
  11. #16Tuck
    3
    brand score
    1first choice
  12. #17University of Florida
    3
    brand score
    1first choice
  13. #18Kellogg School of Management
    2
    brand score
    #16 2first choice
  14. #19UNC Kenan-Flagler
    2
    brand score
    1first choice
  15. #20University of Georgia
    2
    brand score
    1first choice
  16. #21Carnegie Mellon Tepper
    1
    brand score
    1first choice
  17. #22Indian Institute of Management Ahmedabad
    1
    brand score
    1first choice
  18. #23UCLA Anderson
    1
    brand score
    1first choice
  19. #24University of Washington
    1
    brand score
    1first choice

Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.

Swipe left for the full report

Category · MBA​ Programs ROI

Model by model

Median rank each model gave each brand.

Scroll the table sideways to see every model →

claude-opus-5gemini-3.7-flashgpt-5.6-lunagpt-5.6-solgpt-5.6-terrakimi-k3
Stanford Graduate School of Business1.01.04.04.01.0
INSEAD3.52.01.01.01.02.0
Harvard Business School2.03.04.05.03.5
BYU Marriott School of Business5.05.01.02.01.05.0
The Wharton School3.04.05.03.03.0
Texas McCombs2.02.05.0
Chicago Booth4.55.02.52.02.0
Georgia Tech Scheller3.01.03.0
Darden3.03.5
Indiana Kelley4.03.05.0
Gies College of Business1.0
Massachusetts Institute of Technology4.05.0
Rice4.0
Ross4.04.0
Questrom School of Business2.0
Tuck3.0
University of Florida3.0
Kellogg School of Management5.05.0
UNC Kenan-Flagler4.0
University of Georgia4.0
Carnegie Mellon Tepper5.0
Indian Institute of Management Ahmedabad5.0
UCLA Anderson5.0
University of Washington5.0
1st
2
3
4
5th

A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.

Category · MBA​ Programs ROI

Where the models disagree

6 of 24 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.

  1. #1Stanford Graduate School of Business

    gpt-5.6-terra 0 kimi-k3 100 · 1 of 6 never named it

  2. #6Texas McCombs

    claude-opus-5 0 gpt-5.6-luna 80 · 3 of 6 never named it

  3. #3Harvard Business School

    gpt-5.6-luna 0 claude-opus-5 75 · 1 of 6 never named it

  4. #4BYU Marriott School of Business

    gemini-3.7-flash 5 gpt-5.6-luna 75

  5. #2INSEAD

    gpt-5.6-luna 25 gemini-3.7-flash 80

  6. #5The Wharton School

    gpt-5.6-terra 0 kimi-k3 55 · 1 of 6 never named it

A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.

Category · MBA​ Programs ROI

Named, and named first

Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.

the default answera narrow favouritelisted, rarely ledthe long tail12345678
  1. 1Stanford Graduate School of Business67/67
  2. 2INSEAD54/50
  3. 3Harvard Business School58/0
  4. 4BYU Marriott School of Business54/33
  5. 5The Wharton School58/0
  6. 6Texas McCombs33/0
  7. 7Chicago Booth46/7
  8. 8Georgia Tech Scheller13/33

The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.

Category · MBA​ Programs ROI

What the models actually said

Models
6
Answers
24

Stanford Graduate School of Business leads with a score of 58 of 100, ranked by 5 of 6 models.

The category splits cleanly along one fault line: whether "ROI" means prestige-and-earnings or cost-and-payback. The two readings produce almost opposite rankings, and the split runs largely between model families rather than across every brand evenly.

The core disagreement: prestige vs. payback

Stanford GSB is the clearest example. It carries a median rank of 1 across all models, but that headline hides a divide. claude-opus-5, gemini-3.7-flash, and kimi-k3 each rank it first, anchoring on the highest post-MBA compensation and Silicon Valley access. Both GPT models that ranked it — gpt-5.6-luna and gpt-5.6-sol — place it at median 4, weighting attendance cost and forgone earnings more heavily.

However, its high cost and the earnings forgone during the degree make its expected financial ROI less consistently compelling than lower-cost programs, especially for students without scholarship support. — gpt-5.6-luna

The likely driver is source mix rather than a genuine dispute over outcomes. The models ranking Stanford top lean on employment-report and ranking sources; the GPT models pair those with cost-side pages such as Stanford GSB Cost of Attendance. That said, the source lists alone do not prove causation — this is a plausible reading, not something the data confirms.

The starkest split: BYU Marriott

BYU Marriott shows the widest divergence in the category, median rank 4 but Q1 1 and Q3 5. The gpt-5.6 family clusters at the top (luna and terra at 1, sol at 2), treating low tuition and solid six-figure placement as decisive. claude-opus-5, gemini-3.7-flash, and kimi-k3 all put it at 5, accepting the same value logic but treating the capped earnings ceiling as disqualifying for a top slot.

BYU Marriott is one of the strongest value choices because its tuition is unusually low for a nationally recognized full-time MBA, while graduates still access consulting, technology, and finance recruiting. — gpt-5.6-luna

This is the cleanest illustration of the two ROI definitions producing opposite ranks from shared facts.

Chicago Booth: the same pattern, reversed

Booth mirrors the split but with the model families swapped. gpt-5.6-terra (2), kimi-k3 (2), and gpt-5.6-sol (2.5) rank it high; claude-opus-5 (4.5) and gemini-3.7-flash (5) rank it low. Here kimi-k3 sits on the payback side, citing a Forbes ROI-based ranking, while claude-opus-5 emphasizes total cost.

Forbes' ROI-based ranking put Booth at the top for five-year MBA gain, reflecting exceptional salary growth relative to cost. — kimi-k3

The Forbes ROI ranking is a plausible source behind kimi-k3's high placement, though the same model puts BYU at 5 — so its logic is not uniformly cost-first, cautioning against reading any model as applying one fixed definition everywhere.

Where the models agree

INSEAD is the strongest point of consensus, median 2, with all three GPT models at 1 and gemini and kimi at 2. Every model credits the one-year format for halving opportunity cost. Only claude-opus-5 (3.5) qualifies it, on the compressed timeline limiting internships — a difference of emphasis, not of the core cost case.

Wharton also draws fairly tight agreement (median 3.5), with disagreement confined to cost rather than earning power. Kellogg is unanimous where ranked, both gpt-5.6-sol and kimi-k3 at 5, both citing weaker finance exposure.

Harvard: consensus on quality, spread on arithmetic

HBS lands at median 3 with a real spread: claude-opus-5 at 2, gpt-5.6-terra at 5. No model disputes the compensation or alumni network; the divergence is entirely about whether total cost and two-year opportunity cost offset that.

I rank it fourth because its large total cost and two-year opportunity cost can lengthen payback, particularly for candidates who already have lucrative careers or receive limited aid. — gpt-5.6-sol

The single-model tail

A large share of brands — Gies, Questrom, Tuck, University of Florida, Rice, UNC Kenan-Flagler, University of Georgia, Tepper, IIM Ahmedabad, UCLA Anderson, University of Washington — appear in only one model's ranking, so no agreement or disagreement can be measured. Notably, most of the affordability and online-focused picks (Gies, Questrom, Georgia Tech Scheller at rank 1 for sol) come from the GPT family, consistent with its payback-first framing. Whether that reflects the models' reasoning or just which brands each happened to recall is not something the data settles.

Sources behind the splits

The most-recalled sources — Poets&Quants (25×), U.S. News, and the Financial Times Global MBA Ranking — appear across both camps, so they do not explain the divergence on their own. The distinguishing feature is that the payback-leaning rankings more often surface program-specific cost-of-attendance and tuition pages alongside employment reports, and kimi-k3 and gpt-5.6-terra reach for Forbes' ROI-based ranking. The link between those cost-side sources and the lower prestige-school placements is the most consistent pattern in the data, but it remains a hypothesis about what drove each ranking, not a demonstrated cause.

Category · MBA​ Programs ROI

What shaped the answers

74 sources across 338 references, grouped by site from 177 recalled names. The top 5 carry 46% of them.

U.S. News Best Business Schools×50 · 15%

usnews.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · MBA​ Programs ROI
#1

Stanford Graduate School of Business

also named Stanford University

58
brand score
57first choice
  • gpt-5.6-sol 4.0
12345

Summary

Stanford Graduate School of Business lands at a median rank of 1 across all models (Q1 1, Q3 1) over 16 ranked answers from 5 models, though that headline figure masks a clear split. Three models—claude-opus-5, gemini-3.7-flash, and kimi-k3—each place it at a median rank of 1, converging on a single theme: the highest post-MBA compensation of any program, paired with Silicon Valley access to venture capital, technology, private equity, and founder equity, which they argue offsets premium tuition to produce the strongest long-run payback. These models draw on employment-report data (Stanford GSB Employment Report) alongside ranking sources such as the Financial Times Global MBA Ranking, Poets&Quants, Forbes, U.S. News, and Bloomberg Businessweek. Kimi-k3 in particular anchors its case in specific compensation figures.

The disagreement comes from the two GPT models. gpt-5.6-luna places it at a median rank of 4 (Q1 4, Q3 4) over 1 ranked answer, and gpt-5.6-sol at a median rank of 4 (Q1 2.5, Q3 4.5) over 3 ranked answers. Both acknowledge the same upside but weight cost and opportunity cost more heavily, arguing that high attendance costs and uncertain near-term entrepreneurial outcomes make the measurable ROI less predictable than lower-cost programs ranked above it. Their reasoning leans on cost-side sources such as Stanford GSB Cost of Attendance and Stanford GSB Financial Aid alongside the employment reports and Financial Times ranking.

Stanford consistently posts the highest median total compensation of any MBA program, with recent classes exceeding $250,000 thanks to outsized placement in tech, venture capital, and private equity. — kimi-k3

However, its high cost and the earnings forgone during the degree make its expected financial ROI less consistently compelling than lower-cost programs, especially for students without scholarship support. — gpt-5.6-luna

Per model summary

  • claude-opus-5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    4/4

    Consistently emphasizes Stanford's highest median post-MBA compensation and Silicon Valley/VC/tech equity upside, arguing the salary premium and network offset its very high tuition to produce top-ranked long-run payback.

  • gemini-3.7-flash
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    4/4

    Repeatedly points to the world's highest post-graduation compensation and deep Silicon Valley ties (VC, tech, PE, founder equity), which drive long-term wealth creation and quickly offset premium tuition.

  • kimi-k3
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    4/4

    Consistently cites the highest post-MBA median total compensation (often above $230K–$250K) and top placement in tech, VC, and PE, which offset high tuition to yield a short payback and the strongest lifetime earnings uplift.

  • gpt-5.6-luna
    brand score
    10
    first choice
    6
    median rank
    4.0
    present in
    1/4

    Acknowledges strong compensation potential and entrepreneurial access but stresses that high cost and forgone earnings make its financial ROI less consistently compelling than lower-cost programs.

  • gpt-5.6-sol
    brand score
    40
    first choice
    36
    median rank
    4.0
    present in
    3/4

    Highlights extraordinary long-term upside via tech, VC, and entrepreneurship, while noting that high costs and uncertain near-term entrepreneurial outcomes make its measurable ROI less predictable.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-55 sites · 12 references
Stanford GSB Employment Report×4 · 33%

gsb.stanford.edu

gemini-3.7-flash4 sites · 12 references
Financial Times×4 · 33%

rankings.ft.com

kimi-k36 sites · 12 references
Poets&Quants×4 · 33%

poetsandquants.com

gpt-5.6-luna3 sites · 3 references
Financial Times Global MBA Ranking×1 · 33%

no URL recalled

  • ×1named without a URL
gpt-5.6-sol3 sites · 9 references

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · MBA​ Programs ROI
#2

INSEAD

42
brand score
31first choice
12345

Summary

INSEAD lands at a median rank of 2 (Q1 2, Q3 2) across 13 ranked answers from 6 models, placing it consistently near the top of the field. Agreement is strong at the upper end: gpt-5.6-luna, gpt-5.6-sol, and gpt-5.6-terra all rank it first (each median 1), and gemini-3.7-flash and kimi-k3 place it second (each median 2). The main dispersion comes from claude-opus-5, whose median of 3.5 (Q1 2.75, Q3 4) reflects a more qualified view. The dominant theme across every model is the accelerated 10-month, one-year format, which cuts forgone salary and living costs roughly in half versus two-year U.S. programs and produces one of the fastest payback periods among elite schools. This is reinforced by consistent references to strong global consulting and corporate placement, particularly MBB, across Europe, Asia, and the Middle East.

The recalled sources cluster around the same ranking and placement authorities: the Financial Times Global MBA Ranking appears under nearly every model, alongside INSEAD's own Employment Statistics and Employment Report, with Poets&Quants, Forbes, Bloomberg Businessweek, and The Economist supporting the ROI and placement themes. Claude-opus-5's lower placement stems from the trade-offs it foregrounds — the compressed timeline limiting internships and access to U.S. finance recruiting — rather than any disagreement on the core cost advantage.

INSEAD's one-year format is the core of its ROI case: you pay one year of tuition and forfeit only one year of salary, cutting total cost dramatically versus two-year US programs. — kimi-k3

The 10-month format cuts opportunity cost roughly in half compared with two-year US programs while still delivering elite consulting placement (MBB hires a large share of each class). — claude-opus-5

Per model summary

  • gpt-5.6-luna
    brand score
    25
    first choice
    25
    median rank
    1.0
    present in
    1/4

    The roughly 10-month format reduces lost salary and living costs while its international network and mobility benefit candidates in consulting, finance, or multinational careers.

  • gpt-5.6-sol
    brand score
    25
    first choice
    25
    median rank
    1.0
    present in
    1/4

    The accelerated 10-month format minimizes tuition and lost earnings while providing global employer access, producing strong near-term ROI for internationally oriented candidates.

  • gpt-5.6-terra
    brand score
    25
    first choice
    25
    median rank
    1.0
    present in
    1/4

    The one-year format reduces forgone salary and living costs while its global recruiting network supports post-MBA mobility, making it a strong ROI choice for international careers.

  • gemini-3.7-flash
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    4/4

    The intensive 10-month format halves opportunity cost and forgone earnings versus two-year programs, producing one of the fastest payback periods, supported by strong global consulting and corporate placement.

  • kimi-k3
    brand score
    40
    first choice
    25
    median rank
    2.0
    present in
    2/4

    The one-year format is the core ROI advantage, halving total cost by sacrificing only one year of tuition and salary, with strong MBB placement across Europe, Asia, and the Middle East yielding payback periods that beat US peers.

  • claude-opus-5
    brand score
    55
    first choice
    33
    median rank
    3.5
    present in
    4/4

    The 10-12 month format cuts tuition and forgone salary roughly in half, driving fast payback and top FT/Forbes ROI rankings, aided by strong MBB consulting placement, though the compressed timeline limits internships and US finance recruiting access.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gpt-5.6-luna2 sites · 2 references
Financial Times Global MBA Ranking×1 · 50%

rankings.ft.com

gpt-5.6-sol2 sites · 3 references
INSEAD Employment Statistics×2 · 67%

insead.edu

gpt-5.6-terra3 sites · 3 references
Financial Times MBA Rankings×1 · 33%

rankings.ft.com

gemini-3.7-flash5 sites · 12 references
Financial Times×4 · 33%

rankings.ft.com

kimi-k34 sites · 6 references
INSEAD×2 · 33%

insead.edu

claude-opus-56 sites · 12 references
Financial Times Global MBA Ranking×4 · 33%

rankings.ft.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · MBA​ Programs ROI
#3

Harvard Business School

also named Harvard, Harvard University

34
brand score
21first choice
12345

Summary

Harvard Business School lands at a median rank of 3 (Q1 2.25, Q3 3.75) across 14 ranked answers from 5 models, placing it firmly in the upper tier but not at the summit. Model-level positioning spreads notably: claude-opus-5 is the most favorable at median rank 2 (Q1 2, Q3 2.25), gemini-3.7-flash and kimi-k3 sit at median 3 and 3.5 respectively, while gpt-5.6-sol places it at 4 and gpt-5.6-terra at 5. The common thread across all five is agreement on the underlying strengths — elite median compensation (cited near $175K base), the most powerful global alumni network, and generous need-based financial aid — recalled through sources such as the Harvard Business School Employment Data pages, the Financial Times Global MBA Ranking, U.S. News Best Business Schools, and Bloomberg Businessweek Best B-Schools.

The disagreement centers not on quality but on the arithmetic of ROI, where high total cost and two-year opportunity cost weigh against near-term payback. Models ranking HBS lower emphasize this cost drag and, in kimi-k3's case, the diluting effect of a large class size, while higher-ranking models offset the sticker price against fellowships and long-run career compounding. That tension explains why several models place it a notch behind Stanford despite comparable earnings.

HBS combines top-tier median compensation (~$175K base plus bonuses) with arguably the most powerful global alumni network, which compounds earnings over a career. — claude-opus-5

I rank it fourth because its large total cost and two-year opportunity cost can lengthen payback, particularly for candidates who already have lucrative careers or receive limited aid. — gpt-5.6-sol

Per model summary

  • claude-opus-5
    brand score
    75
    first choice
    46
    median rank
    2.0
    present in
    4/4

    HBS is consistently described as combining top-tier median compensation with the most powerful global alumni network and generous need-based fellowships that lower effective cost, while ranking slightly behind Stanford on pure ROI math due to high full price.

  • gemini-3.7-flash
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    4/4

    HBS is portrayed as pairing immense global brand prestige and a vast alumni network with generous need-based financial aid, producing top-tier compensation and strong lifetime compounding returns in fields like private equity and consulting.

  • kimi-k3
    brand score
    55
    first choice
    33
    median rank
    3.5
    present in
    4/4

    HBS delivers elite compensation and arguably the strongest lifetime brand equity and network, but ranks behind peers like Stanford because its very high cost and large class size make near-term salary-to-cost payback slightly less efficient.

  • gpt-5.6-sol
    brand score
    10
    first choice
    6
    median rank
    4.0
    present in
    1/4

    HBS offers extraordinary brand value and long-term career optionality, but is ranked lower because high total cost and two-year opportunity cost can lengthen payback, especially with limited aid.

  • gpt-5.6-terra
    brand score
    5
    first choice
    5
    median rank
    5.0
    present in
    1/4

    HBS provides enormous long-term upside through employer access and alumni network, but high total cost and lost earnings make near-term payback less certain, making it most compelling for those who leverage the network for leadership or venture-building.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-55 sites · 12 references
Harvard Business School Employment Data×4 · 33%

hbs.edu

gemini-3.7-flash5 sites · 12 references
Forbes×4 · 33%

forbes.com

kimi-k35 sites · 12 references
Harvard Business School×4 · 33%

hbs.edu

gpt-5.6-sol2 sites · 3 references
Harvard Business School Financial Aid×2 · 67%

hbs.edu

gpt-5.6-terra1 site · 3 references
Harvard MBA×3 · 100%

hbs.edu

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · MBA​ Programs ROI
#4

BYU Marriott School of Business

also named BYU Marriott, Brigham Young University, Brigham Young University Marriott School of Business

32
brand score
29first choice
12345

Summary

BYU Marriott School of Business lands at a median rank of 4 (Q1 1, Q3 5) across 13 ranked answers from 6 models, a spread that reflects genuine disagreement rather than consensus. The gpt-5.6 family clusters at the top—gpt-5.6-luna and gpt-5.6-terra both at median rank 1, gpt-5.6-sol at median rank 2—while claude-opus-5, gemini-3.7-flash, and kimi-k3 all settle at median rank 5. The unifying theme is that Marriott is a value or payback pick: very low tuition paired with solid six-figure placements produces a strong salary-to-cost ratio, with the reduced rate for LDS members frequently cited. The divergence in placement stems from how each model weighs that against a caveat all of them acknowledge—a narrower national and global recruiting footprint and a lower top-end earnings ceiling than elite M7 schools.

The models that rank it highest frame the low cost and access to consulting, technology, and finance recruiting as decisive, subject to fit around the faith-based environment and geographic concentration; those recalling institutional sources such as BYU Marriott MBA Tuition and the BYU Marriott MBA Employment Report sit under this theme. The models that place it fifth accept the same ROI logic but treat the capped ceiling as a reason to rank it below prestige programs, leaning on ROI-focused sources like Forbes Best Business Schools and Poets&Quants alongside U.S. News Best Business Schools.

BYU Marriott is one of the strongest value choices because its tuition is unusually low for a nationally recognized full-time MBA, while graduates still access consulting, technology, and finance recruiting. — gpt-5.6-luna

It ranks fifth here only because its ceiling for ultra-high-pa… — kimi-k3

Per model summary

  • gpt-5.6-luna
    brand score
    75
    first choice
    75
    median rank
    1.0
    present in
    3/4

    Highlights low tuition alongside solid consulting, tech, and finance recruiting, while repeatedly noting the ROI is most compelling for students comfortable with its faith-based environment and geographic/eligibility constraints.

  • gpt-5.6-terra
    brand score
    50
    first choice
    50
    median rank
    1.0
    present in
    2/4

    Stresses an unusually favorable salary-to-cost payback from low tuition and strong outcomes, while emphasizing fit—its LDS affiliation, location, and recruiting base align best with certain candidates.

  • gpt-5.6-sol
    brand score
    20
    first choice
    13
    median rank
    2.0
    present in
    1/4

    Points to strong placement and compensation at substantially lower cost, especially for students eligible for the reduced tuition, with the caveat that its culture and network may not suit everyone.

  • claude-opus-5
    brand score
    20
    first choice
    16
    median rank
    5.0
    present in
    3/4

    Consistently frames BYU Marriott as the classic value/payback pick: very low tuition (especially for LDS members) paired with strong six-figure placements, offset by a narrower national/global recruiting footprint and lower earnings ceiling than M7 schools.

  • gemini-3.7-flash
    brand score
    5
    first choice
    5
    median rank
    5.0
    present in
    1/4

    Emphasizes an exceptionally favorable salary-to-debt ratio, with very low tuition combined with competitive placements at consulting, tech, and accounting firms.

  • kimi-k3
    brand score
    20
    first choice
    16
    median rank
    5.0
    present in
    3/4

    Describes it as a top ROI standout on a salary-to-debt basis due to heavily subsidized tuition and solid six-figure placements, ranked slightly lower only because its recruiting network and top-end salary ceiling are narrower than elite schools.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gpt-5.6-luna6 sites · 7 references
BYU Marriott Graduate Program Tuition×2 · 29%

marriott.byu.edu

gpt-5.6-terra4 sites · 5 references
BYU Marriott MBA Employment Statistics×2 · 40%

marriott.byu.edu

gpt-5.6-sol3 sites · 3 references
BYU Graduate Tuition and Fees×1 · 33%

no URL recalled

  • ×1named without a URL
claude-opus-55 sites · 9 references
U.S. News Best Business Schools×3 · 33%

usnews.com

gemini-3.7-flash3 sites · 3 references
Forbes×1 · 33%

forbes.com

kimi-k35 sites · 9 references
BYU Marriott MBA×3 · 33%

marriott.byu.edu

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · MBA​ Programs ROI
#5

The Wharton School

also named The Wharton School (University of Pennsylvania), The Wharton School, University of Pennsylvania, University of Pennsylvania, Wharton, University of Pennsylvania - Wharton, University of Pennsylvania – Wharton School

29
brand score
18first choice
  • gpt-5.6-sol 3.0
12345

Summary

The Wharton School lands at a median rank of 3.5 (Q1 3, Q3 4) across 14 ranked answers from 5 models, placing it firmly in the upper tier without reaching the top spot. Agreement is fairly tight through the middle of the distribution: claude-opus-5, gpt-5.6-sol, and kimi-k3 all settle at a median of 3, while gemini-3.7-flash lands at 4 and gpt-5.6-luna sits lowest at 5. The consistent theme is Wharton's finance pipeline — private equity, investment banking, hedge funds, and consulting — paired with compensation the models describe as on par with HBS and Stanford. That case rests on career-outcome and ranking sources including the Wharton MBA Career Report and Career Statistics, Poets&Quants, the Financial Times Global MBA Ranking, U.S. News, Forbes, and Bloomberg Businessweek.

The disagreement is almost entirely about cost rather than earning power. The models that rank Wharton highest still flag that its full sticker tuition and Philadelphia living costs trim the ROI advantage, while gemini-3.7-flash frames those costs as rapidly amortized by signing bonuses and gpt-5.6-luna pushes it down to fifth on the grounds that the premium is easiest to justify only for those targeting the industries that reward the brand most.

Its large class and global alumni base give broad access to high-paying roles across industries. Costs are among the highest, which trims the ROI advantage relative to lower-tuition options. — claude-opus-5

I rank it fifth on value because its total cost is very high and its premium is easiest to justify for applicants targeting industries and employers that reward the brand particularly strongly. — gpt-5.6-luna

Per model summary

  • claude-opus-5
    brand score
    30
    first choice
    17
    median rank
    3.0
    present in
    2/4

    Wharton is framed as delivering top-tier finance placement (PE, banking, hedge funds) with compensation on par with HBS and Stanford, though its high full tuition slightly dilutes ROI relative to lower-cost or better-aid options.

  • gpt-5.6-sol
    brand score
    45
    first choice
    27
    median rank
    3.0
    present in
    3/4

    Wharton is portrayed as offering exceptional compensation and access to finance/consulting/leadership roles plus a strong alumni network, with high tuition and costs making ROI dependable mainly for those making high-paying transitions.

  • kimi-k3
    brand score
    55
    first choice
    31
    median rank
    3.0
    present in
    4/4

    Wharton is credited with elite finance/consulting compensation and strong lifetime earnings, but ranked just below Stanford and HBS because its high tuition, larger class size, and dependence on cyclical finance hiring slightly lengthen payback.

  • gemini-3.7-flash
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    4/4

    Wharton is consistently described as a dominant finance/banking/PE pipeline whose high compensation and signing bonuses rapidly amortize tuition debt, supported by its alumni network and quantitative reputation.

  • gpt-5.6-luna
    brand score
    5
    first choice
    5
    median rank
    5.0
    present in
    1/4

    Wharton is seen as an elite brand for finance, consulting, and leadership pay, but ranked lower on value because its very high cost is best justified for applicants targeting industries that reward the brand most.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-55 sites · 6 references
Poets&Quants×2 · 33%

poetsandquants.com

gpt-5.6-sol4 sites · 9 references
Wharton MBA Career Statistics×3 · 33%

statistics.mbacareers.wharton.upenn.edu

kimi-k36 sites · 12 references
Wharton MBA Career Report×4 · 33%

wharton.upenn.edu

gemini-3.7-flash5 sites · 12 references
gpt-5.6-luna3 sites · 3 references
Financial Times Global MBA Ranking×1 · 33%

no URL recalled

  • ×1named without a URL

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · MBA​ Programs ROI

How this was measured

The question

  • Unaided brand recommendation question: “Which MBA program gives you the best return on investment?
  • The prompt asks each model to return exactly 5 brands, ranked 1 to 5, and for each one a reason for the recommendation and the sources that informed it.

Sampling

  • Models are not deterministic, so the question is asked over and over — 4 answers per model
  • Spellings of the same brand are normalized to the most commonly used form before counting

Measures of Position: Median and Quartiles

  • Median value indicates that in 50% of answers the brand held this position or higher.
  • Quartiles help visualize the spread of rankings per brand. Q1 indicates that 25% of answers had this rank or higher. Q3 means that 75% of answers ranked the brand as X or better.
  • Medians and quartiles are computed only from answers where the brand was present.

Coverage

  • Count of answers in which the given brand was named, as a percentage.

Brand score - Normalized Borda score

  • The brand score is calculated using a normalized Borda score. Borda count is a voting method: each ballot awards points by position instead of naming one winner. Our ranking responses from the LLM are always fixed to 5 answers, so we assign 100 points for rank 1, then 80, 60, 40, 20 — and 0 if the brand is missing.
Brand score=100ni=1n6ri5\text{Brand\ score} = \frac{100}{n} \sum_{i=1}^{n} \frac{6 - r_i}{5}

where

ri={1,,5if the brand appears6if the brand is not mentionedr_i = \begin{cases} 1,\dots,5 & \text{if the brand appears} \\ 6 & \text{if the brand is not mentioned} \end{cases}
  • This way we can calculate a common score for all brands mentioned across model responses. The score reflects both how high the brand was ranked and how frequently it was mentioned.

First choice score - adjusted MRR - top-rank indicator

  • The first choice score is based on Mean Reciprocal Rank (a common search-engine measure), but adapted to measure the position of a specific brand across repeated ranked responses.
  • Each occurrence receives a reciprocal-position score: 100 for 1st, 50 for 2nd, 33.3 for 3rd, 25 for 4th, 20 for 5th, and 0 when the brand is not mentioned. The scores are averaged across all responses.
First choice score=100ni=1nsi\text{First\ choice\ score} = \frac{100}{n} \sum_{i=1}^{n} s_i

where

si={1/riif the brand is present0if the brand is absents_i = \begin{cases} 1/r_i & \text{if the brand is present} \\ 0 & \text{if the brand is absent} \end{cases}
  • No rank below 1st place gets more than 50, so the score is heavily driven by first places.
Category · AI models for coding

Which AI model would you recommend for coding?

7 AI models · 160 answers · ranked by score

ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0

  1. #1Anthropic
    100
    brand score
    99first choice
    12345
  2. #2OpenAI
    80
    brand score
    50first choice
  3. #3Google
    57
    brand score
    32first choice
  4. #4DeepSeek
    35
    brand score
    22first choice
  5. #5Mistral AI
    8
    brand score
    8first choice
Show 7 more
  1. #6Meta
    5
    brand score
    5first choice
  2. #7GitHub
    5
    brand score
    3first choice
  3. #8xAI
    4
    brand score
    4first choice
  4. #9Alibaba
    2
    brand score
    2first choice
  5. #10Qwen
    1
    brand score
    1first choice
  6. #11Cursor
    0
    brand score
    0first choice
  7. #12Microsoft
    0
    brand score
    0first choice

Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.

Swipe left for the full report

Category · AI models for coding

Model by model

Median rank each model gave each brand.

Scroll the table sideways to see every model →

claude-opus-5gemini-3.7-flashgpt-5.6-lunagpt-5.6-solgpt-5.6-terragpt-6-astrakimi-k3
Anthropic1.01.01.01.01.01.01.0
OpenAI2.02.02.02.02.02.02.0
Google3.04.03.03.03.03.03.0
DeepSeek4.03.04.04.04.04.04.0
Mistral AI5.05.05.05.05.05.0
Meta5.05.05.0
GitHub4.04.04.04.0
xAI5.05.0
Alibaba5.05.05.0
Qwen5.0
Cursor3.0
Microsoft4.0
1st
2
3
4
5th

A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.

Category · AI models for coding

Where the models disagree

0 of 12 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.

  1. #4DeepSeek

    gpt-5.6-luna 13 gemini-3.7-flash 49

  2. #7GitHub

    gemini-3.7-flash 0 gpt-5.6-luna 26 · 3 of 7 never named it

  3. #8xAI

    claude-opus-5 0 gpt-5.6-terra 18 · 5 of 7 never named it

  4. #6Meta

    gpt-5.6-luna 0 kimi-k3 16 · 4 of 7 never named it

  5. #5Mistral AI

    kimi-k3 0 gpt-5.6-luna 18 · 1 of 7 never named it

  6. #3Google

    gemini-3.7-flash 46 kimi-k3 60

A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.

Category · AI models for coding

Named, and named first

Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.

the default answera narrow favouritelisted, rarely ledthe long tail12345678
  1. 1Anthropic100/99
  2. 2OpenAI99/1
  3. 3Google99/0
  4. 4DeepSeek87/0
  5. 5Mistral AI37/0
  6. 6Meta27/0
  7. 7GitHub13/0
  8. 8xAI18/0

The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.

Category · AI models for coding

What the models actually said

Models
7
Answers
160

Anthropic leads with a score of 100 of 100, ranked by 7 of 7 models.

The category shows a remarkably stable spine. Anthropic, OpenAI, and Google occupy ranks 1, 2, and 3 with near-total agreement across all seven models, and the disputes only begin at rank 4 and below, where open-weight providers and product-layer tools compete for the same crowded fifth slot.

The undisputed top three

Anthropic takes first place with no dissent whatsoever: a median of 1 across 160 answers, and every individual model also returns a median of 1. OpenAI mirrors this at rank 2 (median 2 for all seven models), and Google holds rank 3 with six models agreeing. The only crack in the top three is gemini-3.7-flash, which places its own maker Google at a median of 4 rather than 3 — the sole model to rate Google below the field.

The framing splits along a benchmark-versus-workflow line, though this is a difference of emphasis rather than ranking. Benchmark-anchored models (claude-opus-5, gemini-3.7-flash, kimi-k3) lean on SWE-bench, Aider, and LMArena, while the gpt-5.6 family and gpt-6-astra frame the same placements through agentic tooling and documentation. The distinction between OpenAI and Anthropic is repeatedly cast as task-dependent rather than absolute.

They are particularly strong at algorithmic problem-solving and explaining code, though some developers find them slightly less consistent than Claude on large, multi-step engineering tasks.

kimi-k3

Where the models actually diverge

The real disagreement concerns DeepSeek. Six models place it at rank 4, but gemini-3.7-flash rates it a full rank higher at 3 — the same model that demoted Google. Its reasoning leans on benchmark parity and self-hosting appeal (SWE-bench, Hugging Face, LMSYS), whereas the six models holding DeepSeek at 4 emphasize ecosystem and enterprise-support gaps drawn from its API docs and GitHub. Whether gemini-3.7-flash's benchmark focus causes both its DeepSeek promotion and its Google demotion is a plausible reading of its source mix, not something the data confirms.

Below that, three fundamentally different kinds of brand pile up at rank 5:

  • Open-weight providers — Mistral AI, Meta, and Alibaba/Qwen — all sit at a median of 5, and the reasoning is strikingly uniform: capable and open, but ranked on out-of-the-box coding quality rather than flexibility. gemini-3.7-flash again shows marginally more spread on Mistral (Q1 4).
  • A frontier-adjacent lab — xAI, placed at 5 by only gpt-5.6-sol and gpt-5.6-terra, judged capable but thin on tooling and track record.
  • Product layers — GitHub (rank 4) and Microsoft (rank 4, one answer), rated as routers or integration layers over other labs' models rather than model providers in their own right.

The GitHub case is notable because its rank-4 placement sits above the open-weight labs at 5, yet every model attributes its ceiling to inherited capability rather than its own model.

GitHub Copilot is the most widely deployed coding assistant, with deep IDE and pull-request integration, enterprise controls and now a model picker that routes to Anthropic, OpenAI and Google models.

claude-opus-5

Single-model entries

Cursor, Microsoft, and Qwen each rest on a single model's answers, so their ranks reflect one perspective rather than any consensus. Cursor (gpt-5.6-luna, rank 3) and Qwen (gpt-6-astra, rank 5) are internally consistent within their lone raters, but carry no cross-model signal. Both are worth reading as individual judgments, not category positions.

A recurring qualifier across the low-ranked open-weight entries is that the placement would rise sharply if self-hosting were the priority — gpt-6-astra says as much for both Alibaba and Qwen, which suggests the rank-5 clustering reflects an implied "typical user" default more than a capability verdict.

I rank it fifth for a typical user seeking an immediately useful coding assistant, but it could rank considerably higher if self-hosting and control over deployment are your priorities.

gpt-6-astra

Shared evidence base

The source pool is dominated by a handful of leaderboards and benchmarks — Aider (152 times), SWE-bench (147), LMArena (79), and Artificial Analysis (48) — which explains why placements agree so tightly across models drawing on a common evidence base. Vendor documentation appears mainly for the lower-ranked and open-weight brands, consistent with those rankings being argued on deployment and ecosystem grounds rather than measured performance.

Category · AI models for coding

What shaped the answers

65 sources across 2145 references, grouped by site from 496 recalled names. The top 5 carry 45% of them.

SWE-bench×311 · 14%

swebench.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · AI models for coding
#1

Anthropic

100
brand score
99first choice
12345

Summary

Anthropic sits at the front of this category with unusual consistency: a median rank of 1 (Q1 1, Q3 1) across 160 ranked answers from all 7 models, and every one of the seven models individually returns a median rank of 1 as well. There is effectively no disagreement about placement — the divergence is only in emphasis. The benchmark-anchored models (claude-opus-5, gemini-3.7-flash, kimi-k3) lean on measured performance, citing SWE-bench Verified leadership and integration into tools like Cursor and GitHub Copilot, drawing on the SWE-bench leaderboard, Aider LLM leaderboards, and LMArena. The workflow-oriented models (gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra) frame the same ranking through repository-scale reasoning, multi-file edits, and agentic tooling such as Claude Code, sourcing Anthropic documentation and the Claude Code overview.

Claude models (Sonnet/Opus 4.x series) are widely regarded as the strongest at real-world software engineering tasks, leading benchmarks like SWE-bench Verified and powering tools such as Claude Code, Cursor and GitHub Copilot.

claude-opus-5

Underlying the top placement is a shared theme of large-codebase competence — long context, coordinated multi-file changes, careful instruction-following, and clear explanations rather than isolated snippet generation. Both camps land on the same conclusion from different angles, and one model even flags its own conflict of interest while retaining the ranking.

Claude models are especially strong at understanding large codebases, making coordinated multi-file changes, and explaining code clearly.

gpt-5.6-sol

Per model summary

  • claude-opus-5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    25/25

    Consistently emphasizes that Claude (Sonnet/Opus 4.x) leads real-world coding benchmarks like SWE-bench Verified and excels at agentic, multi-file refactoring in large codebases. Repeatedly cites Claude Code and integration into tools like Cursor and GitHub Copilot as reinforcing its top position.

  • gemini-3.7-flash
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    25/25

    Repeatedly frames Anthropic's Claude models (especially Claude 3.5 Sonnet) as the industry gold standard for code generation, refactoring, and architectural reasoning. Consistently highlights large context windows, strong instruction-following, and minimal hallucinations.

  • gpt-5.6-luna
    brand score
    98
    first choice
    95
    median rank
    1.0
    present in
    20/20

    Consistently recommends Anthropic for its strength in understanding large codebases, debugging, refactoring, and following detailed instructions to produce maintainable code. Emphasizes careful reasoning, clear explanations, and agentic/repository-level workflows.

  • gpt-5.6-sol
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    20/20

    Consistently positions Anthropic as the strongest overall coding recommendation, citing repository-scale reasoning, multi-file edits, and clear explanations. Emphasizes long context, agentic tooling like Claude Code, and reliable instruction-following.

  • gpt-5.6-terra
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    20/20

    Repeatedly recommends Anthropic for demanding, professional coding work, emphasizing repository-scale reasoning, code review, debugging, and following detailed engineering constraints. Highlights long-context capability and agentic workflows over lowest cost.

  • gpt-6-astra
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    25/25

    Consistently names Anthropic as its default choice for complex, general-purpose coding, focused on understanding existing codebases, coordinated multi-file changes, and explaining design decisions. Emphasizes Claude Code and agentic, repository-oriented workflows over snippet generation.

  • kimi-k3
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    25/25

    Consistently cites Claude models topping coding benchmarks like SWE-bench Verified and strong developer sentiment, with emphasis on multi-file reasoning and agentic tools like Claude Code. Notes precise instruction-following and integration into tools such as Cursor.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-58 sites · 75 references
SWE-bench leaderboard×25 · 33%

swebench.com

  • ×25/
gemini-3.7-flash8 sites · 69 references
SWE-bench×25 · 36%

swebench.com

  • ×25/
gpt-5.6-luna6 sites · 51 references
SWE-bench×18 · 35%

swebench.com

  • ×18/
gpt-5.6-sol7 sites · 60 references
SWE-bench×20 · 33%

swebench.com

  • ×20/
gpt-5.6-terra7 sites · 58 references
SWE-bench×20 · 34%

swebench.com

gpt-6-astra6 sites · 53 references
Anthropic — Claude Code best practices×22 · 42%

anthropic.com

kimi-k36 sites · 75 references
SWE-bench×25 · 33%

swebench.com

  • ×25/

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · AI models for coding
#2

OpenAI

80
brand score
50first choice
12345

Summary

OpenAI lands at a consistent second across the field, with a median rank of 2 (Q1 2, Q3 2) over 159 ranked answers, and every one of the 7 models reports the same median of 2. The agreement is unusually tight: each model frames OpenAI as a close runner-up to Anthropic rather than a distant one. The recurring theme is a split between recognized strengths — algorithmic problem solving, debugging, reasoning-heavy tasks, and the broadest tooling ecosystem — and the reason it stops just short of first, namely that Claude is perceived as more reliable on large-repo, multi-file agentic work. claude-opus-5, kimi-k3, and gemini-3.7-flash all lean on ecosystem breadth, citing sources like the Aider LLM leaderboards, LMArena, GitHub Copilot, and SWE-bench, while the gpt-5.6 variants and gpt-6-astra emphasize Codex tooling and API maturity, backed by SWE-bench and OpenAI Codex documentation.

It is a very close second and often better for algorithmic reasoning and competitive-programming style problems.

claude-opus-5

The second theme is variability: several models qualify their ranking by noting that the best OpenAI choice depends on the specific model, task, or workflow, which is the main reason it settles narrowly behind Anthropic rather than tying or leading. This shows up plainly in the gpt-5.6 reasonings, where the ranking is repeatedly described as task-dependent, and in gpt-6-astra's advice to test both providers on one's own repository rather than assume a universal winner.

They are particularly strong at algorithmic problem-solving and explaining code, though some developers find them slightly less consistent than Claude on large, multi-step engineering tasks.

kimi-k3

Per model summary

  • claude-opus-5
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    25/25

    GPT-5/o-series and Codex models are praised for algorithmic problem solving, debugging, and competitive-programming tasks with the broadest tooling ecosystem, but consistently placed a very close second to Anthropic due to Claude's edge on large-repo agentic coding.

  • gemini-3.7-flash
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    25/25

    Emphasizes GPT-4o and o1 reasoning models as industry-leading for algorithmic problem solving and debugging, with unmatched ecosystem integration into tools like GitHub Copilot and Cursor, while noting it is slightly edged out on nuanced multi-file/full-stack tasks.

  • gpt-5.6-luna
    brand score
    82
    first choice
    55
    median rank
    2.0
    present in
    20/20

    Describes OpenAI as a highly capable, versatile general-purpose option for code generation, debugging, and tool use, ranked slightly below Anthropic because coding quality varies by model and workflow.

  • gpt-5.6-sol
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    20/20

    Highlights strong code generation, debugging, and agentic/Codex tooling backed by a mature API ecosystem, ranking it just behind Anthropic since quality and model selection can vary by task.

  • gpt-5.6-terra
    brand score
    76
    first choice
    48
    median rank
    2.0
    present in
    19/20

    Frames OpenAI as an excellent, safe general-purpose default with strong reasoning, mature APIs, and broad tooling, ranked just below Anthropic because the best choice depends on the specific environment and workflow.

  • gpt-6-astra
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    25/25

    Positions OpenAI as a close second, strong for code generation, debugging, tests, and Codex-supported workflows with a major ecosystem advantage, while slightly favoring Anthropic for repository-heavy editing and recommending testing both on one's own repo.

  • kimi-k3
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    25/25

    Notes GPT/o-series models excel at algorithmic and reasoning-heavy tasks with the most mature ecosystem (GitHub Copilot, API), but rank just behind Anthropic as developers find Claude more reliable on complex, multi-file software engineering.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-56 sites · 75 references
LMArena coding leaderboard×25 · 33%

lmarena.ai

gemini-3.7-flash8 sites · 69 references
OpenAI Research×22 · 32%

openai.com

gpt-5.6-luna6 sites · 52 references
OpenAI×17 · 33%

openai.com

gpt-5.6-sol5 sites · 60 references
SWE-bench×20 · 33%

swebench.com

  • ×20/
gpt-5.6-terra7 sites · 55 references
OpenAI API documentation×19 · 35%

platform.openai.com

gpt-6-astra6 sites · 56 references
OpenAI — Introducing Codex×16 · 29%

openai.com

kimi-k37 sites · 75 references
LMArena×20 · 27%

lmarena.ai

  • ×20/

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · AI models for coding
#3

Google

also named Google DeepMind

57
brand score
32first choice
12345

Summary

Google occupies a stable third position in the models' coding recommendations, with a median rank of 3 (Q1 3, Q3 3) across 159 ranked answers from 7 models. Six of the seven models—claude-opus-5, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra, and kimi-k3—all converge on a median of 3, while gemini-3.7-flash sits slightly lower at a median rank of 4 (Q1 3, Q3 4). The models refer to the brand as both Google and Google DeepMind. The recurring theme behind this placement is a distinctive strength in large context windows paired with a perceived gap in day-to-day coding consistency relative to the top two brands.

The dominant point of agreement is that Gemini's very large context window—cited by several models as reaching one to two million tokens—makes it well suited to reasoning across entire repositories and documentation-heavy projects, reinforced by multimodal capabilities and Google Cloud, Android, and developer-ecosystem integration. This is backed by sources including the Google DeepMind Gemini model page (cited 19 times by claude-opus-5), LMArena (24 times by kimi-k3), SWE-bench, the Aider LLM leaderboards, and Gemini API and Code Assist documentation. The consistent caveat, and the reason Google lands third rather than higher, is that its agentic and iterative coding output is seen as less consistent than Anthropic's or OpenAI's, with gemini-3.7-flash also flagging occasional shortfalls in code precision on complex edge cases.

Gemini 2.5/3 Pro offers a huge context window that is great for reasoning across entire repositories, plus strong performance and generous free tiers via AI Studio and Gemini CLI.

claude-opus-5

They rank third because, while highly capable and improving rapidly, developer consensus generally places their coding output quality just behind Anthropic and OpenAI on day-to-day tasks.

kimi-k3

Per model summary

  • claude-opus-5
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    25/25

    Consistently highlights Gemini 2.5/3 Pro's very large context window for repository-wide reasoning, strong benchmarks, multimodal/frontend strength, and generous free-tier access via AI Studio and Gemini CLI, while ranking it a close third due to less consistent agentic code editing than Claude or GPT.

  • gpt-5.6-luna
    brand score
    58
    first choice
    33
    median rank
    3.0
    present in
    20/20

    Frames Google as a strong option for long-context, multimodal work and for developers in the Google Cloud/Android ecosystem, but consistently ranks it behind Anthropic and OpenAI due to less consistent coding quality across model versions and integrations.

  • gpt-5.6-sol
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    20/20

    Stresses Gemini's very large context windows for analyzing large repositories and documentation plus Google Cloud integration, while noting coding consistency varies across model tiers and trails the top two choices.

  • gpt-5.6-terra
    brand score
    57
    first choice
    32
    median rank
    3.0
    present in
    19/20

    Highlights large context windows, multimodal inputs, and Google Cloud/ecosystem integration as strengths, positioning coding performance as competitive but variable by model version and workflow, generally below the top two for demanding agent tasks.

  • gpt-6-astra
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    25/25

    Consistently recommends Gemini as a strong context-heavy and multimodal option ranked third for general coding, but potentially a first choice for workflows involving large documentation, many files, or Google's developer ecosystem.

  • kimi-k3
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    25/25

    Emphasizes Gemini's exceptionally large context window for reasoning across entire repositories and competitive benchmark scores and pricing, while ranking it third because developer mindshare and day-to-day coding consistency slightly trail Anthropic and OpenAI.

  • gemini-3.7-flash
    brand score
    46
    first choice
    27
    median rank
    4.0
    present in
    25/25

    Repeatedly emphasizes massive multi-million-token context windows for ingesting entire codebases and documentation, plus multimodal and ecosystem integration, while noting occasional shortfalls in code precision or idiomatic output on complex edge cases relative to top competitors.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-510 sites · 75 references
Google DeepMind Gemini model page×25 · 33%

deepmind.google

gpt-5.6-luna7 sites · 51 references
Gemini Code Assist×15 · 29%

cloud.google.com

gpt-5.6-sol11 sites · 60 references
Google Gemini API documentation×19 · 32%

ai.google.dev

gpt-5.6-terra8 sites · 55 references
Google AI for Developers×21 · 38%

ai.google.dev

gpt-6-astra5 sites · 51 references
Google — Gemini API documentation×32 · 63%

ai.google.dev

kimi-k36 sites · 67 references
LMArena×25 · 37%

lmarena.ai

  • ×25/
gemini-3.7-flash9 sites · 61 references
Google DeepMind×24 · 39%

deepmind.google

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · AI models for coding
#4

DeepSeek

35
brand score
22first choice
12345

Summary

DeepSeek settles at a median rank of 4 across all seven models (Q1 4, Q3 4) over 141 ranked answers, marking it as a consistent mid-table recommendation rather than a top pick. Agreement is notably tight: six of the seven models place it at a median of 4, with only gemini-3.7-flash landing higher at a median of 3 (Q1 3, Q3 3). The shared narrative pairs a strong upside — near-frontier coding quality at a fraction of the cost, with open weights for self-hosting — against a recurring limitation around less mature tooling, enterprise support, and reliability on complex agentic tasks. This value-versus-maturity framing runs through claude-opus-5, kimi-k3, and the GPT-family models alike, drawing heavily on sources such as the Aider LLM leaderboards, DeepSeek's API documentation and GitHub repositories, Hugging Face model cards, Artificial Analysis, and SWE-bench.

The cost-and-openness theme is where models converge most tightly, typically citing the V3 and R1 lineage as the anchor for their reasoning.

DeepSeek's V3 and R1 models deliver coding performance that rivals much more expensive proprietary models, making them exceptional value and a leading open-weight option.

kimi-k3

The offsetting theme — why it rarely climbs above rank 4 — centers on deployment effort, governance, and ecosystem maturity, drawn from the API documentation and GitHub-heavy source pool that the GPT models lean on.

I rank it fourth because its surrounding tools, enterprise support, and overall developer ecosystem are less mature than those of the leading providers.

gpt-5.6-sol

Gemini's higher placement reflects a heavier emphasis on benchmark parity and self-hosting appeal, supported by references to the SWE-bench Leaderboard, Hugging Face, and the LMSYS Chatbot Arena.

Per model summary

  • gemini-3.7-flash
    brand score
    49
    first choice
    28
    median rank
    3.0
    present in
    22/25

    Repeatedly presents DeepSeek as an open-weights powerhouse rivaling proprietary models on coding benchmarks at a fraction of the cost, ideal for self-hosting, with a less mature developer tooling ecosystem as its main limitation.

  • claude-opus-5
    brand score
    37
    first choice
    24
    median rank
    4.0
    present in
    25/25

    Consistently frames DeepSeek's V3/R1-lineage models as delivering near-frontier coding performance at a fraction of the cost with open weights for self-hosting, while trailing the top proprietary labs on complex agentic tasks and tooling polish.

  • gpt-5.6-luna
    brand score
    13
    first choice
    9
    median rank
    4.0
    present in
    7/20

    Emphasizes DeepSeek's capable coding performance and cost efficiency with open-weight availability, ranking it below top providers due to less predictable reliability, tooling consistency, and deployment/ecosystem maturity.

  • gpt-5.6-sol
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    20/20

    Consistently describes DeepSeek as offering capable coding and reasoning at competitive prices with open-weight/self-hosting options, ranking it lower because support, governance, privacy, and ecosystem maturity require more evaluation.

  • gpt-5.6-terra
    brand score
    31
    first choice
    21
    median rank
    4.0
    present in
    17/20

    Frames DeepSeek as a value-oriented option strong on cost efficiency, open weights, and self-hosting flexibility, placing it below leading providers because deployment, support, safety controls, and production tooling need more evaluation.

  • gpt-6-astra
    brand score
    38
    first choice
    25
    median rank
    4.0
    present in
    25/25

    Positions DeepSeek as attractive for cost efficiency and access to open model weights, ranking it below the leading integrated assistants because deployment choices, self-hosting, and workflow tooling require more setup effort.

  • kimi-k3
    brand score
    38
    first choice
    25
    median rank
    4.0
    present in
    25/25

    Consistently highlights DeepSeek's V3 and R1 models delivering near-frontier coding at very low cost with open weights for self-hosting, ranking it below the top labs due to less mature tooling, reliability, and consistency on complex tasks.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gemini-3.7-flash12 sites · 58 references
Hugging Face×20 · 34%

huggingface.co

claude-opus-57 sites · 75 references
Hugging Face DeepSeek model cards×24 · 32%

huggingface.co

gpt-5.6-luna7 sites · 18 references
DeepSeek×6 · 33%

deepseek.com

gpt-5.6-sol8 sites · 60 references
DeepSeek API documentation×17 · 28%

api-docs.deepseek.com

  • ×17/
gpt-5.6-terra7 sites · 49 references
DeepSeek GitHub×14 · 29%

github.com

gpt-6-astra3 sites · 55 references
DeepSeek-Coder repository×46 · 84%

github.com

kimi-k310 sites · 66 references
Aider LLM Leaderboards×17 · 26%

aider.chat

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · AI models for coding
#5

Mistral AI

also named Mistral

8
brand score
8first choice
12345

Summary

Mistral AI settles at a median rank of 5 (Q1 5, Q3 5) across 57 ranked answers from all 6 models, and the tight quartiles signal strong agreement: every model places it fifth, with only gemini-3.7-flash showing marginally more spread (Q1 4, Q3 5). The consensus is less about weakness than positioning. Models frame Mistral as a specialist and deployment-oriented pick rather than a top capability choice, repeatedly citing its coding-focused models Codestral and Devstral. claude-opus-5 and gemini-3.7-flash lean on the technical strengths — fast, cheap, open-weight models tuned for autocomplete and fill-in-the-middle — while noting they trail frontier models on hard multi-file work. These themes rest on sources such as the Hugging Face Devstral model card, Mistral AI docs, the Codestral announcement, and Artificial Analysis.

Codestral and Devstral are purpose-built coding models with permissive/open weights and excellent latency for fill-in-the-middle autocomplete, plus EU data-residency appeal.

claude-opus-5

The second recurring theme, most prominent in the gpt-5.6 family, is that Mistral appeals when openness, deployment flexibility, and European hosting matter more than maximizing raw coding performance. gpt-5.6-luna, gpt-5.6-sol, and gpt-5.6-terra consistently rank it fifth on the grounds of a less established ecosystem and weaker demonstrated results on demanding repository-scale tasks, drawing on Mistral AI documentation alongside benchmark references like Aider LLM Leaderboards and SWE-bench. gpt-6-astra echoes the completion-workflow angle, treating it as a targeted rather than general-purpose fit.

Mistral is a good choice when openness, deployment flexibility, and control over infrastructure matter more than achieving the strongest possible coding results.

gpt-5.6-luna

Per model summary

  • claude-opus-5
    brand score
    4
    first choice
    4
    median rank
    5.0
    present in
    5/25

    Consistently highlights Codestral and Devstral as fast, cheap, open-weight models good for autocomplete, fill-in-the-middle, and EU/on-prem deployment, while noting they lag frontier models on hard multi-file tasks, making Mistral a privacy/cost choice rather than a top capability pick.

  • gemini-3.7-flash
    brand score
    13
    first choice
    10
    median rank
    5.0
    present in
    12/25

    Emphasizes Codestral as an efficient, lightweight model optimized for low-latency code completion, fill-in-the-middle, and broad language support, ideal for IDE autocompletion and self-hosting, but ranked fifth because it trails frontier models on complex multi-step reasoning.

  • gpt-5.6-luna
    brand score
    18
    first choice
    18
    median rank
    5.0
    present in
    18/20

    Positions Mistral as appealing for open-weight options, deployment flexibility, and European hosting, but ranks it fifth because its coding ecosystem, tooling, and performance on complex multi-file tasks are less consistently strong than leading providers.

  • gpt-5.6-sol
    brand score
    10
    first choice
    10
    median rank
    5.0
    present in
    10/20

    Frames Mistral as a good option for deployment flexibility, European hosting, and open-weight/Codestral coding models, but ranks it below others due to a less established ecosystem and weaker demonstrated performance on demanding repository-scale tasks.

  • gpt-5.6-terra
    brand score
    6
    first choice
    5
    median rank
    5.0
    present in
    5/20

    Recommends Mistral when European hosting, deployment flexibility, or open-weight options matter, while generally preferring higher-ranked providers first for the hardest end-to-end software-engineering tasks.

  • gpt-6-astra
    brand score
    6
    first choice
    6
    median rank
    5.0
    present in
    7/25

    Values Mistral for specialized code completion and fill-in-the-middle workflows with deployment flexibility, ranking it fifth for broad complex coding assistance while noting it can fit targeted completion-oriented workloads (and advises checking licensing).

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-54 sites · 15 references
Artificial Analysis×5 · 33%

artificialanalysis.ai

gemini-3.7-flash6 sites · 29 references
Hugging Face×11 · 38%

huggingface.co

gpt-5.6-luna5 sites · 46 references
Mistral AI×19 · 41%

mistral.ai

gpt-5.6-sol6 sites · 30 references
Mistral AI documentation×11 · 37%

docs.mistral.ai

gpt-5.6-terra6 sites · 15 references
Mistral AI documentation×5 · 33%

docs.mistral.ai

gpt-6-astra2 sites · 14 references
Mistral AI — Codestral×7 · 50%

mistral.ai

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · AI models for coding

How this was measured

The question

  • Unaided brand recommendation question: “Which AI model would you recommend for coding?
  • The prompt asks each model to return exactly 5 brands, ranked 1 to 5, and for each one a reason for the recommendation and the sources that informed it.

Sampling

  • Models are not deterministic, so the question is asked over and over — 23 answers per model
  • Spellings of the same brand are normalized to the most commonly used form before counting

Measures of Position: Median and Quartiles

  • Median value indicates that in 50% of answers the brand held this position or higher.
  • Quartiles help visualize the spread of rankings per brand. Q1 indicates that 25% of answers had this rank or higher. Q3 means that 75% of answers ranked the brand as X or better.
  • Medians and quartiles are computed only from answers where the brand was present.

Coverage

  • Count of answers in which the given brand was named, as a percentage.

Brand score - Normalized Borda score

  • The brand score is calculated using a normalized Borda score. Borda count is a voting method: each ballot awards points by position instead of naming one winner. Our ranking responses from the LLM are always fixed to 5 answers, so we assign 100 points for rank 1, then 80, 60, 40, 20 — and 0 if the brand is missing.
Brand score=100ni=1n6ri5\text{Brand\ score} = \frac{100}{n} \sum_{i=1}^{n} \frac{6 - r_i}{5}

where

ri={1,,5if the brand appears6if the brand is not mentionedr_i = \begin{cases} 1,\dots,5 & \text{if the brand appears} \\ 6 & \text{if the brand is not mentioned} \end{cases}
  • This way we can calculate a common score for all brands mentioned across model responses. The score reflects both how high the brand was ranked and how frequently it was mentioned.

First choice score - adjusted MRR - top-rank indicator

  • The first choice score is based on Mean Reciprocal Rank (a common search-engine measure), but adapted to measure the position of a specific brand across repeated ranked responses.
  • Each occurrence receives a reciprocal-position score: 100 for 1st, 50 for 2nd, 33.3 for 3rd, 25 for 4th, 20 for 5th, and 0 when the brand is not mentioned. The scores are averaged across all responses.
First choice score=100ni=1nsi\text{First\ choice\ score} = \frac{100}{n} \sum_{i=1}^{n} s_i

where

si={1/riif the brand is present0if the brand is absents_i = \begin{cases} 1/r_i & \text{if the brand is present} \\ 0 & \text{if the brand is absent} \end{cases}
  • No rank below 1st place gets more than 50, so the score is heavily driven by first places.
Category · Business Podcasts

Which business podcast would you recommend?

9 AI models · 42 answers · ranked by score

ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0

  1. #1Acquired
    68
    brand score
    55first choice
    12345
  2. #2How I Built This
    66
    brand score
    54first choice
  3. #3Masters of Scale
    35
    brand score
    21first choice
  4. #4HBR IdeaCast
    34
    brand score
    26first choice
  5. #5The Tim Ferriss Show
    19
    brand score
    17first choice
Show 23 more
  1. #6Planet Money
    18
    brand score
    12first choice
  2. #7How I Built This with Guy Raz
    10
    brand score
    8first choice
  3. #8Harvard Business Review
    8
    brand score
    6first choice
  4. #9Masters of Scale with Reid Hoffman
    8
    brand score
    4first choice
  5. #10The Diary of a CEO
    7
    brand score
    5first choice
  6. #11NPR
    6
    brand score
    4first choice
  7. #12The GaryVee Audio Experience
    5
    brand score
    3first choice
  8. #13The Indicator from Planet Money
    2
    brand score
    1first choice
  9. #14The Wall Street Journal
    2
    brand score
    1first choice
  10. #15My First Million
    2
    brand score
    1first choice
  11. #16Wondery
    2
    brand score
    1first choice
  12. #17Bloomberg
    1
    brand score
    1first choice
  13. #18The Journal
    1
    brand score
    #20 1first choice
  14. #19Pivot
    1
    brand score
    #17 1first choice
  15. #20The $100 MBA
    1
    brand score
    1first choice
  16. #21HubSpot
    1
    brand score
    1first choice
  17. #22The Journal.
    1
    brand score
    1first choice
  18. #23The Prof G Pod
    1
    brand score
    1first choice
  19. #24Business Wars
    0
    brand score
    0first choice
  20. #25Freakonomics Radio
    0
    brand score
    0first choice
  21. #26Masters in Business
    0
    brand score
    0first choice
  22. #27Smart Passive Income
    0
    brand score
    0first choice
  23. #28The Indicator
    0
    brand score
    0first choice

Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.

Swipe left for the full report

Category · Business Podcasts

Model by model

Median rank each model gave each brand.

Scroll the table sideways to see every model →

claude-opus-5deepseek-v4.1-flashgemini-3.7-flashgpt-5.6-lunagpt-5.6-solgpt-5.6-terragpt-6-astrakimi-k3mistral-medium-3.5
Acquired2.01.02.01.01.02.03.01.0
How I Built This1.02.01.04.03.01.02.02.02.0
Masters of Scale4.03.03.52.54.03.05.04.03.0
HBR IdeaCast4.05.02.02.52.05.01.05.05.0
The Tim Ferriss Show5.04.04.04.51.0
Planet Money3.04.03.05.03.54.0
How I Built This with Guy Raz1.02.0
Harvard Business Review1.52.5
Masters of Scale with Reid Hoffman2.03.0
The Diary of a CEO4.55.04.05.05.04.03.0
NPR1.53.0
The GaryVee Audio Experience5.04.0
The Indicator from Planet Money3.04.0
The Wall Street Journal3.5
My First Million5.03.0
Wondery4.0
Bloomberg4.5
The Journal3.0
Pivot5.0
The $100 MBA5.0
HubSpot5.0
The Journal.5.0
The Prof G Pod5.0
Business Wars5.0
Freakonomics Radio5.0
Masters in Business5.0
Smart Passive Income5.0
The Indicator5.0
1st
2
3
4
5th

A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.

Category · Business Podcasts

Where the models disagree

5 of 28 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.

  1. #1Acquired

    mistral-medium-3.5 0 gpt-5.6-luna 100 · 1 of 9 never named it

  2. #5The Tim Ferriss Show

    gemini-3.7-flash 0 mistral-medium-3.5 100 · 4 of 9 never named it

  3. #4HBR IdeaCast

    deepseek-v4.1-flash 4 gpt-5.6-sol 85

  4. #2How I Built This

    gpt-5.6-luna 30 claude-opus-5 100

  5. #7How I Built This with Guy Raz

    claude-opus-5 0 mistral-medium-3.5 48 · 7 of 9 never named it

  6. #3Masters of Scale

    claude-opus-5 8 gpt-5.6-luna 65

A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.

Category · Business Podcasts

Named, and named first

Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.

the default answera narrow favouritelisted, rarely ledthe long tail12345678
  1. 1Acquired78/42
  2. 2How I Built This82/39
  3. 3Masters of Scale73/0
  4. 4HBR IdeaCast60/9
  5. 5The Tim Ferriss Show36/20
  6. 6Planet Money42/0
  7. 7How I Built This with Guy Raz11/50
  8. 8Harvard Business Review10/25

The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.

Category · Business Podcasts

What the models actually said

Models
9
Answers
42

Acquired leads with a score of 68 of 100, ranked by 8 of 9 models.

The business podcast category divides into a small tier of broadly recognised shows and a long tail of brands that surface in a single model's list. Two titles anchor the top: Acquired (Borda 67.56, ranked by 8 of 9) and How I Built This (66.22, ranked by all 9). Both draw near-universal recognition, and both split the panel not on quality but on placement.

The sharpest fault line runs through Acquired. Models agree on its deeply researched, long-form company histories, but disagree on whether episode length keeps it out of the top slot.

gpt-5.6-luna
100
rates it highest
mistral-medium-3.5
0
never named it

Acquired · 100 points apart on a 0–100 scale

How I Built This divides along a parallel axis — narrative accessibility versus analytical depth. claude-opus-5 placed it first in every run for a perfect 100; gpt-5.6-luna settled it at rank 4 for a Borda of 30, judging it better for inspiration than rigour. This mirrors Acquired's pattern: the models converge on what a show is and diverge on what to weigh it against.

That same tension governs the middle tier. HBR IdeaCast (34.44) is the clearest case: gpt-5.6-sol and gpt-6-astra treat its research-grounded management substance as the strongest all-around pick (Borda 85 and 84), while deepseek-v4.1-flash, mistral-medium-3.5 and gpt-5.6-terra rank it near last (4, 8, 10), citing an academic tone and short format. The disagreement is entirely about weighting rigour against listenability, not about what the show delivers.

HBR IdeaCast

per-model scores 485 of 100 · mean 34 across 9 models

Two brands are pushed high by a single dissenting model. The Tim Ferriss Show (18.89) owes its standing almost entirely to mistral-medium-3.5, which ranked it first in all five runs; the other four models that named it placed it fourth or fifth on grounds of topical fit — whether a show ranging across health and self-optimisation counts as business at all. mistral-medium-3.5 also lifts The GaryVee Audio Experience and, uniquely, ranks a cluster of solo titles (Smart Passive Income, The $100 MBA, Masters of Scale with Reid Hoffman), suggesting a broader, entrepreneur-facing frame than its peers.

Where the panel does agree without caveat is at the bottom. Masters of Scale (34.67) and Planet Money (18.44) both draw the "broad agreement" label — every model that ranks them places them mid-table, discounting scaling-specific or news-explainer scope against direct operator guidance. Below that, a long tail of single-model brands (Pivot, Bloomberg, Freakonomics Radio, Masters in Business, Business Wars, The Journal, The Prof G Pod) carries the "models agree" label, but this is agreement by absence rather than conviction.

A recurring artefact worth flagging: several brands appear twice under near-duplicate names — How I Built This and How I Built This with Guy Raz, Masters of Scale and …with Reid Hoffman, The Indicator and The Indicator from Planet Money, The Journal and The Journal. The split versions each sit low precisely because coverage fragments across the variants. This is a naming effect in the recalled data, not evidence that models see these as distinct shows.

The sourcing is strikingly uniform across the whole category and does little to explain the splits. Aggregators dominate — Apple Podcasts×50 and Spotify×31 lead by a wide margin — followed by each show's own site and publisher pages. Because nearly every model reaches for the same directory and official-page material, the recalled sources plausibly explain what the models know about each show, but not why they rank the same show so differently. That divergence tracks editorial judgement — rigour versus accessibility, focus versus breadth — more than any evidence base. The one exception where a source hints at a view is claude-opus-5 citing BBC News reporting on health claims behind its cautious read of The Diary of a CEO; that link is suggestive, not established by the data.

The panel agrees on what each podcast is and splits on how to value it — a judgement gap the shared, aggregator-heavy sources cannot account for.
Category · Business Podcasts

What shaped the answers

55 sources across 473 references, grouped by site from 141 recalled names. The top 5 carry 68% of them.

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Business Podcasts
#1

Acquired

68
brand score
55first choice
12345

Summary

Acquired scores a Borda average of 67.56 of 100, ranked by 8 of the 9 models, with a first-choice score of 55.19. The panel converges on what the show offers — deeply researched, long-form company histories that function as strategy case studies — but diverges sharply on where to place it, a split the report labels "sharply split" and reflected in a standard deviation of 33.51. That range runs from gpt-5.6-luna's borda of 100, ranking it first in all four runs, down to gemini-3.7-flash's 16 off a single 20% coverage run, with claude-opus-5, gpt-5.6-terra and gpt-6-astra clustering in the middle at second or third.

gpt-5.6-luna
100
rates it highest
mistral-medium-3.5
0
never named it

Acquired · 100 points apart on a 0–100 scale

The recurring theme behind both the praise and the hesitation is episode length: models value the multi-hour depth for founders, investors and strategy-minded listeners while flagging the time commitment as the reason to rank it just below a more accessible pick. The sourcing sits heavily on the brand's own material — Acquiredwww.acquired.fm×19 recurs across most models — supplemented by Apple Podcastspodcasts.apple.com/us/podcast/acquired/id1050462261×10, with kimi-k3 and claude-opus-5 also reaching to press and reference coverage.

Near-universal agreement on Acquired's analytical depth, but genuine disagreement on whether its length keeps it out of the top spot.

Its episodes function almost like mini-MBA case studies, and its production quality and host chemistry are consistently praised.

kimi-k3

The depth is a major strength, but its often very long episodes make it a less convenient default recommendation for a general listener.

gpt-6-astra

Per model summary

  • deepseek-v4.1-flash
    brand score
    72
    first choice
    67
    median rank
    1.0
    present in
    4/5

    Repeatedly highlights deeply researched, long-form company narratives that serve as strategy case studies, valuable for understanding business models and moats, with episode length noted as a limitation for casual listeners.

  • gpt-5.6-luna
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    4/4

    Consistently frames it as offering exceptionally detailed, well-researched long-form analysis of company strategy and finance, valuable for those seeking rigorous insight over quick motivational content.

  • gpt-5.6-sol
    brand score
    95
    first choice
    88
    median rank
    1.0
    present in
    4/4

    Repeatedly stresses exceptionally detailed, well-researched examinations of companies and strategic decisions, with unusually long episodes offset by depth for serious business students.

  • kimi-k3
    brand score
    92
    first choice
    80
    median rank
    1.0
    present in
    5/5

    Consistently describes exceptionally deep, multi-hour, well-researched company breakdowns functioning like mini-MBA case studies, praised for substance though demanding in length for casual listeners.

  • claude-opus-5
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    5/5

    Consistently emphasizes hosts Ben Gilbert and David Rosenthal's deeply researched, multi-hour company histories and unmatched strategic depth, noting episode length as a drawback and strong credibility among investors and operators.

  • gemini-3.7-flash
    brand score
    16
    first choice
    10
    median rank
    2.0
    present in
    1/5

    Points to exhaustive, multi-hour deep dives into company history, strategy, and financials by Gilbert and Rosenthal, positioning it as essential for investors and strategists.

  • gpt-5.6-terra
    brand score
    85
    first choice
    63
    median rank
    2.0
    present in
    4/4

    Consistently emphasizes rigorous, long-form analysis of companies, strategy, and market structure, rewarding listeners who want durable insight despite the larger time commitment.

  • gpt-6-astra
    brand score
    68
    first choice
    40
    median rank
    3.0
    present in
    5/5

    Repeatedly praises detailed company histories and analysis of competitive advantages and strategy, while noting that unusually long episodes make it a less convenient default recommendation.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

deepseek-v4.1-flash5 sites · 11 references
Acquired×3 · 27%

acquired.fm

gpt-5.6-luna2 sites · 7 references
Acquired×4 · 57%

acquired.fm

gpt-5.6-sol3 sites · 9 references
Acquired×4 · 44%

acquired.fm

kimi-k38 sites · 13 references
Acquired×5 · 38%

acquired.fm

claude-opus-57 sites · 15 references
Acquired×5 · 33%

acquired.fm

gemini-3.7-flash2 sites · 2 references
Acquired.fm×1 · 50%

acquired.fm

gpt-5.6-terra2 sites · 8 references
Acquired×4 · 50%

acquired.fm

gpt-6-astra1 site · 5 references
Acquired×5 · 100%

acquired.fm

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Business Podcasts
#2

How I Built This

66
brand score
54first choice
12345

Summary

How I Built This posts a Borda score of

Stat unavailable

and was ranked by all 9 of the 9 models, giving it broad reach even as opinions on its standing diverge. The per-model Borda spans 30 to 100 with a standard deviation of 24.83, which the report labels "models split," and that gap is real: claude-opus-5 placed it first in all five of its runs for a perfect 100, while gpt-5.6-luna settled it at a median rank 4 for a Borda of 30.

claude-opus-5
100
rates it highest
gpt-5.6-luna
30
rates it lowest

How I Built This · 70 points apart on a 0–100 scale

The disagreement traces to a single recurring theme — accessibility and narrative craft versus analytical depth. Models that reward Guy Raz's NPR-produced founder storytelling rank it at or near the top, while those weighing rigorous strategy analysis push it down, several explicitly seating it just below more study-focused shows like HBR IdeaCast or Acquired. The recalled sources cluster tightly under the show's own footprint: NPR appears across nearly every model, alongside Apple Podcasts, Wondery, Spotify, and Wikipedia. A first-choice score of 53.93 confirms it is frequently named near the top yet not uniformly led.

Near-universal recognition, but the models divide on whether its storytelling substitutes for analytical depth.

It's the safest all-around recommendation for someone wanting an engaging business podcast.

claude-opus-5

It is better for entrepreneurial inspiration than for rigorous analysis of markets, operations, or competitive strategy.

gpt-5.6-luna

Per model summary

  • claude-opus-5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Consistently emphasizes Guy Raz's NPR-produced, narrative-driven founder interviews with high production quality, a large back catalog of major companies, and broad accessibility, making it the safest all-around recommendation.

  • gemini-3.7-flash
    brand score
    60
    first choice
    60
    median rank
    1.0
    present in
    3/5

    Focuses on Guy Raz's narrative-driven, deep-dive interviews with founders of world-renowned companies, stressing storytelling depth, resilience, and universal accessibility.

  • gpt-5.6-terra
    brand score
    90
    first choice
    83
    median rank
    1.0
    present in
    4/4

    Presents it as a strong all-around choice with accessible, non-technical founder interviews and clear lessons, while noting it is less analytical than more study-focused shows.

  • deepseek-v4.1-flash
    brand score
    48
    first choice
    30
    median rank
    2.0
    present in
    3/5

    Highlights Guy Raz's candid, well-produced founder interviews that balance emotional origin stories with practical entrepreneurship lessons, praising broad accessibility.

  • gpt-6-astra
    brand score
    88
    first choice
    70
    median rank
    2.0
    present in
    5/5

    Consistently rates it as an accessible, engaging entry point for entrepreneurship, but caveats that retrospective success stories are inspirational rather than a reliable, actionable playbook, ranking it just below HBR IdeaCast.

  • kimi-k3
    brand score
    88
    first choice
    70
    median rank
    2.0
    present in
    5/5

    Emphasizes Guy Raz's engaging, well-produced founder interviews revealing the messy reality of company-building, calling it the most accessible/universal choice though more inspirational than analytical.

  • mistral-medium-3.5
    brand score
    32
    first choice
    20
    median rank
    2.0
    present in
    2/5

    Repeatedly cites NPR's inspiring entrepreneur stories with an engaging narrative style that is both educational and useful for aspiring business leaders.

  • gpt-5.6-sol
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    4/4

    Describes it as making entrepreneurship accessible through polished founder storytelling, strong on inspiration but offering less analytical depth than top-ranked choices.

  • gpt-5.6-luna
    brand score
    30
    first choice
    19
    median rank
    4.0
    present in
    3/4

    Consistently frames it as engaging and accessible for founder stories and inspiration, while noting it offers less rigorous analysis of strategy, markets, and operations.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-56 sites · 15 references
Apple Podcasts Business charts×5 · 33%

podcasts.apple.com

gemini-3.7-flash3 sites · 8 references
Apple Podcasts×3 · 38%

podcasts.apple.com

gpt-5.6-terra2 sites · 8 references
Apple Podcasts×4 · 50%

podcasts.apple.com

deepseek-v4.1-flash3 sites · 8 references
Apple Podcasts×3 · 38%

podcasts.apple.com

gpt-6-astra2 sites · 5 references
kimi-k35 sites · 13 references
NPR×5 · 38%

npr.org

mistral-medium-3.53 sites · 6 references
Apple Podcasts×2 · 33%

no URL recalled

  • ×2named without a URL
gpt-5.6-sol4 sites · 9 references
gpt-5.6-luna2 sites · 5 references
NPR×3 · 60%

npr.org

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Business Podcasts
#3

Masters of Scale

35
brand score
21first choice
12345

Summary

All nine models ranked Masters of Scale, giving it a Borda score of

Stat unavailable

and a first-choice score of 21.39 of 100. No model placed it first in any run, and the per-model Borda spread runs from 8 to 65 with a standard deviation of 17.64 — a range the report labels "broad agreement." That spread tracks a consistent split in placement rather than in sentiment: gpt-5.6-luna sits it at a median rank of 2.5 while gpt-6-astra parks it at a median rank of 5 across all five of its runs.

gpt-5.6-luna
65
rates it highest
claude-opus-5
8
rates it lowest

Masters of Scale · 57 points apart on a 0–100 scale

The disagreement is one of degree rather than direction. Models converge on the same picture — Reid Hoffman's polished, thesis-driven interviews with prominent founders and executives, strong on growth and scaling themes — and then diverge on how much that framing counts against it. Several note the format can feel more inspirational, promotional, or theoretical than analytically rigorous, and gpt-6-astra's fifth-place placements rest specifically on its high-growth focus being less applicable to a general audience. Recalled sources cluster tightly around the show's own footprint: Masters of Scalemastersofscale.com×21 recurs across models, supported by directory and platform listings such as Apple Podcasts×50 and producer references to WaitWhat.

It balances inspiring stories with practical lessons about leadership, scaling, and decision-making.

gpt-5.6-luna

It ranks fifth for a general audience because its emphasis on scaling high-growth companies is less applicable to many everyday business situations.

gpt-6-astra

Per model summary

  • gpt-5.6-luna
    brand score
    65
    first choice
    40
    median rank
    2.5
    present in
    4/4

    Highlights thoughtful founder and executive interviews organized around clear growth principles, with polished production, while noting the inspirational framing can feel more polished than operational or analytical.

  • deepseek-v4.1-flash
    brand score
    32
    first choice
    18
    median rank
    3.0
    present in
    3/5

    Consistently frames it as valuable for scaling and leadership insights from prominent operators and founders, but notes some episodes feel promotional or more interview-driven than analytically rigorous.

  • gpt-5.6-terra
    brand score
    55
    first choice
    31
    median rank
    3.0
    present in
    4/4

    Describes polished, thematic conversations with recognizable operators on growth and scaling, useful for leadership insights but less analytically deep or granular than company-specific case-study shows.

  • mistral-medium-3.5
    brand score
    24
    first choice
    13
    median rank
    3.0
    present in
    2/5

    Frames it as Reid Hoffman exploring how companies grow from zero to a billion with insights from successful founders, while noting it can feel more theoretical than other shows.

  • gemini-3.7-flash
    brand score
    20
    first choice
    12
    median rank
    3.5
    present in
    2/5

    Focuses on Reid Hoffman deconstructing unconventional strategies for growing startups into large enterprises, praising high production and access to top executives as providing practical frameworks.

  • claude-opus-5
    brand score
    8
    first choice
    5
    median rank
    4.0
    present in
    1/5

    Emphasizes Reid Hoffman's scaling framework and strong, high-profile guest list with polished production, while noting the scripted, thesis-driven format can feel less authentic.

  • gpt-5.6-sol
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    4/4

    Points to prominent founders and strong production centered on growth and organizational challenges, but notes the content is thesis-driven and more inspirational and celebratory than critically probing or analytical.

  • kimi-k3
    brand score
    48
    first choice
    28
    median rank
    4.0
    present in
    5/5

    Emphasizes Reid Hoffman's LinkedIn credibility, strong guests, and polished production, but notes the scripted, thesis-driven format can feel promotional of his portfolio and less candid than long-form interviews.

  • gpt-6-astra
    brand score
    20
    first choice
    20
    median rank
    5.0
    present in
    5/5

    Recommends it for startup growth, leadership, and founder perspectives, but consistently ranks it fifth because its focus on high-growth companies is less applicable to a general or everyday-business audience.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gpt-5.6-luna3 sites · 7 references
Masters of Scale×4 · 57%

mastersofscale.com

deepseek-v4.1-flash3 sites · 8 references
Apple Podcasts×3 · 38%

podcasts.apple.com

gpt-5.6-terra2 sites · 8 references
Apple Podcasts×4 · 50%

podcasts.apple.com

mistral-medium-3.53 sites · 6 references
Apple Podcasts×2 · 33%

no URL recalled

  • ×2named without a URL
gemini-3.7-flash2 sites · 4 references
Apple Podcasts×2 · 50%

podcasts.apple.com

claude-opus-53 sites · 3 references
Apple Podcasts×1 · 33%

podcasts.apple.com

gpt-5.6-sol4 sites · 9 references
Masters of Scale×4 · 44%

mastersofscale.com

kimi-k35 sites · 10 references
Masters of Scale×5 · 50%

mastersofscale.com

  • ×4/
  • ×1Masters of Scale official site /
gpt-6-astra1 site · 5 references
Masters of Scale×5 · 100%

mastersofscale.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Business Podcasts
#4

HBR IdeaCast

34
brand score
26first choice
12345

Summary

HBR IdeaCast lands in the middle of the pack with a Borda score of

Stat unavailable

and a first-choice score of 25.93, ranked by all 9 of 9 models but with widely divergent enthusiasm. The panel divides along a consistent fault line: models converge on its credibility and research-grounded management substance, but split sharply on how much that matters against entertainment value. gpt-5.6-sol (borda 85) and gpt-6-astra (borda 84) treat it as the strongest all-around pick for practical management learning, while deepseek-v4.1-flash (borda 4), mistral-medium-3.5 (borda 8) and gpt-5.6-terra (borda 10) rank it near last, citing an academic tone and short, topic-led format that feels less immersive than founder-story shows.

gpt-5.6-sol
85
rates it highest
deepseek-v4.1-flash
4
rates it lowest

HBR IdeaCast · 81 points apart on a 0–100 scale

The recalled sources cluster tightly around the show's own home and directory listings — the Harvard Business Review — HBR IdeaCasthbr.org/podcasts/ideacast×3 page and Apple Podcastspodcasts.apple.com/us/podcast/hbr-ideacast/id152022135×8 listing recur across most models — which underlines that the disagreement is not about what the podcast is but about how to weigh rigor against listenability.

HBR IdeaCast

per-model scores 485 of 100 · mean 34 across 9 models

HBR IdeaCast is the strongest all-around recommendation because it delivers concise, research-informed conversations on leadership, strategy, innovation, and management.

gpt-5.6-sol

It ranks last here mainly because its academic tone and shorter format make it less entertaining and less story-driven than the others, even though the substance is solid.

kimi-k3

Per model summary

  • gpt-6-astra
    brand score
    84
    first choice
    73
    median rank
    1.0
    present in
    5/5

    Repeatedly names it the strongest all-around recommendation for practical leadership and management learning, especially for people managers, while ranking it below story-driven shows for entertainment.

  • gemini-3.7-flash
    brand score
    40
    first choice
    25
    median rank
    2.0
    present in
    3/5

    Emphasizes its Harvard Business Review pedigree, academic rigor, and research-backed, concise episodes offering actionable management insights for corporate professionals.

  • gpt-5.6-sol
    brand score
    85
    first choice
    63
    median rank
    2.0
    present in
    4/4

    Consistently rates it a top all-around pick for translating management research into concise, practical conversations with broad usefulness, while noting less depth per topic than narrative shows.

  • gpt-5.6-luna
    brand score
    35
    first choice
    21
    median rank
    2.5
    present in
    2/4

    Describes it as a credible, evidence-informed source on management and leadership with expert guests and concise episodes, though less entertaining and variable in quality by topic.

  • claude-opus-5
    brand score
    28
    first choice
    17
    median rank
    4.0
    present in
    3/5

    Consistently frames it as a credible, research-grounded management interview show with concise episodes, ideal for managers and professionals, while noting its dry, interview-heavy format lowers its entertainment value.

  • deepseek-v4.1-flash
    brand score
    4
    first choice
    4
    median rank
    5.0
    present in
    1/5

    Views it as credible and research-backed but ranks it lowest for general listeners due to its academic tone and shorter, less engaging format.

  • gpt-5.6-terra
    brand score
    10
    first choice
    10
    median rank
    5.0
    present in
    2/4

    Sees it as a dependable, research-informed source on management and leadership, but ranks it lower because its short, topic-led format feels less immersive than founder-story shows.

  • kimi-k3
    brand score
    16
    first choice
    13
    median rank
    5.0
    present in
    3/5

    Regards it as substantively solid and research-grounded management content, but consistently ranks it near last because its academic tone and brevity make it less entertaining and narrative-driven.

  • mistral-medium-3.5
    brand score
    8
    first choice
    8
    median rank
    5.0
    present in
    2/5

    Characterizes it as well-researched and credible HBR content that can feel more academic and less actionable or engaging than other options.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gpt-6-astra1 site · 5 references
Harvard Business Review — HBR IdeaCast×5 · 100%

hbr.org

gemini-3.7-flash2 sites · 6 references
Apple Podcasts×3 · 50%

podcasts.apple.com

gpt-5.6-sol3 sites · 9 references
Apple Podcasts×4 · 44%

podcasts.apple.com

gpt-5.6-luna2 sites · 3 references
Harvard Business Review×2 · 67%

hbr.org

claude-opus-52 sites · 6 references
Apple Podcasts×3 · 50%

podcasts.apple.com

deepseek-v4.1-flash3 sites · 3 references
Apple Podcasts×1 · 33%

podcasts.apple.com

gpt-5.6-terra2 sites · 4 references
Apple Podcasts×2 · 50%

podcasts.apple.com

kimi-k32 sites · 6 references
Apple Podcasts×3 · 50%

podcasts.apple.com

mistral-medium-3.53 sites · 6 references
Apple Podcasts×2 · 33%

no URL recalled

  • ×2named without a URL

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Business Podcasts
#5

The Tim Ferriss Show

19
brand score
17first choice
12345

Summary

The Tim Ferriss Show scores a Borda of

Stat unavailable

across the nine models, ranked by five of them but with a standard deviation the report labels "sharply split." That divergence turns on one recurring question: whether the show counts as a business podcast at all. mistral-medium-3.5 ranked it first in all five of its runs, praising the in-depth interviews and actionable insights, while claude-opus-5, kimi-k3, gpt-5.6-terra and deepseek-v4.1-flash all placed it lower on grounds of topical fit rather than quality — the show, in their reading, drifts into health, psychedelics, lifestyle and self-optimization, and the long episode formats dilute the business focus.

The Tim Ferriss Show

per-model scores 0100 of 100 · mean 19 across 9 models

The sources sit under two themes that mirror this split: the platform and canonical listings that establish reach — Apple Podcasts×50 and Spotify×31 — and the show's own presence via Tim Ferriss' blog and podcast pages, cited across mistral-medium-3.5, kimi-k3 and claude-opus-5. The agreement is on influence and guest quality; the disagreement is entirely on whether that value maps to business specifically.

Panellists broadly respect the show's depth and roster, but split hard on whether its wide scope qualifies it as a focused business recommendation.

Tim Ferriss consistently delivers high-quality, in-depth interviews with world-class performers across various fields, offering actionable insights.

mistral-medium-3.5

Ranked last here purely on topical fit rather than quality.

claude-opus-5

Per model summary

  • mistral-medium-3.5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Consistently praises Ferriss's high-quality, in-depth interviews with top performers and the actionable insights across diverse topics, presenting it as highly valuable for business and personal growth.

  • deepseek-v4.1-flash
    brand score
    16
    first choice
    10
    median rank
    4.0
    present in
    2/5

    Emphasizes that Ferriss extracts tactics and routines from top performers, but that the long, wide-ranging episodes dilute the business focus, lowering its ranking for pure business listeners.

  • gpt-5.6-terra
    brand score
    10
    first choice
    6
    median rank
    4.0
    present in
    1/4

    Highlights deep conversations with founders, investors, and thinkers yielding useful mental models, while noting its scope extends beyond business, making it less focused for purely business-oriented listeners.

  • kimi-k3
    brand score
    28
    first choice
    20
    median rank
    4.5
    present in
    4/5

    Repeatedly credits Ferriss with pioneering the long-form format and landing world-class guests with practical tactics, but ranks it lower because it drifts into lifestyle and self-optimization, with variable episode quality and very long formats.

  • claude-opus-5
    brand score
    16
    first choice
    16
    median rank
    5.0
    present in
    4/5

    Consistently frames it as a pioneering long-form interview show with an enormous archive of high-profile guests and tactical insights, but notes its scope drifts beyond business into health, psychedelics, and self-improvement, making it less focused and best browsed selectively.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

mistral-medium-3.53 sites · 15 references
Apple Podcasts×5 · 33%

no URL recalled

  • ×5named without a URL
deepseek-v4.1-flash4 sites · 6 references
Apple Podcasts×2 · 33%

podcasts.apple.com

gpt-5.6-terra2 sites · 2 references
Apple Podcasts×1 · 50%

podcasts.apple.com

kimi-k34 sites · 8 references
Tim Ferriss×4 · 50%

tim.blog

claude-opus-54 sites · 12 references
Apple Podcasts×4 · 33%

podcasts.apple.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Business Podcasts

How this was measured

The question

  • Unaided brand recommendation question: “Which business podcast would you recommend?
  • The prompt asks each model to return exactly 5 brands, ranked 1 to 5, and for each one a reason for the recommendation and the sources that informed it.

Sampling

  • Models are not deterministic, so the question is asked over and over — 5 answers per model
  • Spellings of the same brand are normalized to the most commonly used form before counting

Measures of Position: Median and Quartiles

  • Median value indicates that in 50% of answers the brand held this position or higher.
  • Quartiles help visualize the spread of rankings per brand. Q1 indicates that 25% of answers had this rank or higher. Q3 means that 75% of answers ranked the brand as X or better.
  • Medians and quartiles are computed only from answers where the brand was present.

Coverage

  • Count of answers in which the given brand was named, as a percentage.

Brand score - Normalized Borda score

  • The brand score is calculated using a normalized Borda score. Borda count is a voting method: each ballot awards points by position instead of naming one winner. Our ranking responses from the LLM are always fixed to 5 answers, so we assign 100 points for rank 1, then 80, 60, 40, 20 — and 0 if the brand is missing.
Brand score=100ni=1n6ri5\text{Brand\ score} = \frac{100}{n} \sum_{i=1}^{n} \frac{6 - r_i}{5}

where

ri={1,,5if the brand appears6if the brand is not mentionedr_i = \begin{cases} 1,\dots,5 & \text{if the brand appears} \\ 6 & \text{if the brand is not mentioned} \end{cases}
  • This way we can calculate a common score for all brands mentioned across model responses. The score reflects both how high the brand was ranked and how frequently it was mentioned.

First choice score - adjusted MRR - top-rank indicator

  • The first choice score is based on Mean Reciprocal Rank (a common search-engine measure), but adapted to measure the position of a specific brand across repeated ranked responses.
  • Each occurrence receives a reciprocal-position score: 100 for 1st, 50 for 2nd, 33.3 for 3rd, 25 for 4th, 20 for 5th, and 0 when the brand is not mentioned. The scores are averaged across all responses.
First choice score=100ni=1nsi\text{First\ choice\ score} = \frac{100}{n} \sum_{i=1}^{n} s_i

where

si={1/riif the brand is present0if the brand is absents_i = \begin{cases} 1/r_i & \text{if the brand is present} \\ 0 & \text{if the brand is absent} \end{cases}
  • No rank below 1st place gets more than 50, so the score is heavily driven by first places.
Category · Business Class airlines

Which airline would you recommend for business-class travel?

4 AI models · 5 answers per model · ranked by score

ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0

  1. #1Qatar Airways
    100
    brand score
    100first choice
    12345
  2. #2Singapore Airlines
    80
    brand score
    50first choice
  3. #3ANA
    48
    brand score
    28first choice
  4. #4Emirates
    41
    brand score
    25first choice
  5. #5Cathay Pacific
    14
    brand score
    10first choice
Show 4 more
  1. #6Delta Air Lines
    9
    brand score
    9first choice
  2. #7Japan Airlines
    4
    brand score
    3first choice
  3. #8EVA Air
    2
    brand score
    2first choice
  4. #9Turkish Airlines
    2
    brand score
    2first choice

Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.

Swipe left for the full report

Category · Business Class airlines

Model by model

Median rank each model gave each brand.

claude-opus-5gemini-3.7-flashgpt-6-astrakimi-k3
Qatar Airways1.01.01.01.0
Singapore Airlines2.02.02.02.0
ANA4.03.03.03.0
Emirates3.04.05.04.0
Cathay Pacific5.04.05.0
Delta Air Lines5.05.0
Japan Airlines4.0
EVA Air5.0
Turkish Airlines5.0
1st
2
3
4
5th

A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.

Category · Business Class airlines

Where the models disagree

0 of 9 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.

  1. #4Emirates

    gpt-6-astra 12 claude-opus-5 60

  2. #5Cathay Pacific

    claude-opus-5 0 gpt-6-astra 40 · 1 of 4 never named it

  3. #3ANA

    claude-opus-5 24 gpt-6-astra 60

  4. #6Delta Air Lines

    gpt-6-astra 0 claude-opus-5 20 · 2 of 4 never named it

A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.

Category · Business Class airlines

Named, and named first

Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.

the default answera narrow favouritelisted, rarely ledthe long tail12345678
  1. 1Qatar Airways100/100
  2. 2Singapore Airlines100/0
  3. 3ANA90/0
  4. 4Emirates90/0
  5. 5Cathay Pacific45/0
  6. 6Delta Air Lines45/0
  7. 7Japan Airlines10/0
  8. 8EVA Air10/0

The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.

Category · Business Class airlines

What the models actually said

Models
4
Answers
20

Qatar Airways leads with a score of 100 of 100, ranked by 4 of 4 models.

The category shows strong consensus at the top and widening disagreement further down. All four models place Qatar Airways at a median rank of 1 and Singapore Airlines at 2, with no model dissenting on either. The splits emerge among the middle and lower tiers, and much of the divergence is about how far a carrier trails rather than the direction of the judgment.

Where the models agree

Qatar and Singapore are locked. Every model lands on 1 and 2 respectively, and the reasoning converges too: the Qatar Qsuite is treated as the benchmark cabin, and Singapore's second place is consistently attributed to seat privacy and diagonal sleeping position rather than any soft-product weakness. The recalled sources reinforce the agreement — Skytrax World Airline Awards, The Points Guy, and One Mile at a Time appear across all four models for both carriers.

The one wrinkle at the top is framing rather than placement: gpt-6-astra shares Qatar's first rank but qualifies it, flagging that the Qsuite is not fleet-wide.

Qsuite is not available on every aircraft, so confirm the scheduled seat configuration before booking.

— gpt-6-astra

Where the models split

The clearest divergence is on Emirates. claude-opus-5 rates it highest at a median of 3, gemini-3.7-flash and kimi-k3 sit at 4, and gpt-6-astra is the low outlier at 5. All four cite the same tension — strong A380 and ground experience against the older 2-3-2 777 seating — but weigh the seating variability very differently. gpt-6-astra lets that flaw set the ceiling, while claude-opus-5 treats the new 777 suites as a mitigating factor.

It ranks fifth among these options because the seating experience varies substantially: some older Boeing 777 cabins lack direct aisle access for every passenger and offer less privacy than the leading alternatives.

— gpt-6-astra

ANA shows a milder split: three models place it at 3, but claude-opus-5 alone sits at 4. Notably, claude-opus-5 rates Emirates above ANA (3 vs 4), while gpt-6-astra reverses that ordering (ANA 3, Emirates 5). The underlying complaint for ANA is uniform — The Room is exceptional but flight-dependent — so the disagreement is about how heavily to penalize limited availability versus Emirates' seat inconsistency.

Cathay Pacific splits by degree among the three models that ranked it: gpt-6-astra at 4, gemini-3.7-flash and kimi-k3 at 5. kimi-k3 pushes it lowest on post-pandemic service and catering recovery, a concern the other models do not foreground.

Single-model tails

Several carriers were ranked by only one model, so they carry no cross-model signal:

  • Japan Airlines (rank 4) and Delta (rank 5) — Delta ranked by claude-opus-5 and gemini-3.7-flash, both at 5, but no other model placed it.
  • EVA Air (rank 5) — gpt-6-astra only.
  • Turkish Airlines (rank 5) — kimi-k3 only.

These reflect differences in which brands each model chose to recall as much as differences in evaluation. That claude-opus-5 surfaced Japan Airlines while kimi-k3 surfaced Turkish Airlines, for instance, may be a coverage difference rather than a disagreement about relative quality — but the data does not let us test that, since the other models did not rank these carriers at all.

On the sources

The recalled source base is highly uniform across the category: Skytrax World Airline Awards (57×), The Points Guy (45×), and One Mile at a Time (39×) dominate every brand. This makes it difficult to attribute any specific model's split to a distinctive source. One plausible pattern — and it is a hypothesis, not something the data confirms — is that the US-inflected placements lean on region-specific sources: Delta's ranking draws on J.D. Power North America Airline Satisfaction Study, which appears nowhere else in the category. Whether that source drives claude-opus-5 and gemini-3.7-flash to surface Delta at all, versus merely accompanying a judgment reached otherwise, cannot be determined from these numbers.

Category · Business Class airlines

What shaped the answers

35 sources across 293 references, grouped by site from 53 recalled names. The top 5 carry 77% of them.

Skytrax World Airline Awards×75 · 26%

worldairlineawards.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Business Class airlines
#1

Qatar Airways

100
brand score
100first choice
12345

Summary

Qatar Airways occupies an unambiguous top position in this category, holding a median rank of 1 (Q1 1, Q3 1) across all 20 ranked answers from 4 models, with each individual model also landing on a median rank of 1. Agreement is complete: no model placed it below first, and the reasoning behind that consensus is consistent, centering almost entirely on the Qsuite business-class product. Across claude-opus-5, gemini-3.7-flash, and kimi-k3, the Qsuite is repeatedly described as the benchmark for the cabin — fully enclosed suites with doors, double-bed and quad configurations, and dine-on-demand service — supplemented by the Al Mourjan lounge at Doha's Hamad International and repeated Skytrax World's Best Business Class wins. The recalled sources under these themes converge as well, with Skytrax World Airline Awards, The Points Guy, and One Mile at a Time appearing across all four models, alongside Business Traveller and Qatar's own site.

Its Qsuite business class offers fully enclosed suites with doors and quad configurations, widely regarded as the best business class product in the world.

— claude-opus-5

The one point of nuance comes from gpt-6-astra, which shares the top ranking but qualifies its endorsement by flagging that the flagship product is not universally deployed, advising travelers to verify the aircraft and seat map before booking.

Qsuite is not available on every aircraft, so confirm the scheduled seat configuration before booking.

— gpt-6-astra

Per model summary

  • claude-opus-5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Consistently centers on the Qsuite as the benchmark business-class product (enclosed suites with doors, quad configurations), alongside top-rated dining, the Doha/Al Mourjan lounge, global connectivity, and repeated Skytrax World's Best Business Class wins.

  • gemini-3.7-flash
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Repeatedly frames Qatar Airways as setting the industry benchmark via its Qsuite (privacy doors, double-bed configurations), combined with dine-on-demand fine dining and world-class lounge/ground service at Doha's Hamad International.

  • gpt-6-astra
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Recommends Qatar Airways as the top choice for long-haul business class based on Qsuite's privacy, flexible layouts, and dining, while consistently cautioning that Qsuite is not on every aircraft and advising travelers to check the seat map.

  • kimi-k3
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Describes the Qsuite as the benchmark for business class (enclosed suites, doors, double-bed/quad options, dine-on-demand), reinforced by consistent fleet-wide quality, the Doha hub/lounge, and Skytrax awards.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-54 sites · 15 references
Skytrax World Airline Awards×5 · 33%

worldairlineawards.com

gemini-3.7-flash5 sites · 15 references
Skytrax World Airline Awards×5 · 33%

worldairlineawards.com

gpt-6-astra5 sites · 15 references
Skytrax World’s Best Business Class Airlines×5 · 33%

worldairlineawards.com

kimi-k33 sites · 15 references
One Mile at a Time×5 · 33%

onemileatatime.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Business Class airlines
#2

Singapore Airlines

80
brand score
50first choice
12345

Summary

Singapore Airlines occupies a remarkably stable position across all four models, holding a median rank of 2 (Q1 2, Q3 2) over 20 ranked answers. Every model — claude-opus-5, gemini-3.7-flash, gpt-6-astra, and kimi-k3 — landed on the same median of 2, indicating broad agreement that the carrier is a near-top choice but not the single strongest recommendation. The shared themes are consistent: polished, world-renowned cabin service, exceptionally wide business-class seats, and the Book the Cook dining program, with several models also crediting the Changi hub experience. These recurring praises draw on sources such as the Skytrax World Airline Awards, One Mile at a Time, The Points Guy, and Condé Nast Traveler.

The disagreement is less about placement than about the reason it sits at second rather than first. Both gpt-6-astra and kimi-k3 explicitly frame it as trailing Qatar's Qsuite, citing seat privacy and sleeping configuration as the deciding factors, while claude-opus-5 and kimi-k3 emphasize the manual flip and diagonal sleeping position of the seat as its recurring drawback. Those hardware caveats, drawn from reviewer sources including One Mile at a Time and The Points Guy, temper otherwise strong soft-product assessments.

Extremely wide business class seats, excellent Book the Cook dining and famously polished cabin crew make it a benchmark carrier.

— claude-opus-5

It loses the top spot mainly because most of its business seats lack privacy doors and require sleeping diagonally, making it slightly less private than Qatar's Qsuite.

— kimi-k3

Per model summary

  • claude-opus-5
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    5/5

    Consistently praised for polished cabin service, very wide business seats, and Book the Cook dining, while repeatedly noting the drawback that some seats require manual flipping to lie flat and older configurations lack direct aisle access.

  • gemini-3.7-flash
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    5/5

    Emphasizes world-renowned cabin crew hospitality, wide lie-flat seating, the Book the Cook dining program, and consistency, along with the Changi hub experience.

  • gpt-6-astra
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    5/5

    Regards it as an excellent all-round choice for polished service, dining, and spacious seats, but consistently ranks it just behind Qatar because seat privacy and sleeping comfort vary by aircraft and trail the Qsuite.

  • kimi-k3
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    5/5

    Highlights exceptionally wide seats, polished service, and Book the Cook dining, while noting that seats require sleeping diagonally and often lack privacy doors, keeping it slightly behind Qatar's Qsuite.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-56 sites · 15 references
Skytrax World Airline Awards×5 · 33%

worldairlineawards.com

gemini-3.7-flash6 sites · 15 references
Skytrax World Airline Awards×5 · 33%

worldairlineawards.com

gpt-6-astra5 sites · 15 references
Skytrax World’s Best Business Class Airlines×5 · 33%

worldairlineawards.com

kimi-k35 sites · 15 references
Skytrax World Airline Awards×5 · 33%

worldairlineawards.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Business Class airlines
#3

ANA

also named ANA (All Nippon Airways), All Nippon Airways, ANA All Nippon Airways

48
brand score
28first choice
12345

Summary

ANA lands consistently in the upper-middle of the field, with a median rank of 3 across 18 ranked answers from 4 models (Q1 3, Q3 4). Three of the four models — gemini-3.7-flash, gpt-6-astra, and kimi-k3 — each place it at a median of 3, while claude-opus-5 sits slightly lower at a median of 4 (Q1 4, Q3 4). The models agree closely on the underlying reasoning: the standout attraction is The Room, described as one of the widest, door-equipped business suites in commercial aviation, paired with attentive Japanese hospitality and strong dining. The recurring caveat that keeps ANA from ranking higher is limited availability — the best seat appears only on select aircraft and routes, making the experience flight-dependent — alongside a network described as more Japan- or Asia-centric.

That tension between an exceptional hard product and inconsistent access runs through every model's assessment. gpt-6-astra ties its third-place ranking directly to that limitation, while kimi-k3 and claude-opus-5 point to both fleet inconsistency and a thinner global network. The recalled sources cluster around a small set of aviation-review outlets, most frequently Skytrax World Airline Awards, One Mile at a Time, and The Points Guy, which appear under both the praise for The Room and the notes on route availability.

ANA pairs attentive service and excellent Japanese dining with an outstanding seat on aircraft equipped with The Room. It ranks third because that exceptionally spacious, enclosed configuration is available only on selected aircraft, making the recommendation more flight-dependent.

— gpt-6-astra

It ranks third because the product is only on select aircraft and routes, and its global network is thinner than the Gulf and Singapore carriers.

— kimi-k3

Per model summary

  • gemini-3.7-flash
    brand score
    56
    first choice
    32
    median rank
    3.0
    present in
    5/5

    Repeatedly emphasizes 'The Room' as one of the widest, most private business suites on select 777-300ERs, paired with Japanese hospitality and refined dining, with the main caveat being limited availability of the top product.

  • gpt-6-astra
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    5/5

    Focuses on attentive service, excellent Japanese dining, and the standout spacious seat on aircraft with The Room, while ranking ANA third due to that cabin's limited availability making the experience flight-dependent.

  • kimi-k3
    brand score
    52
    first choice
    30
    median rank
    3.0
    present in
    5/5

    Consistently praises 'The Room' on 777s as one of the widest, door-equipped seats with strong Japanese service and catering, but notes fleet inconsistency, limited route availability, and a thinner network outside Asia-Pacific.

  • claude-opus-5
    brand score
    24
    first choice
    15
    median rank
    4.0
    present in
    3/5

    Consistently highlights 'The Room' as arguably the most spacious business seat with doors, plus excellent Japanese service and catering, while noting its limitation to select aircraft/routes and a more Japan/Asia-centric network.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gemini-3.7-flash7 sites · 15 references
Skytrax World Airline Awards×4 · 27%

worldairlineawards.com

gpt-6-astra6 sites · 15 references
One Mile at a Time×4 · 27%

onemileatatime.com

  • ×3/
  • ×1named without a URL
kimi-k33 sites · 12 references
Skytrax World Airline Awards×5 · 42%

worldairlineawards.com

claude-opus-54 sites · 9 references
One Mile at a Time×3 · 33%

onemileatatime.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Business Class airlines
#4

Emirates

41
brand score
25first choice
12345

Summary

Emirates settles in the upper-middle of the field, with an overall median rank of 4 (Q1 3, Q3 4) across 18 ranked answers from 4 models. Agreement on placement is fairly tight but not uniform: claude-opus-5 puts it highest at a median rank of 3 (Q1 3, Q3 3), gemini-3.7-flash and kimi-k3 both land at a median of 4, while gpt-6-astra is the outlier with a median of 5 (Q1 5, Q3 5). The narrative behind the numbers is remarkably consistent across all four — the brand is credited for its A380 onboard bar and lounge, ICE entertainment, chauffeur service, Dubai ground experience and vast network, then held back by one recurring flaw.

That flaw is the older 2-3-2 seating on much of the 777 fleet, which lacks direct aisle access and drives the perception of product inconsistency. This tension is what separates the more generous placements from gpt-6-astra's fifth-place ranking, where the seating variability weighs heaviest.

The A380 business cabin's 1-2-1 layout is good but older 777 cabins in 2-3-2 are a weak point, though the new 777 suites are much improved.

— claude-opus-5

It ranks fifth among these options because the seating experience varies substantially: some older Boeing 777 cabins lack direct aisle access for every passenger and offer less privacy than the leading alternatives.

— gpt-6-astra

The recalled sources cluster around industry review and awards bodies that underpin both the praise and the caveat: Skytrax World Airline Awards appears under every model, alongside One Mile at a Time and The Points Guy, with claude-opus-5 and gpt-6-astra also citing Emirates directly, and gemini-3.7-flash and kimi-k3 leaning on Business Traveller.

Per model summary

  • claude-opus-5
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    5/5

    Emirates is praised for its A380 onboard bar/lounge, strong entertainment, excellent Dubai ground experience and vast network, but consistently criticized for the older 2-3-2 777 business cabins lacking direct aisle access, which newer retrofits are improving.

  • gemini-3.7-flash
    brand score
    44
    first choice
    27
    median rank
    4.0
    present in
    5/5

    Emirates is highlighted for its glamorous A380 onboard bar/lounge, ICE entertainment, chauffeur service and extensive Dubai network, with the recurring caveat that inconsistent 2-3-2 seating on older 777s keeps it from ranking higher.

  • kimi-k3
    brand score
    48
    first choice
    28
    median rank
    4.0
    present in
    5/5

    Emirates is noted for its glamorous A380 onboard bar/lounge, strong soft product, chauffeur service and Dubai network, but repeatedly flagged for inconsistency due to older 2-3-2 777 seats without direct aisle access.

  • gpt-6-astra
    brand score
    12
    first choice
    12
    median rank
    5.0
    present in
    3/5

    Emirates is valued for its extensive network, entertainment and A380 onboard lounge, but ranks fifth because seating varies substantially by aircraft, with older 777 cabins offering less privacy and no direct aisle access for every seat.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-56 sites · 15 references
Emirates×4 · 27%

emirates.com

gemini-3.7-flash8 sites · 15 references
Business Traveller×4 · 27%

businesstraveller.com

kimi-k35 sites · 13 references
Business Traveller×4 · 31%

businesstraveller.com

gpt-6-astra4 sites · 9 references
One Mile at a Time×3 · 33%

onemileatatime.com

  • ×2/
  • ×1named without a URL

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Business Class airlines
#5

Cathay Pacific

14
brand score
10first choice
12345

Summary

Cathay Pacific occupies a mid-table position across the three models, landing at a median rank of 4 (Q1 4, Q3 5) over 9 ranked answers. The models agree on the underlying strengths — comfortable reverse-herringbone seating, the newer Aria Suite, and highly regarded Hong Kong lounges such as The Pier — but diverge on placement. gpt-6-astra holds it steadiest at a median of 4 across 5 answers, while gemini-3.7-flash and kimi-k3 both place it at 5. The gap is one of degree rather than direction: all three cite the same qualities, but weigh the shortcomings differently.

For gpt-6-astra, the ceiling is set by cabin generation, noting that older business-class cabins offer less privacy than leading enclosed suites and that the Aria Suite is not fleet-wide. kimi-k3 pushes it lower on service and recovery concerns, pointing to an aging hard product and inconsistent post-pandemic performance. The recalled sources cluster around Skytrax World Airline Awards, The Points Guy, and One Mile at a Time, with Cathay Pacific's own site and outlets like Business Traveller and Executive Traveller feeding the seat, lounge, and dining themes.

Its newer Aria Suite strengthens the offering, but varying cabin generations keep it below the top three for an unspecified itinerary.

— gpt-6-astra

It ranks last here because the hard product is aging, and service and catering have been inconsistent during its post-pandemic rebuild compared with the carriers above.

— kimi-k3

Per model summary

  • gpt-6-astra
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    5/5

    Consistently frames Cathay Pacific as a well-rounded long-haul option with comfortable seating and excellent Hong Kong lounges, ranking it fourth because older cabin generations offer less privacy than leading suites and the newer Aria Suite is not fleet-wide.

  • gemini-3.7-flash
    brand score
    4
    first choice
    4
    median rank
    5.0
    present in
    1/5

    Emphasizes the well-designed reverse-herringbone seats and new Aria Suite, along with top-rated Hong Kong lounges like The Pier and quality bedding and dining.

  • kimi-k3
    brand score
    12
    first choice
    12
    median rank
    5.0
    present in
    3/5

    Notes comfortable reverse-herringbone seats and excellent Hong Kong lounges like The Pier, but ranks it last/fifth citing an aging hard product, inconsistent post-pandemic service, and slower recovery compared with rivals.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gpt-6-astra6 sites · 15 references
Cathay Pacific×4 · 27%

cathaypacific.com

gemini-3.7-flash3 sites · 3 references
Executive Traveller×1 · 33%

executivetraveller.com

kimi-k33 sites · 7 references
Skytrax World Airline Awards×3 · 43%

worldairlineawards.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Business Class airlines

How this was measured

The question

  • Unaided brand recommendation question: “Which airline would you recommend for business-class travel?
  • The prompt asks each model to return exactly 5 brands, ranked 1 to 5, and for each one a reason for the recommendation and the sources that informed it.

Sampling

  • Models are not deterministic, so the question is asked over and over — 5 answers per model
  • Spellings of the same brand are normalized to the most commonly used form before counting

Measures of Position: Median and Quartiles

  • Median value indicates that in 50% of answers the brand held this position or higher.
  • Quartiles help visualize the spread of rankings per brand. Q1 indicates that 25% of answers had this rank or higher. Q3 means that 75% of answers ranked the brand as X or better.
  • Medians and quartiles are computed only from answers where the brand was present.

Coverage

  • Count of answers in which the given brand was named, as a percentage.

Brand score - Normalized Borda score

  • The brand score is calculated using a normalized Borda score. Borda count is a voting method: each ballot awards points by position instead of naming one winner. Our ranking responses from the LLM are always fixed to 5 answers, so we assign 100 points for rank 1, then 80, 60, 40, 20 — and 0 if the brand is missing.
Brand score=100ni=1n6ri5\text{Brand\ score} = \frac{100}{n} \sum_{i=1}^{n} \frac{6 - r_i}{5}

where

ri={1,,5if the brand appears6if the brand is not mentionedr_i = \begin{cases} 1,\dots,5 & \text{if the brand appears} \\ 6 & \text{if the brand is not mentioned} \end{cases}
  • This way we can calculate a common score for all brands mentioned across model responses. The score reflects both how high the brand was ranked and how frequently it was mentioned.

First choice score - adjusted MRR - top-rank indicator

  • The first choice score is based on Mean Reciprocal Rank (a common search-engine measure), but adapted to measure the position of a specific brand across repeated ranked responses.
  • Each occurrence receives a reciprocal-position score: 100 for 1st, 50 for 2nd, 33.3 for 3rd, 25 for 4th, 20 for 5th, and 0 when the brand is not mentioned. The scores are averaged across all responses.
First choice score=100ni=1nsi\text{First\ choice\ score} = \frac{100}{n} \sum_{i=1}^{n} s_i

where

si={1/riif the brand is present0if the brand is absents_i = \begin{cases} 1/r_i & \text{if the brand is present} \\ 0 & \text{if the brand is absent} \end{cases}
  • No rank below 1st place gets more than 50, so the score is heavily driven by first places.
Category · EV vehicles

Which EV brand would you recommend?

6 AI models · 5 answers per model · ranked by score

ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0

  1. #1Tesla
    89
    brand score
    78first choice
    12345
  2. #2Hyundai
    75
    brand score
    61first choice
  3. #3Kia
    48
    brand score
    28first choice
  4. #4Rivian
    25
    brand score
    16first choice
  5. #5BMW
    22
    brand score
    14first choice
Show 6 more
  1. #6Ford
    19
    brand score
    14first choice
  2. #7BYD
    15
    brand score
    10first choice
  3. #8Nissan
    3
    brand score
    3first choice
  4. #9Chevrolet
    2
    brand score
    2first choice
  5. #10Volvo
    1
    brand score
    1first choice
  6. #11Lucid
    1
    brand score
    1first choice

Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.

Swipe left for the full report

Category · EV vehicles

Model by model

Median rank each model gave each brand.

Scroll the table sideways to see every model →

claude-opus-5deepseek-v4.1-flashgemini-3.7-flashgpt-6-astrakimi-k3mistral-medium-3.5
Tesla2.01.01.03.01.01.0
Hyundai1.02.02.01.02.04.0
Kia3.03.03.02.03.0
Rivian5.05.04.04.03.0
BMW4.04.04.04.0
Ford5.04.05.05.04.54.0
BYD5.05.02.0
Nissan5.0
Chevrolet5.05.0
Volvo5.0
Lucid5.0
1st
2
3
4
5th

A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.

Category · EV vehicles

Where the models disagree

5 of 11 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.

  1. #2Hyundai

    mistral-medium-3.5 8 gpt-6-astra 100

  2. #7BYD

    claude-opus-5 0 mistral-medium-3.5 80 · 3 of 6 never named it

  3. #3Kia

    mistral-medium-3.5 0 gpt-6-astra 80 · 1 of 6 never named it

  4. #4Rivian

    gpt-6-astra 0 mistral-medium-3.5 60 · 1 of 6 never named it

  5. #5BMW

    deepseek-v4.1-flash 0 gemini-3.7-flash 44 · 2 of 6 never named it

  6. #1Tesla

    gpt-6-astra 60 mistral-medium-3.5 100

A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.

Category · EV vehicles

Named, and named first

Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.

the default answera narrow favouritelisted, rarely ledthe long tail12345678
  1. 1Tesla100/63
  2. 2Hyundai87/37
  3. 3Kia73/0
  4. 4Rivian63/0
  5. 5BMW53/0
  6. 6Ford63/0
  7. 7BYD23/0
  8. 8Nissan17/0

The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.

Category · EV vehicles

What the models actually said

Models
6
Answers
30

Tesla leads with a score of 89 of 100, ranked by 6 of 6 models.

Tesla anchors the category, ranked by every model and scoring 88.67, but the panel divides on how much weight to give its well-documented caveats rather than on the caveats themselves. Four models put it first almost always; claude-opus-5 and gpt-6-astra pull it down, with gpt-6-astra sitting furthest away at a median rank 3, placing it behind Hyundai and Kia on interface and build-quality grounds.

Tesla

per-model scores 60100 of 100 · mean 89 across 6 models

The sharpest genuine split in the field is Hyundai, whose per-model scores run the full width from 8 to 100. Claude-opus-5 and gpt-6-astra rank it first outright, three models settle it second, and mistral-medium-3.5 names it just once at rank 4. The high scorers converge on the E-GMP platform and value story, so the divide is one of enthusiasm rather than of substance.

gpt-6-astra
100
rates it highest
mistral-medium-3.5
8
rates it lowest

Hyundai · 92 points apart on a 0–100 scale

Kia tracks Hyundai closely as its E-GMP near-peer, but the divergence there is coverage, not sentiment — gemini-3.7-flash names it in only 40% of runs while others hold it steadily at rank 3, and no model ranks it first. Rivian and BMW sit in similar territory: both are read consistently on the merits, with the spread driven by how far each model lets caveats (young-company risk for Rivian, price for BMW) pull the brand down, and by silent models scoring zero.

The most instructive disagreement is BYD, where mistral-medium-3.5's rank-2 conviction collides with two models placing it fifth. All three agree on the battery and value strengths; they split on whether limited Western availability should cap the ranking — a geography question, not an engineering one.

Across the field the models rarely disagree about what a brand is; they disagree about how heavily to price its known weaknesses.

The lower tier — Ford, Nissan, Chevrolet, Volvo, Lucid — shows "agreement" that is mostly omission. Ford is the exception, named by all six but never first, its rank-4-to-5 clustering a true consensus. The others register with one or two models each, so their low standard deviations reflect near-universal absence rather than a shared verdict.

The recalled source base is uniform enough that it explains little of the divergence: Car and Driverwww.caranddriver.com×32 and Edmundswww.edmunds.com×29 dominate nearly every brand, with Consumer Reportswww.consumerreports.org×23 recurring on reliability points. Because the same outlets sit behind both high and low placements, any link between a specific source and a model's harsher or softer read is a hypothesis the data does not confirm — the models appear to weight shared evidence differently rather than to draw on different evidence.

Category · EV vehicles

What shaped the answers

37 sources across 431 references, grouped by site from 141 recalled names. The top 5 carry 66% of them.

Car and Driver×92 · 21%

caranddriver.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · EV vehicles
#1

Tesla

89
brand score
78first choice
12345

Summary

Tesla scores 88.67 of 100 on Borda and was ranked by all 6 of the 6 models, placing it among the strongest recommendations in the category. The panel reaches broad agreement on its merits: the recurring theme across every model is the Supercharger network as the anchor of an EV ownership ecosystem, paired with energy efficiency, range, and over-the-air software updates. Where models diverge, they do so on how much weight to give the widely noted caveats — inconsistent build quality, minimalist touchscreen-heavy controls, and uneven service — rather than on whether those caveats exist.

Tesla

per-model scores 60100 of 100 · mean 89 across 6 models

That split is visible in the per-model placements. Four models (deepseek-v4.1-flash, gemini-3.7-flash, kimi-k3, mistral-medium-3.5) rank it first most or all of the time, treating build quality as a minor caveat within a seamless ecosystem. Claude-opus-5 settles it at a median rank 2, framing it as the top pick specifically for charging convenience and long-distance travel, while gpt-6-astra sits furthest away at a median rank 3, consistently ranking it below Hyundai and Kia because the interface and fit-and-finish trade-offs make it, in its words, a less universal recommendation. The recalled sources cluster consistently under these themes — Consumer Reportswww.consumerreports.org×23 and Edmunds recur behind the reliability and ownership caveats, while EPA fueleconomy.gov and InsideEVs underpin the efficiency and range claims.

Tesla remains the benchmark for EVs due to its Supercharger network, long range, and advanced software.

deepseek-v4.1-flash

I rank it below Hyundai and Kia because its touchscreen-heavy controls, reported build-quality inconsistencies, and uneven service experiences are meaningful trade-offs.

gpt-6-astra

Per model summary

  • deepseek-v4.1-flash
    brand score
    96
    first choice
    90
    median rank
    1.0
    present in
    5/5

    Repeatedly positions Tesla as the EV benchmark for most buyers based on its Supercharger network, range, performance, and over-the-air updates, while noting mixed build quality and service or reliability concerns.

  • gemini-3.7-flash
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Consistently calls Tesla the industry benchmark for its Supercharger network, energy efficiency, software integration, and OTA updates, treating variable build quality as a minor caveat within a seamless ownership ecosystem.

  • kimi-k3
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Repeatedly describes Tesla as the most complete and easiest EV to own thanks to efficiency, Supercharger access, software/OTA updates, competitive pricing, and resale value, while citing inconsistent build quality and service.

  • mistral-medium-3.5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Consistently frames Tesla as the leader in EV technology, battery innovation, and charging infrastructure via the Supercharger network, emphasizing performance, range, and software with no drawbacks noted.

  • claude-opus-5
    brand score
    76
    first choice
    47
    median rank
    2.0
    present in
    5/5

    Consistently frames Tesla as the top pick for charging convenience and long-distance travel due to its mature Supercharger network, efficiency, range, and software/OTA updates, while flagging inconsistent build quality, minimalist interiors, and variable service.

  • gpt-6-astra
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    5/5

    Highlights Tesla's efficiency, integrated route planning, and convenient charging for road trips, but consistently ranks it below Hyundai and Kia due to touchscreen-heavy controls, fit-and-finish variability, and uneven service.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

deepseek-v4.1-flash6 sites · 15 references
Car and Driver×4 · 27%

caranddriver.com

  • ×2/
  • ×1/tesla
  • ×1named without a URL
gemini-3.7-flash4 sites · 15 references
Car and Driver×5 · 33%

caranddriver.com

kimi-k36 sites · 15 references
Consumer Reports×5 · 33%

consumerreports.org

mistral-medium-3.53 sites · 15 references
Consumer Reports×5 · 33%

no URL recalled

  • ×5named without a URL
claude-opus-56 sites · 15 references
Consumer Reports×5 · 33%

consumerreports.org

gpt-6-astra5 sites · 14 references
Car and Driver×5 · 36%

caranddriver.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · EV vehicles
#2

Hyundai

75
brand score
61first choice
12345

Summary

Hyundai lands at a Borda score of

Stat unavailable

and was ranked by all 6 models, yet the panel is far from unanimous. Per-model borda spans 8 to 100 with a standard deviation of 31.28, which the report labels "sharply split." Two models (claude-opus-5 and gpt-6-astra) placed it first in all five of their runs, three others (deepseek-v4.1-flash, gemini-3.7-flash, kimi-k3) settled it at a median rank of 2, while mistral-medium-3.5 named it only once at rank 4 for a coverage of 20% and a borda of 8. The recurring theme across the high scorers is Hyundai's E-GMP platform: 800V ultra-fast charging, competitive pricing, long range, and a generous warranty, framed as the best all-around value alternative to Tesla.

gpt-6-astra
100
rates it highest
mistral-medium-3.5
8
rates it lowest

Hyundai · 92 points apart on a 0–100 scale

The sources under these themes cluster on mainstream automotive review outlets — Car and Driver, Edmunds, MotorTrend, IIHS, and Consumer Reports recur across models — with award citations such as World Car Awards backing the "well-rounded" framing. The caveats that hold Hyundai just behind the top are consistent too: less mature software, historically weaker charging network access (partly offset by NACS adoption), and ICCU recalls that several models flag as worth verifying before purchase. Mistral's lone, lower placement rests on a similar read but positions Hyundai as reliable rather than cutting-edge.

It's the best all-around balance of value, tech, and dependability for most buyers.

claude-opus-5

While not as cutting-edge as Tesla or BYD, it provides reliable and well-rounded options.

mistral-medium-3.5

Per model summary

  • claude-opus-5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Consistently highlights the E-GMP platform's 800V ultra-fast charging, strong efficiency/range, competitive pricing, and excellent warranty, framing Ioniq 5/6 as the best all-around value, with NACS access and awards noted; ICCU recalls and software glitches mentioned as caveats.

  • gpt-6-astra
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Frames Hyundai as the best all-around recommendation for balancing efficiency, comfort, value, and fast charging, while repeatedly cautioning that charging performance varies by model and buyers should verify recall completion.

  • deepseek-v4.1-flash
    brand score
    84
    first choice
    60
    median rank
    2.0
    present in
    5/5

    Emphasizes the E-GMP platform's 800V ultra-fast charging, long range, competitive pricing, and generous warranty, positioning Hyundai as a well-rounded value choice; a smaller charging network is noted as the main improving downside.

  • gemini-3.7-flash
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    5/5

    Repeatedly centers on the advanced 800V E-GMP platform enabling ultra-fast charging at mainstream prices, paired with distinctive styling, generous warranty, and strong overall value.

  • kimi-k3
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    5/5

    Focuses on the E-GMP 800V platform's very fast charging, distinctive design, long battery warranty, and price advantage, ranking Hyundai just behind Tesla due to less mature software and charging network access.

  • mistral-medium-3.5
    brand score
    8
    first choice
    5
    median rank
    4.0
    present in
    1/5

    Describes Hyundai as offering a diverse EV lineup with solid range and competitive pricing, reliable and well-rounded but not as cutting-edge as Tesla or BYD.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-56 sites · 15 references
Car and Driver×5 · 33%

caranddriver.com

gpt-6-astra4 sites · 11 references
Car and Driver×5 · 45%

caranddriver.com

deepseek-v4.1-flash5 sites · 15 references
Car and Driver×5 · 33%

caranddriver.com

gemini-3.7-flash5 sites · 15 references
MotorTrend×5 · 33%

motortrend.com

kimi-k35 sites · 15 references
Car and Driver×4 · 27%

caranddriver.com

mistral-medium-3.53 sites · 3 references
Hyundai Official Website×1 · 33%

hyundai.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · EV vehicles
#3

Kia

48
brand score
28first choice
12345

Summary

Kia lands at

Stat unavailable

on Borda, ranked by 5 of the 6 models, but the panel does not settle on a single reading. Per-model borda spans 0–80 with a standard deviation of 27.23, which the report labels "models split." The divide is largely one of coverage rather than sentiment: gpt-6-astra places it consistently at median rank 2 for a borda of 80, while gemini-3.7-flash ranks it in only 2 of 5 runs, holding coverage to 40% and borda to 24. The three remaining models cluster tightly at median rank 3, scoring between 60 and 64.

Kia

per-model scores 080 of 100 · mean 48 across 6 models

The recurring theme is that Kia is a near-peer to Hyundai on the shared 800V E-GMP platform, valued for the three-row EV9, fast charging, and warranty coverage, but held just behind on pricing at higher trims, inconsistent dealer experience, and availability. Those points recur across the models' recalled sources, with Car and Driver×10 and Edmunds×6 appearing most often, alongside award citations and reliability references such as J.D. Power, IIHS, and Consumer Reports. First-choice score sits at 27.78, and no model ranked Kia first in any run, consistent with the repeated framing of it as a strong second or co-equal rather than the lead pick.

Kia is a close second, offering compelling electric family vehicles with practical interiors, strong warranties, and excellent charging speeds on its newer dedicated EV platforms.

gpt-6-astra

Essentially a co-equal to Hyundai, ranked just behind on value per dollar.

claude-opus-5

Per model summary

  • gpt-6-astra
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    5/5

    Kia is consistently a close second to Hyundai, strongest for families needing space and three-row seating, but placed behind because larger, better-equipped models become expensive and value depends on needs.

  • claude-opus-5
    brand score
    64
    first choice
    37
    median rank
    3.0
    present in
    5/5

    Kia is framed as a co-equal to Hyundai on the shared 800V E-GMP platform, praised for the three-row EV9, fast charging, and strong warranty, but ranked just behind Hyundai due to rising prices on top trims and inconsistent dealer experience.

  • deepseek-v4.1-flash
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    5/5

    Kia is positioned as sharing Hyundai's E-GMP platform while adding bolder styling and value, with the three-row EV9 as a standout, though dealer markups and availability are noted drawbacks.

  • gemini-3.7-flash
    brand score
    24
    first choice
    13
    median rank
    3.0
    present in
    2/5

    Kia is highlighted for exceptional value, fast charging, and family-friendly practicality, with models consistently earning awards for ergonomics, range, and price-to-performance.

  • kimi-k3
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    5/5

    Kia shares Hyundai's excellent E-GMP platform with strong styling and the standout three-row EV9, but ranks just below Hyundai due to higher pricing on top trims and inconsistent dealer/availability.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gpt-6-astra2 sites · 10 references
Car and Driver×5 · 50%

caranddriver.com

claude-opus-57 sites · 15 references
Edmunds×4 · 27%

edmunds.com

deepseek-v4.1-flash8 sites · 15 references
Car and Driver×5 · 33%

caranddriver.com

  • ×3/
  • ×1/kia
  • ×1named without a URL
gemini-3.7-flash4 sites · 6 references
Edmunds×2 · 33%

edmunds.com

kimi-k37 sites · 15 references
Car and Driver×3 · 20%

caranddriver.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · EV vehicles
#4

Rivian

25
brand score
16first choice
12345

Summary

Rivian lands at a modest

Stat unavailable

and is named by 5 of the 6 models, yet the panel is far from settled on where it belongs. Per-model borda runs from 0 to 60 with a standard deviation of 20.58, which the report labels "models split" — mistral-medium-3.5 ranked it median 3 across all five runs, while claude-opus-5 and deepseek-v4.1-flash placed it at median rank 5 where they named it at all. The gap reflects a shared read applied with different weight rather than genuine disagreement about the product: every model describes the R1T and R1S as capable, adventure-focused vehicles with strong software, then discounts them for premium pricing, a thin service network, and the financial and reliability uncertainty of a young automaker.

Rivian

per-model scores 060 of 100 · mean 25 across 6 models

The evidence sits under a consistent set of automotive review sources, with MotorTrendwww.motortrend.com×21 and Edmundswww.edmunds.com×29 recurring across models, alongside Consumer Reports where reliability risk is raised. What separates the rankings is how far each model lets the caveats pull the brand down: for the more cautious models the young-company risk and thin support network reduce it to a niche recommendation, while mistral-medium-3.5 treats the same limitations as constraints on breadth rather than on quality. No model ranked it first in any run.

The R1T and R1S are arguably the best-driving adventure EVs available, with excellent software, clever packaging, and some of the highest owner-satisfaction scores in the industry.

kimi-k3

Rivian excels in adventure-focused EVs with impressive off-road capabilities and a strong brand identity.

mistral-medium-3.5

Per model summary

  • mistral-medium-3.5
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    5/5

    Rivian excels in adventure-focused EVs with off-road capability and premium build quality, though a limited model lineup, higher prices, and still-growing production scale restrict broader appeal.

  • gemini-3.7-flash
    brand score
    40
    first choice
    26
    median rank
    4.0
    present in
    5/5

    Rivian excels in premium, adventure-focused trucks and SUVs with off-road capability, refined interiors, and strong software, but higher prices and a still-expanding service and charging network limit its appeal for everyday drivers.

  • kimi-k3
    brand score
    28
    first choice
    19
    median rank
    4.0
    present in
    4/5

    The R1T and R1S are described as outstanding adventure EVs with clever design and high owner satisfaction, but high prices, a thin service network, and financial/reliability risks for a young company make it a compelling yet riskier choice.

  • claude-opus-5
    brand score
    8
    first choice
    8
    median rank
    5.0
    present in
    2/5

    Rivian's R1T and R1S are praised as capable, well-engineered adventure vehicles with strong software and loyal owners, but the recommendation is qualified by premium pricing, thin service networks, and young-automaker financial and reliability uncertainty.

  • deepseek-v4.1-flash
    brand score
    12
    first choice
    12
    median rank
    5.0
    present in
    3/5

    The R1T and R1S are praised for off-road capability and performance, but high prices and a still-growing service network make Rivian a niche choice for adventure enthusiasts rather than most buyers.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

mistral-medium-3.55 sites · 15 references
Rivian Official Site×5 · 33%

rivian.com

gemini-3.7-flash8 sites · 15 references
MotorTrend×4 · 27%

motortrend.com

kimi-k36 sites · 12 references
Consumer Reports×3 · 25%

consumerreports.org

claude-opus-53 sites · 6 references
Consumer Reports×2 · 33%

consumerreports.org

  • ×1/
  • ×1Consumer Reports reliability rankings /cars
deepseek-v4.1-flash4 sites · 9 references
Car and Driver×3 · 33%

caranddriver.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · EV vehicles
#5

BMW

22
brand score
14first choice
12345

Summary

BMW lands at a Borda score of 22 of 100, held down by being ranked by only 4 of the 6 models rather than by any hostility from those that did name it. Among the four that ranked it, placement is strikingly stable: gemini-3.7-flash gave it 44, claude-opus-5 and gpt-6-astra both 40, with each of those three covering it in every run and consistently settling on a median rank of 4. The gap that the report labels "models split" comes almost entirely from the two silent models scoring 0 and from kimi-k3's thin 20% coverage yielding a borda of 8.

BMW

per-model scores 044 of 100 · mean 22 across 6 models

The themes are consistent across the models that engaged with it: refined driving dynamics, quiet and well-built premium cabins, and above-average reliability, offset uniformly by higher purchase prices, costly options, and 400V-class charging that trails the 800V Korean rivals. That is why the brand clusters at fourth rather than higher — the models frame the ranking as a value trade-off, not an EV weakness. The recalled sources reflect this enthusiast-and-ownership lens, with Car and Driver×10 and Edmunds×6 recurring alongside Consumer Reports×7 on the reliability point.

The models agree BMW is a strong premium pick whose price, not its engineering, keeps it at fourth.

BMW is my strongest premium pick here for refined cabins, ride quality, and engaging driving dynamics across its electric range.

gpt-6-astra

They rank lower mainly on value: prices are high, options add up quickly, and charging speeds are good rather than class-leading.

kimi-k3

Per model summary

  • claude-opus-5
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    5/5

    Consistently frames BMW's i4/iX/i5 as refined, well-built premium EVs with above-average reliability and traditional luxury feel, while noting higher costs and slower charging than 800V Korean rivals.

  • gemini-3.7-flash
    brand score
    44
    first choice
    27
    median rank
    4.0
    present in
    5/5

    Repeatedly emphasizes that BMW translates its signature driving dynamics and luxury craftsmanship into EVs with quiet, refined cabins, but cites higher prices and lower efficiency/slower charging as drawbacks.

  • gpt-6-astra
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    5/5

    Consistently positions BMW as a premium pick for refinement, cabin quality, and engaging driving dynamics, ranking it fourth mainly due to high purchase prices and costly options rather than any EV weakness.

  • kimi-k3
    brand score
    8
    first choice
    5
    median rank
    4.0
    present in
    1/5

    Describes BMW's i4 and iX as refined, well-built premium EVs with respectable range, ranking lower on value due to high prices, costly options, and merely good charging speeds.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-56 sites · 15 references
Car and Driver×5 · 33%

caranddriver.com

gemini-3.7-flash7 sites · 15 references
Car and Driver×4 · 27%

caranddriver.com

gpt-6-astra2 sites · 10 references
Car and Driver×5 · 50%

caranddriver.com

kimi-k33 sites · 3 references
Car and Driver×1 · 33%

caranddriver.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · EV vehicles

How this was measured

The question

  • Unaided brand recommendation question: “Which EV brand would you recommend?
  • The prompt asks each model to return exactly 5 brands, ranked 1 to 5, and for each one a reason for the recommendation and the sources that informed it.

Sampling

  • Models are not deterministic, so the question is asked over and over — 5 answers per model
  • Spellings of the same brand are normalized to the most commonly used form before counting

Measures of Position: Median and Quartiles

  • Median value indicates that in 50% of answers the brand held this position or higher.
  • Quartiles help visualize the spread of rankings per brand. Q1 indicates that 25% of answers had this rank or higher. Q3 means that 75% of answers ranked the brand as X or better.
  • Medians and quartiles are computed only from answers where the brand was present.

Coverage

  • Count of answers in which the given brand was named, as a percentage.

Brand score - Normalized Borda score

  • The brand score is calculated using a normalized Borda score. Borda count is a voting method: each ballot awards points by position instead of naming one winner. Our ranking responses from the LLM are always fixed to 5 answers, so we assign 100 points for rank 1, then 80, 60, 40, 20 — and 0 if the brand is missing.
Brand score=100ni=1n6ri5\text{Brand\ score} = \frac{100}{n} \sum_{i=1}^{n} \frac{6 - r_i}{5}

where

ri={1,,5if the brand appears6if the brand is not mentionedr_i = \begin{cases} 1,\dots,5 & \text{if the brand appears} \\ 6 & \text{if the brand is not mentioned} \end{cases}
  • This way we can calculate a common score for all brands mentioned across model responses. The score reflects both how high the brand was ranked and how frequently it was mentioned.

First choice score - adjusted MRR - top-rank indicator

  • The first choice score is based on Mean Reciprocal Rank (a common search-engine measure), but adapted to measure the position of a specific brand across repeated ranked responses.
  • Each occurrence receives a reciprocal-position score: 100 for 1st, 50 for 2nd, 33.3 for 3rd, 25 for 4th, 20 for 5th, and 0 when the brand is not mentioned. The scores are averaged across all responses.
First choice score=100ni=1nsi\text{First\ choice\ score} = \frac{100}{n} \sum_{i=1}^{n} s_i

where

si={1/riif the brand is present0if the brand is absents_i = \begin{cases} 1/r_i & \text{if the brand is present} \\ 0 & \text{if the brand is absent} \end{cases}
  • No rank below 1st place gets more than 50, so the score is heavily driven by first places.
Category · Watch brands

Which watch brand would you wear on a first date?

5 AI models · 5 answers per model · ranked by score

ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0

  1. #1Cartier
    78
    brand score
    55first choice
    12345
  2. #2Omega
    77
    brand score
    63first choice
    • claude-opus-5 3.0
  3. #3Seiko
    54
    brand score
    37first choice
  4. #4Rolex
    38
    brand score
    36first choice
  5. #5NOMOS Glashütte
    17
    brand score
    13first choice
Show 6 more
  1. #6Tudor
    16
    brand score
    9first choice
  2. #7Tissot
    13
    brand score
    8first choice
    • claude-opus-5 3.0
  3. #8Casio
    5
    brand score
    5first choice
  4. #9Apple
    1
    brand score
    1first choice
  5. #10Hublot
    1
    brand score
    1first choice
  6. #11Timex
    1
    brand score
    1first choice

Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.

Swipe left for the full report

Category · Watch brands

Model by model

Median rank each model gave each brand.

Scroll the table sideways to see every model →

claude-opus-5gemini-3.7-flashgpt-6-astrakimi-k3mistral-medium-3.5
Cartier2.02.02.02.03.0
Omega3.01.04.01.02.0
Seiko2.03.03.04.04.0
Rolex5.05.05.05.01.0
NOMOS Glashütte2.0
Tudor3.03.0
Tissot3.05.0
Casio5.05.0
Apple5.0
Hublot5.0
Timex5.0
1st
2
3
4
5th

A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.

Category · Watch brands

Where the models disagree

5 of 11 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.

  1. #4Rolex

    gpt-6-astra 16 mistral-medium-3.5 100

  2. #7Tissot

    gpt-6-astra 0 claude-opus-5 60 · 3 of 5 never named it

  3. #2Omega

    gpt-6-astra 44 gemini-3.7-flash 100

  4. #6Tudor

    claude-opus-5 0 kimi-k3 48 · 3 of 5 never named it

  5. #3Seiko

    kimi-k3 36 claude-opus-5 80

  6. #1Cartier

    mistral-medium-3.5 60 gpt-6-astra 88

A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.

Category · Watch brands

Named, and named first

Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.

the default answera narrow favouritelisted, rarely ledthe long tail12345678
  1. 1Cartier100/20
  2. 2Omega100/40
  3. 3Seiko96/12
  4. 4Rolex96/20
  5. 5NOMOS Glashütte20/40
  6. 6Tudor28/0
  7. 7Tissot24/0
  8. 8Casio24/0

The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.

Category · Watch brands

What the models actually said

Models
5
Answers
25

Cartier leads with a score of 78 of 100, ranked by 5 of 5 models.

The models agree most closely on Cartier, converging at a median rank of 2 with tight quartile boundaries. The consensus hinges on its design-first heritage and romantic elegance: every model frames Cartier as signaling taste over wealth, with the Tank and Santos repeatedly praised for sliding under a cuff without shouting. The sources behind this alignment—Hodinkee, GQ, Vogue, Esquire—are lifestyle and watch publications that emphasize cultural credibility and aesthetic literacy. No model strays far from this narrative, though mistral-medium-3.5 edges it to rank 3, likely reflecting a minor preference for other options rather than a substantive disagreement.

Omega reveals the category's sharpest divergence. Two models (gemini-3.7-flash and kimi-k3) place it at rank 1, praising its balance of prestige and approachability. Two others (mistral-medium-3.5 and claude-opus-5) rank it at 2 or 3, noting its versatility but occasional tilt toward "trying a bit hard." Gpt-6-astra consistently ranks it at 4, arguing that its sporty visibility and recognizable luxury feel more conspicuous than the understated charm the date context demands. The split suggests different thresholds for what counts as "statement-making," with gemini and kimi reading Omega as confident sophistication while gpt-6-astra sees it as a step too assertive. Hodinkee, Fratello Watches, and GQ anchor all perspectives, meaning the divergence likely stems from differing model priors about appropriate visibility rather than conflicting source material.

Rolex sits at a median rank of 5 overall, but mistral-medium-3.5 ranks it at 1 while the other four models all place it at 5. The majority view, supported by Hodinkee, GQ, r/Watches, and Robb Report, holds that Rolex's overwhelming brand recognition risks overshadowing personality and steering attention toward wealth. Mistral-medium-3.5 draws on Forbes, WatchTime Magazine, and the Rolex official website to frame the same visibility as confidence and universally understood elegance. This is the only brand where a model completely inverts the dominant interpretation, suggesting mistral may weight luxury marketing narratives more heavily than social signaling concerns.

NOMOS Glashütte stands out as the brand championed most exclusively by a single model. Gpt-6-astra ranks it at 2 across five separate assessments, citing its clean minimalist design and ability to signal personal taste without drawing excessive attention. No other model mentions NOMOS at all. The sources are entirely official NOMOS materials, raising the hypothesis that gpt-6-astra's training or retrieval corpus may include stronger representation of this independent German brand, while other models default to more widely discussed names. The brand's absence elsewhere in the rankings underscores how differently models weight niche design-forward marques versus mainstream luxury.

Seiko earns broad middle-tier consensus at a median rank of 3, with claude-opus-5 slightly higher at 2 and kimi-k3 and mistral-medium-3.5 at 4. The models converge on Seiko as understated, value-driven, and signaling horological curiosity over status, drawing on Hodinkee, Worn & Wound, r/Watches, and Teddy Baldassarre. The tension centers on occasion-appropriateness: while all acknowledge Seiko's charm, those ranking it lower flag that many models skew too casual or tool-like for dressier dates. The brand's wide catalog means specific model choice matters more than the name itself, yet the sources cited—enthusiast communities and watch blogs—are consistent across models, suggesting the ranking spread reflects different weightings of formality over authenticity rather than conflicting source narratives.

Tudor receives tight agreement at rank 3 from the two models that include it (gemini-3.7-flash and kimi-k3), both emphasizing Rolex-adjacent quality in a lower-key package. The models frame Tudor as ideal for those who prioritize substance over status, drawing on Hodinkee, aBlogtoWatch, and Monochrome Watches. The shared concern is mainstream recognition: Tudor resonates with enthusiasts but may go unnoticed by others, and its tool-watch aesthetic can feel utilitarian for formal dates. The narrow spread suggests strong alignment on Tudor's positioning, with the main variation being whether a given scenario skews casual enough to favor its grounded confidence.

Tissot occupies a median rank of 3.5, with claude-opus-5 at 3 and gemini-3.7-flash at 5. Both models praise the PRX as polished and versatile, supported by Fratello Watches, Worn & Wound, GQ, and Hodinkee. The disagreement centers on distinctiveness: claude views Tissot as a "low-risk, high-polish option," while gemini faults it for lacking "the distinct character or conversational intrigue of higher-ranked heritage brands." The PRX's ubiquity is flagged by both—claude notes it "says less about you personally" due to widespread popularity, while gemini sees it as "more like a sensible default." This suggests models diverge on whether safety or individuality matters more for first dates, even when drawing on overlapping sources.

Casio lands uniformly at rank 5 from both models that assess it, reflecting agreement that its utilitarian aesthetic undercuts sophistication. Kimi-k3 and mistral-medium-3.5 draw on TechRadar, the Casio Official Website, and Gear Patrol to acknowledge reliability and occasional conversational charm (vintage digital models), but both conclude that typical resin designs read as too casual. The consensus is tight, with the only variation being kimi's acknowledgment that specific lines like the Edifice might work for relaxed settings. The shared sources and reasoning suggest this is one of the category's clearest points of agreement.

The remaining brands—Apple, Hublot, Timex—appear only once each, all at rank 5. Apple is flagged by kimi-k3 as too distracting and insufficiently formal, supported by Apple's own site and The Verge. Hublot, assessed by gemini-3.7-flash, is deemed too bold and polarizing, with Time and Tide Watches and Watchuseek cited as community perspectives on its loudness. Timex, ranked by gpt-6-astra, is praised as unpretentious but noted as less distinctive than preferred alternatives, with Timex's official site as the sole source. These single-model assessments reflect idiosyncratic inclusion rather than category-wide themes.

The most frequently recalled sources—Hodinkee (53 citations), GQ (17), Fratello Watches (15), Worn & Wound (12), and r/Watches (11)—dominate across models, anchoring the consensus that first-date watches should balance quality with restraint. The hypothesis that these outlets shape model views more than official brand sites is supported by the fact that even when models cite manufacturer websites, they pair them with editorial perspectives that emphasize social appropriateness over technical specs. The exception is gpt-6-astra's exclusive reliance on NOMOS's official materials, suggesting that where a brand lacks strong third-party coverage in a model's training data, official sources may fill the gap and produce outlier rankings.

Category · Watch brands

What shaped the answers

55 sources across 313 references, grouped by site from 99 recalled names. The top 5 carry 44% of them.

Hodinkee×56 · 18%

hodinkee.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Watch brands
#1

Cartier

78
brand score
55first choice
12345

Summary

Cartier achieves a median rank of 2 across all models, with strong consensus around that position—the first and third quartiles sit at 2 and 3, respectively. Four of the five models place it at rank 2, while one (mistral-medium-3.5) ranks it at 3. The models converge on a consistent narrative: Cartier signals design literacy and romantic elegance rather than technical prowess or wealth display. Its iconic models—particularly the Tank and Santos—are repeatedly framed as jewelry-adjacent styling choices that read as thoughtful and refined. The brand is seen as recognizable enough to non-enthusiasts that it can start conversations without feeling like ostentation, though several models note its luxury associations and dressier character may skew slightly formal for very casual first encounters.

The sources supporting these themes are predominantly watch and lifestyle publications including Hodinkee, GQ, Vogue, Esquire, and A Blog to Watch, along with direct references to Cartier's own website. These outlets anchor the perception of Cartier as a design-first heritage brand with cultural credibility extending beyond collector circles. The emphasis on romantic Parisian elegance, timeless style, and aesthetic over specification is consistent across models, with only minor variation in how strongly each flags the potential formality trade-off.

A Tank or Santos reads as taste rather than money — slim, elegant, and it slides under a shirt cuff without shouting.

claude-opus-5

Cartier offers timeless elegance and an effortlessly romantic design aesthetic suited for a date setting.

gemini-3.7-flash

Per model summary

  • claude-opus-5
    brand score
    76
    first choice
    62
    median rank
    2.0
    present in
    5/5

    Consistently emphasizes that Cartier reads as taste and design rather than wealth-flexing, noting it's elegant, recognizable to non-enthusiasts, and slides discreetly under a cuff while carrying design credibility from cultural icons.

  • gemini-3.7-flash
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    5/5

    Repeatedly highlights Cartier's romantic elegance, artistic design pedigree, and timeless sophistication that signals refined personal style and cultural appreciation rather than technical specifications or wealth display.

  • gpt-6-astra
    brand score
    88
    first choice
    70
    median rank
    2.0
    present in
    5/5

    Focuses on Cartier's distinctive, jewelry-like design aesthetic as thoughtful styling rather than technical showmanship, though notes the luxury association can feel formal or conspicuous for casual settings.

  • kimi-k3
    brand score
    84
    first choice
    60
    median rank
    2.0
    present in
    5/5

    Emphasizes Cartier's romantic, Parisian elegance and design-first appeal that reads as refined and thoughtful rather than status-seeking, though consistently notes its dressier character may be slightly formal for very casual dates.

  • mistral-medium-3.5
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    5/5

    Repeatedly describes Cartier as refined, understated luxury with sleek elegance, highlighting its sophistication and heritage in fine jewelry and watchmaking as ideal for classy, upscale dates.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-57 sites · 15 references
Hodinkee×5 · 33%

hodinkee.com

gemini-3.7-flash7 sites · 14 references
Esquire×4 · 29%

esquire.com

gpt-6-astra1 site · 5 references
Cartier×5 · 100%

cartier.com

kimi-k35 sites · 14 references
GQ×5 · 36%

gq.com

mistral-medium-3.54 sites · 15 references
Cartier Official Website×5 · 33%

cartier.com

  • ×4named without a URL
  • ×1/

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Watch brands
#2

Omega

77
brand score
63first choice
  • claude-opus-5 3.0
12345

Summary

Omega lands consistently near the top of model recommendations, with a median rank of 2 overall. Two models rank it at the top (median rank 1), two place it slightly lower at rank 2 and 3, and one model consistently positions it at rank 4. The core theme uniting most perspectives is that Omega occupies a strategic middle ground: prestigious enough to signal taste and success, yet sufficiently understated to avoid appearing ostentatious or status-driven on a first date. Models emphasize its versatility across casual and formal settings, its rich heritage through associations with space exploration and James Bond, and its ability to serve as a natural conversation starter without demanding attention. Sources supporting this view cluster around enthusiast publications like Hodinkee, Fratello Watches, and GQ, alongside official Omega materials.

The primary point of divergence appears in how models weigh visibility against subtlety. While gemini-3.7-flash and kimi-k3 prize Omega's balance as ideal for communicating refined taste without trying too hard, gpt-6-astra consistently ranks it lower, citing its sporty aesthetic and visual prominence as more conspicuous than preferred for an understated first-date impression. Claude-opus-5 occupies the middle, acknowledging Omega's quality and story but occasionally noting it can register as "slightly more statement than charm" or risk tipping into "trying a bit hard" territory depending on context. The tension centers on whether Omega's recognizable luxury presence reads as confident sophistication or as a step beyond what some models consider optimally low-key.

Omega hits the sweet spot for a first date: genuinely respected by watch people yet not a status symbol that screams money, and models like the Aqua Terra or Speedmaster work with both a blazer and a t-shirt.

claude-opus-5

Omega strikes the ideal balance between prestige, timeless design, and understated sophistication. It communicates refined taste without appearing ostentatious or trying too hard on an initial meeting.

gemini-3.7-flash

Per model summary

  • gemini-3.7-flash
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Omega repeatedly represents an optimal balance between prestige and restraint, communicating refined taste, success, and sophistication without crossing into ostentation or appearing to try too hard.

  • kimi-k3
    brand score
    96
    first choice
    90
    median rank
    1.0
    present in
    5/5

    Omega reliably hits the sweet spot of signaling genuine taste and success without ostentation, with its Speedmaster and Seamaster heritage providing natural conversation starters that read as confident rather than flashy.

  • mistral-medium-3.5
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    5/5

    Omega consistently blends luxury with sportiness and approachability, while its associations with space exploration and James Bond add elements of intrigue, adventure, and charm that support both refinement and conversational interest.

  • claude-opus-5
    brand score
    64
    first choice
    47
    median rank
    3.0
    present in
    5/5

    Omega consistently occupies a strategic middle ground—respected by enthusiasts yet not overtly status-driven, versatile enough for varied settings, and offering natural conversation hooks through its Moonwatch and Bond heritage without demanding attention.

  • gpt-6-astra
    brand score
    44
    first choice
    27
    median rank
    4.0
    present in
    5/5

    Omega is recognized as offering good versatility and conversation potential through its technical heritage, but its sporty aesthetic and visual prominence consistently register as more conspicuous than this model's preference for understated first-date choices.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gemini-3.7-flash6 sites · 14 references
GQ×5 · 36%

gq.com

kimi-k36 sites · 15 references
Fratello Watches×5 · 33%

fratellowatches.com

mistral-medium-3.55 sites · 15 references
Hodinkee×5 · 33%

hodinkee.com

  • ×4/
  • ×1named without a URL
claude-opus-55 sites · 15 references
Fratello Watches×5 · 33%

fratellowatches.com

gpt-6-astra1 site · 5 references
Omega×5 · 100%

omegawatches.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Watch brands
#3

Seiko

54
brand score
37first choice
12345

Summary

Seiko lands at a median rank of 3 across models, with fairly strong consensus (Q1 3, Q3 4), though claude-opus-5 rates it slightly higher (median rank 2) while kimi-k3 and mistral-medium-3.5 place it at rank 4. The models converge on a common set of themes: Seiko signals understated confidence, genuine appreciation for craftsmanship over branding, and accessible quality that avoids pretense or status-seeking. It earns consistent praise for models like the Presage and Alpinist that deliver visual refinement without requiring luxury spending, making them safe, approachable choices that keep attention on the wearer rather than the watch. The brand draws support primarily from enthusiast sources including Hodinkee, Worn & Wound, r/Watches, Fratello Watches, and Teddy Baldassarre, alongside official Seiko materials.

The main tension in the rankings centers on versatility versus occasion-appropriateness. Models note that while Seiko offers solid value and thoughtful design, many pieces skew too casual or tool-like for dressier date settings, and the brand lacks the romantic recognition or elevated finish of Swiss alternatives. The broad range within Seiko's catalog also means the specific model chosen matters more than the brand name itself. Still, the consensus holds that Seiko communicates substance, humility, and horological curiosity—traits the models view as charming rather than flashy for a first-date context.

A Seiko (think Presage, Alpinist, or a slim SUR-series dress piece) looks genuinely good without announcing a price tag, which is exactly the tone you want on a first date.

claude-opus-5

Seiko is the enthusiast's choice — wearing one suggests you care about craftsmanship and value rather than logos, which can come across as genuinely charming.

kimi-k3

Per model summary

  • claude-opus-5
    brand score
    80
    first choice
    63
    median rank
    2.0
    present in
    5/5

    Seiko delivers understated charm and horological credibility without signaling wealth or status, keeping attention on the wearer rather than the watch. Models like the Presage and Alpinist punch above their price point while avoiding the risk of seeming pretentious.

  • gemini-3.7-flash
    brand score
    52
    first choice
    30
    median rank
    3.0
    present in
    5/5

    Seiko represents grounded confidence and appreciation for craftsmanship without relying on luxury branding or pretense. It communicates practical elegance and authenticity, though it lacks the elevated polish that creates a special-occasion feeling.

  • gpt-6-astra
    brand score
    64
    first choice
    45
    median rank
    3.0
    present in
    5/5

    Seiko provides attractive, understated options that work for relaxed dates without requiring luxury spending, though its broad range makes the specific model choice more critical than the brand name. It allows intentional style on an accessible budget.

  • kimi-k3
    brand score
    36
    first choice
    22
    median rank
    4.0
    present in
    4/5

    Seiko signals genuine enthusiasm for craftsmanship and value over logos, which reads as charming and unpretentious to those who appreciate substance. It ranks lower because many models skew too casual or tool-like for dressier date settings and it lacks the romantic recognition of Swiss brands.

  • mistral-medium-3.5
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    5/5

    Seiko balances quality craftsmanship, reliability, and affordability while offering versatile styling, particularly in lines like Presage and Cocktail Time. Though not prestigious, it demonstrates practical good taste and substance over flash.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-56 sites · 15 references
r/Watches×5 · 33%

reddit.com

gemini-3.7-flash8 sites · 13 references
Worn & Wound×4 · 31%

wornandwound.com

gpt-6-astra1 site · 5 references
Seiko×5 · 100%

seikowatches.com

kimi-k35 sites · 11 references
Hodinkee×4 · 36%

hodinkee.com

mistral-medium-3.56 sites · 15 references
Seiko Official Website×5 · 33%

seikowatches.com

  • ×4named without a URL
  • ×1/

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Watch brands
#4

Rolex

38
brand score
36first choice
12345

Summary

Rolex sits at a median rank of 5 across the models, revealing a sharp divide in perception. One model consistently places it first, drawing on sources including Forbes, WatchTime Magazine, and the Rolex official website to argue that the brand signals timeless elegance, confidence, and universally recognized quality. The remaining four models all assign it a median rank of 5, citing sources such as Hodinkee, GQ, r/Watches, and Robb Report to emphasize that while Rolex offers exceptional craftsmanship, its overwhelming brand recognition risks overshadowing personality and steering attention toward wealth rather than genuine connection. The disagreement centers on whether iconic status is an asset or a liability in first-date contexts.

The majority theme across models is that Rolex's visibility as a luxury symbol creates conversational friction before rapport is established. Three models warn it may read as conspicuous consumption or status-flexing, and two note that its price is immediately recognizable to most observers, inviting assumptions about the wearer's priorities. Even the dissenting model acknowledges the brand's strong luxury association, though it frames this as confidence rather than distraction. The split reflects tension between Rolex's undisputed prestige and the social risk of leading with a watch that, as multiple models note, speaks before the wearer does.

Rolex is superbly made, but it's the one brand almost everyone can price at a glance, so it tends to speak before you do.

claude-opus-5

While Rolex is the most universally recognized luxury watchmaker, its overt prestige can inadvertently make you seem status-focused or ostentatious on a first outing.

gemini-3.7-flash

Per model summary

  • mistral-medium-3.5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Rolex represents timeless elegance and sophistication that signals confidence, success, and quality while making a strong impression that is universally recognized and respected.

  • claude-opus-5
    brand score
    20
    first choice
    20
    median rank
    5.0
    present in
    5/5

    Rolex is excellent in quality but too recognizable as a wealth signal on first dates, risking distraction from genuine connection by broadcasting status before conversation begins.

  • gemini-3.7-flash
    brand score
    28
    first choice
    22
    median rank
    5.0
    present in
    5/5

    Despite exceptional craftsmanship, Rolex's overwhelming brand recognition risks projecting pretension or status-seeking that can overshadow personality and distract from authentic first-date conversation.

  • gpt-6-astra
    brand score
    16
    first choice
    16
    median rank
    5.0
    present in
    4/5

    Rolex offers quality and classic design, but its strong luxury association draws unwanted attention to price and status rather than keeping focus on the conversation during a first meeting.

  • kimi-k3
    brand score
    28
    first choice
    22
    median rank
    5.0
    present in
    5/5

    Rolex is iconic and well-made, but on first dates it carries high risk of appearing as conspicuous consumption or wealth-flexing that overshadows personality before conversation starts.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

mistral-medium-3.54 sites · 15 references
Forbes×5 · 33%

forbes.com

claude-opus-57 sites · 15 references
r/Watches×5 · 33%

reddit.com

gemini-3.7-flash6 sites · 13 references
Hodinkee×4 · 31%

hodinkee.com

gpt-6-astra1 site · 4 references
Rolex×4 · 100%

rolex.com

kimi-k39 sites · 15 references
GQ×3 · 20%

gq.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Watch brands
#5

NOMOS Glashütte

17
brand score
13first choice
12345

Summary

NOMOS Glashütte receives consistently strong support from the model that ranked it, landing at a median rank of 2 across five separate assessments. The quartile range from 1 to 2 shows remarkable agreement, with the brand sometimes claiming the top position and otherwise placing second. This narrow spread reflects a coherent perception of NOMOS as a first-date choice that balances aesthetic distinction with restraint. The model draws on five separate references to the official NOMOS Glashütte website to support its evaluations.

The recurring themes center on the brand's clean, minimalist design language and its ability to signal personal taste without drawing excessive attention. The model values NOMOS for typography, understated details, and occasional playful color touches that convey thoughtfulness and design awareness rather than status signaling. When NOMOS ranks second rather than first, the reasoning typically involves comparison to more romantic or versatile options, with the model noting that the minimalist aesthetic may feel slightly more casual or design-focused than alternatives, and potentially less suited to rugged or ornate personal styles.

My first choice would be NOMOS: its clean, understated designs suggest personal taste without making the watch the center of attention. That balance feels especially right for a first date, whether the setting is coffee or dinner.

gpt-6-astra

NOMOS would be my choice for understated style: clean typography and restrained designs suggest attention to detail rather than an obvious status display.

gpt-6-astra

Per model summary

  • gpt-6-astra
    brand score
    84
    first choice
    67
    median rank
    2.0
    present in
    5/5

    This model consistently emphasizes NOMOS's clean, understated design and restrained aesthetic that suggests thoughtfulness and personal taste without demanding attention. It values the brand's ability to balance distinctive style with a quiet, minimalist character appropriate for first dates.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gpt-6-astra1 site · 5 references
NOMOS Glashütte×5 · 100%

nomos-glashuette.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Watch brands

How this was measured

The question

  • Unaided brand recommendation question: “Which watch brand would you wear on a first date?
  • The prompt asks each model to return exactly 5 brands, ranked 1 to 5, and for each one a reason for the recommendation and the sources that informed it.

Sampling

  • Models are not deterministic, so the question is asked over and over — 5 answers per model
  • Spellings of the same brand are normalized to the most commonly used form before counting

Measures of Position: Median and Quartiles

  • Median value indicates that in 50% of answers the brand held this position or higher.
  • Quartiles help visualize the spread of rankings per brand. Q1 indicates that 25% of answers had this rank or higher. Q3 means that 75% of answers ranked the brand as X or better.
  • Medians and quartiles are computed only from answers where the brand was present.

Coverage

  • Count of answers in which the given brand was named, as a percentage.

Brand score - Normalized Borda score

  • The brand score is calculated using a normalized Borda score. Borda count is a voting method: each ballot awards points by position instead of naming one winner. Our ranking responses from the LLM are always fixed to 5 answers, so we assign 100 points for rank 1, then 80, 60, 40, 20 — and 0 if the brand is missing.
Brand score=100ni=1n6ri5\text{Brand\ score} = \frac{100}{n} \sum_{i=1}^{n} \frac{6 - r_i}{5}

where

ri={1,,5if the brand appears6if the brand is not mentionedr_i = \begin{cases} 1,\dots,5 & \text{if the brand appears} \\ 6 & \text{if the brand is not mentioned} \end{cases}
  • This way we can calculate a common score for all brands mentioned across model responses. The score reflects both how high the brand was ranked and how frequently it was mentioned.

First choice score - adjusted MRR - top-rank indicator

  • The first choice score is based on Mean Reciprocal Rank (a common search-engine measure), but adapted to measure the position of a specific brand across repeated ranked responses.
  • Each occurrence receives a reciprocal-position score: 100 for 1st, 50 for 2nd, 33.3 for 3rd, 25 for 4th, 20 for 5th, and 0 when the brand is not mentioned. The scores are averaged across all responses.
First choice score=100ni=1nsi\text{First\ choice\ score} = \frac{100}{n} \sum_{i=1}^{n} s_i

where

si={1/riif the brand is present0if the brand is absents_i = \begin{cases} 1/r_i & \text{if the brand is present} \\ 0 & \text{if the brand is absent} \end{cases}
  • No rank below 1st place gets more than 50, so the score is heavily driven by first places.
Category · Neobanks uk

What neobank would you recommend for a UK startup?

6 AI models · 4 answers per model · ranked by score

ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0

  1. #1Starling Bank
    90
    brand score
    76first choice
    12345
  2. #2Revolut
    69
    brand score
    49first choice
  3. #3Tide
    65
    brand score
    49first choice
  4. #4Monzo
    52
    brand score
    33first choice
  5. #5Wise
    17
    brand score
    14first choice
Show 3 more
  1. #6Anna Money
    5
    brand score
    5first choice
  2. #7Mercury
    2
    brand score
    1first choice
  3. #8Mettle
    1
    brand score
    1first choice

Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.

Swipe left for the full report

Category · Neobanks uk

Model by model

Median rank each model gave each brand.

Scroll the table sideways to see every model →

claude-opus-5deepseek-v4.1-flashgemini-3.7-flashgpt-6-astrakimi-k3mistral-medium-3.5
Starling Bank2.01.01.01.02.02.0
Revolut3.03.02.53.53.01.0
Tide1.04.03.53.51.53.0
Monzo4.52.03.52.04.04.0
Wise5.05.05.05.04.0
Anna Money5.05.05.0
Mercury4.0
Mettle5.0
1st
2
3
4
5th

A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.

Category · Neobanks uk

Where the models disagree

2 of 8 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.

  1. #3Tide

    deepseek-v4.1-flash 45 claude-opus-5 100

  2. #4Monzo

    claude-opus-5 30 gpt-6-astra 80

  3. #2Revolut

    gpt-6-astra 50 mistral-medium-3.5 100

  4. #1Starling Bank

    claude-opus-5 75 gpt-6-astra 100

  5. #5Wise

    mistral-medium-3.5 0 kimi-k3 20 · 1 of 6 never named it

  6. #6Anna Money

    claude-opus-5 0 mistral-medium-3.5 20 · 3 of 6 never named it

A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.

Category · Neobanks uk

Named, and named first

Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.

the default answera narrow favouritelisted, rarely ledthe long tail12345678
  1. 1Starling Bank100/54
  2. 2Revolut100/21
  3. 3Tide100/25
  4. 4Monzo100/0
  5. 5Wise67/0
  6. 6Anna Money25/0
  7. 7Mercury4/0
  8. 8Mettle4/0

The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.

Category · Neobanks uk

What the models actually said

Models
6
Answers
24

Starling Bank leads with a score of 90 of 100, ranked by 6 of 6 models.

Starling Bank is the clearest point of consensus in this category, ranked by all six models and placed first by three of them. The split is narrow and structural: deepseek-v4.1-flash, gemini-3.7-flash and gpt-6-astra rank it first every time, while claude-opus-5, kimi-k3 and mistral-medium-3.5 hold it at rank 2 on the same reasoning — a full UK licence and FSCS protection, offset by weaker multi-currency and FX support. The disagreement is about ceiling, not floor.

Starling Bank

per-model scores 75100 of 100 · mean 90 across 6 models

Revolut and Tide are where the models diverge most sharply, and both divides turn on the same fault line: whether an internationally strong but non-fully-licensed provider belongs at the top. mistral-medium-3.5 ranks Revolut first in every run, while gpt-6-astra sits it lowest at a median rank of 3.5, docking it for plan fees on domestic use. Tide shows the widest spread of any brand (standard deviation 19.15), with claude-opus-5 placing it first and deepseek-v4.1-flash placing it fourth or lower — the same ClearBank e-money structure read as reassurance by one and as a gap by the other.

claude-opus-5
100
rates it highest
deepseek-v4.1-flash
45
rates it lowest

Tide · 55 points apart on a 0–100 scale

Monzo produces the most model-specific reversal. gpt-6-astra treats it as a close second (borda 80) suited to founder-led startups, while claude-opus-5 ranks it near-last (borda 30), treating its thinner multi-currency and integration set as disqualifying for scaling teams. No model ranked it first. This is the clearest case where two models applied the same facts to opposite conclusions about who Monzo is for.

MonzoRevolut
  • 80gpt-6-astra50
  • 70deepseek-v4.1-flash60
  • 55gemini-3.7-flash70
  • 30claude-opus-565
  • 35kimi-k370
  • 40mistral-medium-3.5100

Monzo ahead on 2 of 6 models

Where the panel converges is at the bottom. Wise draws unusual agreement (spread 0–20): five models rank it, none place it first, and all frame it as a treasury complement rather than a primary account, citing its lack of FSCS protection. Anna Money, Mercury and Mettle barely register, each pulled up only by isolated models — Mercury solely by claude-opus-5 on US-incorporation grounds, Mettle solely by kimi-k3.

The source patterns are broadly shared rather than distinguishing. Trustpilot×10 and brand-owned pages appear behind nearly every model's reasoning, which makes it hard to attribute the ranking splits to differing evidence. A plausible but unverified reading is that the licence-versus-features divide reflects how each model weighed FSCS status against international capability, not which sources it happened to recall — the recalled sources overlap too heavily to explain the gaps on their own.

The models agree on the shortlist; they disagree on how much a full banking licence should outrank international features.
Category · Neobanks uk

What shaped the answers

43 sources across 313 references, grouped by site from 141 recalled names. The top 5 carry 43% of them.

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Neobanks uk
#1

Starling Bank

also named Starling

90
brand score
76first choice
12345

Summary

Starling Bank is one of the strongest performers in this category, scoring

Stat unavailable

and ranked by all six models. The agreement is broad: per-model borda spans 75–100, and the split runs cleanly between two camps. Three models — deepseek-v4.1-flash, gemini-3.7-flash and gpt-6-astra — placed it first in every run, drawn to a consistent set of themes: a full UK banking licence with FSCS protection, a fee-free business account, accounting integrations (Xero, QuickBooks) and strong UK customer support. These points are anchored in recalled sources such as Trustpilot×10, Which?×6 and the Starling Bank business account page.

Starling Bank

per-model scores 75100 of 100 · mean 90 across 6 models

The remaining three — claude-opus-5, kimi-k3 and mistral-medium-3.5 — held it at rank 2, applying the same strengths but noting weaker multi-currency and international/FX support and less startup-specific tooling relative to rivals such as Tide, Revolut and Wise. That reservation, rather than any doubt about safety or day-to-day fitness, is what separates a first-place vote from a second.

Starling is a fully licensed UK bank with FSCS protection, free everyday business banking, and consistently top-rated customer service in the CMA service quality surveys.

claude-opus-5

It ranks second rather than first because its multi-currency and international payment capabilities are weaker than Revolut's or Wise's.

kimi-k3

Per model summary

  • deepseek-v4.1-flash
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    4/4

    Highlights Starling as a fully licensed, FSCS-protected UK bank with a free business account, strong app, and accounting integrations (Xero, QuickBooks), while noting weaker international/FX support.

  • gemini-3.7-flash
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    4/4

    Repeatedly stresses full UK regulatory protection (FSCS up to £85,000), zero monthly fees, accounting software integrations, and award-winning UK customer support as making it the top choice for startups.

  • gpt-6-astra
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    4/4

    Frames Starling as the first choice for pound-focused UK startups given its no-fee business account, accounting integrations, and FSCS protection, while suggesting international-heavy businesses may need a specialist provider.

  • claude-opus-5
    brand score
    75
    first choice
    46
    median rank
    2.0
    present in
    4/4

    Emphasizes Starling's full UK banking licence with FSCS protection and free, top-rated business account, while consistently noting its weaker multi-currency support and less startup-specific/VC-oriented features versus rivals.

  • kimi-k3
    brand score
    85
    first choice
    63
    median rank
    2.0
    present in
    4/4

    Points to Starling as a fully licensed UK bank with FSCS protection and a free, well-reviewed business account, but ranks it just behind rivals due to less startup-specific tooling and weaker international/multi-currency capabilities.

  • mistral-medium-3.5
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    4/4

    Consistently describes Starling as an FCA-regulated, fee-free business account with fast payments, strong API/integrations, and an emphasis on transparency and customer support for startups.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

deepseek-v4.1-flash5 sites · 11 references
Reddit r/UKPersonalFinance×3 · 27%

reddit.com

gemini-3.7-flash4 sites · 12 references
Trustpilot×4 · 33%

trustpilot.com

gpt-6-astra1 site · 4 references
Starling Bank — Business account×4 · 100%

starlingbank.com

claude-opus-55 sites · 12 references
Starling Bank×6 · 50%

starlingbank.com

kimi-k37 sites · 12 references
Starling Bank×4 · 33%

starlingbank.com

mistral-medium-3.54 sites · 12 references
Starling Bank×4 · 33%

starlingbank.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Neobanks uk
#2

Revolut

also named Revolut Business

69
brand score
49first choice
12345

Summary

Revolut scores a Borda average of

Stat unavailable

and is ranked by all 6 of the 6 models, placing it firmly among the recommended options. The report labels the per-model spread of 50–100 "broad agreement," though the models do diverge on how high to place it: mistral-medium-3.5 ranked it first in all 4 of its runs (borda 100), while gpt-6-astra sat it lowest with a median rank of 3.5 (borda 50).

mistral-medium-3.5
100
rates it highest
gpt-6-astra
50
rates it lowest

Revolut · 50 points apart on a 0–100 scale

The theme uniting the models is Revolut's strength for internationally oriented startups — multi-currency accounts, competitive FX, team cards and API-driven expense management recur across every reasoning. The reservations are equally consistent: several models flag the absence of full FSCS deposit protection, the still-maturing UK banking licence in its mobilisation phase, and recurring reports of account freezes and slow support. gpt-6-astra adds that plan fees and usage allowances make it less compelling for a mainly domestic startup. These themes lean on recalled sources including Trustpilot×10 for support and reliability signals, Siftedsifted.eu×7 and TechCrunch×3 for sector context, with Financial Times coverage and the FCA Register cited on the licence question.

Revolut is the strongest option for startups with international operations, offering multi-currency accounts, near-interbank FX, bulk payments, expense management and a solid API.

claude-opus-5

For a mainly domestic startup, plan fees and usage allowances make it less compelling than the higher-ranked options; check the UK account's current legal provider and deposit-protection terms before holding substantial cash

gpt-6-astra

Per model summary

  • mistral-medium-3.5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    4/4

    Revolut is presented as a top choice for UK startups, emphasizing multi-currency support, competitive FX, expense management, and accounting-tool integrations suited to scaling businesses.

  • gemini-3.7-flash
    brand score
    70
    first choice
    42
    median rank
    2.5
    present in
    4/4

    Revolut excels for startups with international/cross-border operations via multi-currency accounts and competitive FX, but is held back by the lack of full UK FSCS deposit protection and support/account-freeze issues.

  • claude-opus-5
    brand score
    65
    first choice
    38
    median rank
    3.0
    present in
    4/4

    Revolut is positioned as the strongest option for internationally operating startups (multi-currency, interbank FX, strong API), with its UK banking licence in mobilisation improving credibility, but weighed down by expensive tiers, weak support, and account freezes.

  • deepseek-v4.1-flash
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    4/4

    Revolut is favored for multi-currency accounts and international payments, but repeatedly flagged for lacking FSCS protection (e-money status) and inconsistent customer support.

  • kimi-k3
    brand score
    70
    first choice
    50
    median rank
    3.0
    present in
    4/4

    Revolut is the strongest choice for internationally ambitious startups (multi-currency, FX, team cards, API), but its recently granted, restricted UK banking licence in mobilisation plus account freezes and slow support add operational risk.

  • gpt-6-astra
    brand score
    50
    first choice
    29
    median rank
    3.5
    present in
    4/4

    Revolut is attractive for startups paying overseas suppliers or managing multiple currencies, but less compelling for domestic-focused startups due to plan fees and allowances, with repeated advice to verify the legal entity and deposit protection.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

mistral-medium-3.54 sites · 12 references
Revolut Business×4 · 33%

revolut.com

gemini-3.7-flash6 sites · 12 references
Trustpilot×4 · 33%

trustpilot.com

claude-opus-55 sites · 12 references
Revolut Business×4 · 33%

revolut.com

deepseek-v4.1-flash6 sites · 11 references
Revolut Business×4 · 36%

revolut.com

kimi-k37 sites · 12 references
Revolut×4 · 33%

revolut.com

gpt-6-astra1 site · 4 references
Revolut — Business×4 · 100%

revolut.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Neobanks uk
#3

Tide

65
brand score
49first choice
12345

Summary

Tide scores 65 of 100 on Borda and is ranked by all 6 of the 6 models, but the recommendation is far from unanimous. Per-model scores span 45 to 100 with a standard deviation of 19.15, which the report labels "models split." That gap traces a clean divide between claude-opus-5, which ranked Tide first in all four of its runs (borda 100), and the cluster of gpt-6-astra (50) and deepseek-v4.1-flash (45), which consistently placed it fourth or lower.

claude-opus-5
100
rates it highest
deepseek-v4.1-flash
45
rates it lowest

Tide · 55 points apart on a 0–100 scale

The theme uniting every model is that Tide is purpose-built for UK small businesses and startups — fast onboarding, invoicing, expense management, company formation and accounting integrations recur across all six. Where they diverge is on how much its e-money status weighs against it: claude-opus-5 and kimi-k3 treat the ClearBank partnership and FSCS protection as reassurance, while deepseek-v4.1-flash reads the same structure as a lack of protection and gpt-6-astra and gemini-3.7-flash dock it for transaction fees and paid extras as activity scales. Recalled sources concentrate on Trustpilotuk.trustpilot.com/review/tide.co×6 and Tide's own site, with comparison write-ups from Startups.co.uk×4 and MoneySavingExpert supporting the feature and pricing claims.

Tide is the model's default for early-stage UK admin, but its e-money status and per-transaction fees drive the split over whether it belongs as a primary account.

Tide is purpose-built for UK small businesses and startups, with free company incorporation bundles, fast onboarding, invoicing, expense cards and accounting integrations.

claude-opus-5

It sits at rank four because per-transaction fees and a leaner feature set make it less attractive once a startup begins to scale.

kimi-k3

Per model summary

  • claude-opus-5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    4/4

    Tide is consistently described as purpose-built for UK small businesses and startups, with free company incorporation, fast onboarding, invoicing/expense tools, accounting integrations, and the largest UK SME challenger market share. The recurring caveat is that it is an e-money account rather than a full bank (FSCS protection via ClearBank).

  • kimi-k3
    brand score
    80
    first choice
    69
    median rank
    1.5
    present in
    4/4

    Tide is repeatedly described as purpose-built for UK SMEs and startups with fast onboarding, invoicing, expense management and accounting integrations, noting funds held via ClearBank give FSCS protection, though it is an e-money platform and fees limit it as a startup scales.

  • mistral-medium-3.5
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    4/4

    Tide is consistently characterized as tailored for small businesses and startups with quick setup and invoicing tools, valued for simplicity and affordability but noted as lacking some advanced or scalable features of competitors.

  • gemini-3.7-flash
    brand score
    55
    first choice
    33
    median rank
    3.5
    present in
    4/4

    Tide is consistently framed as tailored specifically for UK small businesses and sole traders with rapid onboarding and built-in bookkeeping/invoicing tools, but ranked lower due to transaction fees and its status as an e-money institution relying on partner banks.

  • gpt-6-astra
    brand score
    50
    first choice
    29
    median rank
    3.5
    present in
    4/4

    Tide is consistently positioned as strong for business administration (invoicing, bookkeeping, company formation) but ranked below Starling and Monzo because of transaction charges and paid extras, with emphasis that it is a financial platform rather than a bank itself.

  • deepseek-v4.1-flash
    brand score
    45
    first choice
    30
    median rank
    4.0
    present in
    4/4

    Tide is portrayed as a fast, low-cost account popular with small businesses offering invoicing and expense features, but repeatedly flagged as an e-money institution without FSCS protection, best suited to very early-stage needs rather than a scaling primary account.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-56 sites · 12 references
Tide×4 · 33%

tide.co

kimi-k36 sites · 12 references
Tide×4 · 33%

tide.co

mistral-medium-3.54 sites · 12 references
MoneySavingExpert×4 · 33%

no URL recalled

  • ×4named without a URL
gemini-3.7-flash6 sites · 12 references
Trustpilot×4 · 33%

trustpilot.com

gpt-6-astra1 site · 4 references
Tide×4 · 100%

tide.co

deepseek-v4.1-flash6 sites · 11 references
Tide×4 · 36%

tide.co

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Neobanks uk
#4

Monzo

also named Monzo Business

52
brand score
33first choice
12345

Summary

Monzo lands mid-table with a Borda score of 51.67 of 100, ranked by all 6 of the 6 models but with notable disagreement on how highly to place it. The per-model span runs from 30 to 80 with a standard deviation of 18.41, which the report labels "models split."

Monzo

per-model scores 3080 of 100 · mean 52 across 6 models

The divide tracks a consistent theme: every model credits Monzo as a fully licensed, FSCS-protected UK bank with a polished app and useful touches like tax pots and invoicing, but they weigh its narrower business feature set differently. gpt-6-astra positions it as a close second to Starling and scored it 80, while claude-opus-5 sees the thinner multi-currency, FX and integration offering as disqualifying for scaling startups and scored it 30. Between them sit deepseek-v4.1-flash at 70 and the more skeptical kimi-k3 (35) and mistral-medium-3.5 (40), the latter pair emphasising that Monzo suits solo founders and micro-startups more than growing teams. No model ranked it first in any run. The recalled sources cluster around review aggregators and Monzo's own pages, with Trustpilotuk.trustpilot.com/review/monzo.com×5 and Which?www.which.co.uk×1 among those cited under these themes.

Monzo is a close second for a small, founder-led startup because of its straightforward app, spending visibility and useful business administration features.

gpt-6-astra

It's ranked last here because its feature set is thinner for scaling startups — weaker multi-currency/FX, fewer accounting and API integrations, and the Pro tier costs extra for capabilities rivals include free.

claude-opus-5

Per model summary

  • deepseek-v4.1-flash
    brand score
    70
    first choice
    44
    median rank
    2.0
    present in
    4/4

    Monzo is portrayed as a fully licensed, FSCS-protected bank with an intuitive app, tax pots and easy setup ideal for small founder-led teams, but with fewer features and integrations than Starling or Tide and paid tiers for fuller functionality.

  • gpt-6-astra
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    4/4

    Monzo is consistently positioned as a close second to Starling, praised for its intuitive app and spending visibility with FSCS protection, but with useful business features locked behind paid plans, prompting a cost comparison against Starling.

  • gemini-3.7-flash
    brand score
    55
    first choice
    33
    median rank
    3.5
    present in
    4/4

    Monzo is framed as a fully licensed, FSCS-protected UK bank with a polished, award-winning app and features like tax pots and invoicing, but ranked behind Starling because multi-user access and advanced tools sit behind paid subscription tiers.

  • kimi-k3
    brand score
    35
    first choice
    24
    median rank
    4.0
    present in
    4/4

    Monzo is described as a licensed UK bank with excellent UX and features like tax pots that suit solo founders and early-stage teams, but with a thinner, less mature business feature set than Tide or Starling that limits its fit for scaling or international startups.

  • mistral-medium-3.5
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    4/4

    Monzo is characterized as user-friendly and well-integrated, ideal for sole traders and micro-startups with basic needs, but lacking advanced business features like multi-currency support compared to Revolut or Starling.

  • claude-opus-5
    brand score
    30
    first choice
    23
    median rank
    4.5
    present in
    4/4

    Monzo is consistently described as a fully licensed, FSCS-protected UK bank with an excellent app and features like tax pots and invoicing, but with a narrower business feature set (limited multi-currency/FX, fewer integrations) that suits UK-only solo founders and small teams over scaling startups.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

deepseek-v4.1-flash5 sites · 11 references
Monzo Business×4 · 36%

monzo.com

gpt-6-astra1 site · 4 references
Monzo — Business banking×4 · 100%

monzo.com

gemini-3.7-flash6 sites · 12 references
Trustpilot×4 · 33%

trustpilot.com

kimi-k35 sites · 11 references
Monzo×4 · 36%

monzo.com

mistral-medium-3.54 sites · 12 references
Monzo×4 · 33%

monzo.com

claude-opus-55 sites · 12 references
Monzo×4 · 33%

monzo.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Neobanks uk
#5

Wise

also named Wise Business

17
brand score
14first choice
12345

Summary

Wise scores 16.67 of 100 on Borda, ranked by 5 of the 6 models but never placed first by any of them. The agreement here is unusually tight: per-model Borda spans 0–20 with a standard deviation of 7.45, which the report labels "models agree." Every model that ranked it landed on the same conclusion, differing only in how far down the list they placed it — kimi-k3 gave it a median rank of 4, while claude-opus-5, deepseek-v4.1-flash, gemini-3.7-flash and gpt-6-astra all settled on a median rank of 5.

Wise

per-model scores 020 of 100 · mean 17 across 6 models

The theme behind that consensus is consistent across all five: Wise is valued for low-cost international payments, mid-market FX rates and multi-currency holding, but is treated as a complement rather than a primary account because it is an e-money or payments institution lacking FSCS protection, credit and overdrafts. Recalled sources cluster around user-review and comparison material, with Trustpilot×10 appearing most often across models, alongside the Wise Businesswise.com/gb/business×5 official site. The distinction models draw is not about quality but role — a specialist treasury tool paired with a UK business bank rather than the bank itself.

Best used as a complement rather than the primary account.

kimi-k3

For most UK startups, it serves best as a secondary treasury tool alongside a primary UK clearing bank.

gemini-3.7-flash

Per model summary

  • kimi-k3
    brand score
    20
    first choice
    13
    median rank
    4.0
    present in
    2/4

    Wise is described as outstanding for low-cost international payments and multi-currency accounts, but ranked last as a primary bank because it is an e-money institution without FSCS protection or credit products, best used as a complement.

  • claude-opus-5
    brand score
    20
    first choice
    16
    median rank
    5.0
    present in
    3/4

    Wise is positioned as an excellent supplementary account for cross-border payments and multi-currency holding, but ranked lowest as a primary option because it is not a bank—lacking FSCS protection, credit, and overdrafts.

  • deepseek-v4.1-flash
    brand score
    20
    first choice
    16
    median rank
    5.0
    present in
    3/4

    Wise is presented as a strong companion account for international transfers and multi-currency holding, but not a full UK bank—no FSCS protection or lending—so it is best used as a secondary tool rather than a primary account.

  • gemini-3.7-flash
    brand score
    20
    first choice
    20
    median rank
    5.0
    present in
    4/4

    Wise is highlighted for mid-market FX rates and multi-currency international payments, but as an electronic money/payment institution rather than a full bank it lacks credit facilities and FSCS coverage, making it a complement to a primary domestic bank.

  • gpt-6-astra
    brand score
    20
    first choice
    20
    median rank
    5.0
    present in
    4/4

    Wise is valued for international revenue collection and overseas payments with transparent pricing, but ranked fifth as a primary account because it is a payments provider with safeguarded rather than FSCS-protected balances, best paired with a UK business bank.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

kimi-k34 sites · 6 references
Monito×2 · 33%

monito.com

claude-opus-54 sites · 9 references
Wise×5 · 56%

wise.com

deepseek-v4.1-flash5 sites · 9 references
Wise Business×3 · 33%

wise.com

gemini-3.7-flash6 sites · 12 references
Trustpilot×4 · 33%

trustpilot.com

gpt-6-astra1 site · 4 references
Wise — Business×4 · 100%

wise.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Neobanks uk

How this was measured

The question

  • Unaided brand recommendation question: “What neobank would you recommend for a UK startup?
  • The prompt asks each model to return exactly 5 brands, ranked 1 to 5, and for each one a reason for the recommendation and the sources that informed it.

Sampling

  • Models are not deterministic, so the question is asked over and over — 4 answers per model
  • Spellings of the same brand are normalized to the most commonly used form before counting

Measures of Position: Median and Quartiles

  • Median value indicates that in 50% of answers the brand held this position or higher.
  • Quartiles help visualize the spread of rankings per brand. Q1 indicates that 25% of answers had this rank or higher. Q3 means that 75% of answers ranked the brand as X or better.
  • Medians and quartiles are computed only from answers where the brand was present.

Coverage

  • Count of answers in which the given brand was named, as a percentage.

Brand score - Normalized Borda score

  • The brand score is calculated using a normalized Borda score. Borda count is a voting method: each ballot awards points by position instead of naming one winner. Our ranking responses from the LLM are always fixed to 5 answers, so we assign 100 points for rank 1, then 80, 60, 40, 20 — and 0 if the brand is missing.
Brand score=100ni=1n6ri5\text{Brand\ score} = \frac{100}{n} \sum_{i=1}^{n} \frac{6 - r_i}{5}

where

ri={1,,5if the brand appears6if the brand is not mentionedr_i = \begin{cases} 1,\dots,5 & \text{if the brand appears} \\ 6 & \text{if the brand is not mentioned} \end{cases}
  • This way we can calculate a common score for all brands mentioned across model responses. The score reflects both how high the brand was ranked and how frequently it was mentioned.

First choice score - adjusted MRR - top-rank indicator

  • The first choice score is based on Mean Reciprocal Rank (a common search-engine measure), but adapted to measure the position of a specific brand across repeated ranked responses.
  • Each occurrence receives a reciprocal-position score: 100 for 1st, 50 for 2nd, 33.3 for 3rd, 25 for 4th, 20 for 5th, and 0 when the brand is not mentioned. The scores are averaged across all responses.
First choice score=100ni=1nsi\text{First\ choice\ score} = \frac{100}{n} \sum_{i=1}^{n} s_i

where

si={1/riif the brand is present0if the brand is absents_i = \begin{cases} 1/r_i & \text{if the brand is present} \\ 0 & \text{if the brand is absent} \end{cases}
  • No rank below 1st place gets more than 50, so the score is heavily driven by first places.
Category · No Code Platforms

What are the best no-code platforms for a company to build internal workflows and apps without developers?

6 AI models · 10 answers per model · ranked by score

ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0

  1. #1Airtable
    90
    brand score
    81first choice
    12345
  2. #2Microsoft
    49
    brand score
    35first choice
    • gemini-3.7-flash 3.0
  3. #3Zapier
    39
    brand score
    24first choice
  4. #4Retool
    32
    brand score
    26first choice
  5. #5Glide
    28
    brand score
    18first choice
Show 11 more
  1. #6AppSheet
    19
    brand score
    12first choice
  2. #7Quickbase
    11
    brand score
    7first choice
  3. #8Make
    6
    brand score
    4first choice
  4. #9Softr
    6
    brand score
    4first choice
  5. #10Kissflow
    5
    brand score
    4first choice
  6. #11Bubble
    5
    brand score
    4first choice
  7. #12Appsmith
    3
    brand score
    3first choice
  8. #13Smartsheet
    2
    brand score
    2first choice
  9. #14Google
    2
    brand score
    1first choice
    • gpt-6-astra 3.0
  10. #15Monday.com
    2
    brand score
    1first choice
  11. #16Zoho Creator
    1
    brand score
    1first choice

Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.

Swipe left for the full report

Category · No Code Platforms

Model by model

Median rank each model gave each brand.

Scroll the table sideways to see every model →

claude-opus-5deepseek-v4.1-flashgemini-3.7-flashgpt-6-astrakimi-k3mistral-medium-3.5
Airtable3.02.01.01.01.01.0
Microsoft2.01.03.03.03.0
Zapier4.05.03.02.02.0
Retool1.04.04.53.0
Glide2.02.04.0
AppSheet2.55.02.55.04.0
Quickbase4.04.05.05.0
Make5.04.05.0
Softr5.04.04.0
Kissflow5.04.0
Bubble5.05.0
Appsmith5.0
Smartsheet5.05.0
Google3.0
Monday.com4.04.5
Zoho Creator5.0
1st
2
3
4
5th

A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.

Category · No Code Platforms

Where the models disagree

5 of 16 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.

  1. #4Retool

    deepseek-v4.1-flash 0 claude-opus-5 100 · 2 of 6 never named it

  2. #2Microsoft

    mistral-medium-3.5 0 deepseek-v4.1-flash 94 · 1 of 6 never named it

  3. #5Glide

    claude-opus-5 0 gpt-6-astra 68 · 3 of 6 never named it

  4. #3Zapier

    gpt-6-astra 0 kimi-k3 74 · 1 of 6 never named it

  5. #6AppSheet

    claude-opus-5 0 deepseek-v4.1-flash 52 · 1 of 6 never named it

  6. #1Airtable

    claude-opus-5 62 mistral-medium-3.5 100

A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.

Category · No Code Platforms

Named, and named first

Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.

the default answera narrow favouritelisted, rarely ledthe long tail12345678
  1. 1Airtable100/70
  2. 2Microsoft70/16
  3. 3Zapier70/0
  4. 4Retool45/25
  5. 5Glide48/0
  6. 6AppSheet40/0
  7. 7Quickbase28/0
  8. 8Make15/0

The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.

Category · No Code Platforms

What the models actually said

Models
6
Answers
60

Airtable leads with a score of 90 of 100, ranked by 6 of 6 models.

The models divide sharply over which platforms best serve companies building internal workflows without developers, with disagreement centered on whether "no developers" means zero technical literacy or simply no dedicated engineering staff. That distinction drives the biggest splits in the data.

The panel's deepest split is over platforms that require semi-technical skill: one model ranks Retool first while another omits it entirely, and Microsoft's scores span from 94 down to 28.

Airtable earns the clearest consensus, appearing in every model's top three and securing a Borda score of 89.67. Four models award it a perfect 100 and place it first, citing its spreadsheet-familiar interface, shallow learning curve, and ability to ship operational trackers and approval workflows in hours. Yet two models rank it second or third, stressing that it "can get expensive and hit performance limits at high record volumes or complex logic." The agreement is broad but not absolute—every model accepts Airtable for departmental use, but assessments diverge on how far those use cases stretch before requiring something heavier.

Airtable

per-model scores 62100 of 100 · mean 90 across 6 models

Microsoft provokes the sharpest split in the category. Deepseek-v4.1-flash ranks Power Platform first with a Borda of 94, emphasizing native Office 365 integration, enterprise governance, and Dataverse. At the other end, gemini-3.7-flash scores it at only 28, and three other models place it third. The divide turns on whether existing Microsoft investment and IT oversight are assets or barriers: models citing Gartner Magic Quadrant for Enterprise Low-Code Application Platforms×10 frame governance and security as strengths for large organizations, while others flag licensing complexity and formula-based customization as pushing citizen developers back toward IT support. Kimi-k3 sums up the tension, noting Power Platform's "steeper learning curve, confusing licensing tiers, and technical complexity often require citizen developer champions or semi-technical staff, which cuts against the pure no-code brief."

deepseek-v4.1-flash
94
rates it highest
mistral-medium-3.5
0
never named it

Microsoft · 94 points apart on a 0–100 scale

Retool triggers a similar but even more dramatic split. Claude-opus-5 ranks it first with a Borda of 100, calling it "the strongest option for building internal tools and workflow apps quickly on top of existing databases and APIs." Two other models omit it entirely, yielding a standard deviation above 38 points. The core disagreement is explicit: models valuing SQL and JavaScript access as power-user features rank Retool highly, while those interpreting the question as requiring zero scripting ability dismiss it. Kimi-k3 acknowledges it as "arguably the most powerful internal-tool builder here" yet notes "real value typically requires some SQL or JavaScript, so it fits a company 'without developers' poorly."

RetoolMicrosoft
  • 68mistral-medium-3.50
  • 100claude-opus-578
  • 4gemini-3.7-flash28
  • 0gpt-6-astra38
  • 18kimi-k356
  • 0deepseek-v4.1-flash94

Retool ahead on 2 of 6 models

Zapier divides the panel more subtly. Two models place it second, praising its 6,000-plus integrations and the addition of Tables, Interfaces, and Canvas for lightweight apps. Two others rank it fourth or fifth, arguing it "excels more at connecting systems than serving as a full internal app platform" and warning that task-based pricing "can become costly as task volume grows." The split reflects whether automation breadth compensates for weaker data modeling and UI capabilities. Sources cluster around G2www.g2.com/products/zapier/reviews×17 and the Zapier site itself, reinforcing both its leadership in workflow glue and its secondary status as an app builder.

Glide, AppSheet, and Quickbase show narrower but meaningful disagreement. Glide earns second-place ranks from gemini-3.7-flash and gpt-6-astra, which highlight its "exceptionally fast" spreadsheet-to-app conversion and mobile polish, but kimi-k3 places it fourth, noting "complex logic, permissions, and large-scale data needs can push it past its limits." Three models omit Glide entirely. AppSheet's deepseek score of 52 contrasts with kimi's 2, a split likely tied to Google Workspace penetration and tolerance for expression syntax. Quickbase appears in four models but never first, with all citing strong governance and relational data yet flagging dated UX and enterprise pricing as barriers for smaller or casual teams.

The long tail of brands—Make, Softr, Kissflow, Bubble, Appsmith, Smartsheet, Monday.com, and Zoho Creator—shows broad agreement on irrelevance or narrow fit. Make, Softr, and Bubble each appear in three models or fewer, and none breaks a Borda of 6. Models converge on reasons: Make is "powerful visual automation" but "steeper learning curve"; Softr is "one of the fastest ways to turn Airtable…into client portals" yet "relies heavily on external databases"; Bubble offers "unmatched visual design flexibility" but "may be overkill for simple workflow automation." These platforms occupy specialist lanes—complex branching logic, portal frontends, full-stack custom apps—that models judge peripheral to the internal-workflow mandate.

Across the category, G2www.g2.com/categories/no-code-development-platforms×46 is the most frequently recalled source, appearing 46 times, followed by product-specific G2 pages for Airtable, Zapier, and Microsoft. Gartner and Forrester analyst research cluster around Microsoft and Quickbase, reinforcing enterprise positioning, while Product Hunt and Hacker News references appear primarily for developer-adjacent tools like Retool and Appsmith. The hypothesis that sources drive model splits is plausible but not demonstrated: deepseek's preference for Microsoft and claude's for Retool both draw on similar G2 and vendor-documentation pools, suggesting the divergence reflects weighting of technical complexity over raw capability rather than distinct source sets.

Category · No Code Platforms

What shaped the answers

79 sources across 828 references, grouped by site from 254 recalled names. The top 5 carry 53% of them.

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · No Code Platforms
#1

Airtable

90
brand score
81first choice
12345

Summary

Airtable commands the strongest consensus in the category, earning a Borda score of 89.67 and ranking first among four of the six models. Every model placed it in its top three, yet the span of scores is wide: four models awarded it a perfect 100, while one scored it at 62, yielding broad agreement overall with a standard deviation above 15 points. The core narrative is consistent across all six: Airtable delivers an intuitive spreadsheet-database hybrid that non-technical teams can extend into custom interfaces, automations, and operational workflows without developer involvement. Models cite its shallow learning curve, rich template library, and broad integrations as key reasons business users can ship real internal tools in hours or days. Sources cluster around review platforms—G2www.g2.com/products/airtable/reviews×28, Capterra×19, and TechCrunch×17—alongside Airtable's own documentation and product pages, reinforcing a picture drawn from both user feedback and vendor materials.

Airtable
12345

Where models diverge is on the severity of Airtable's limits at scale. The four that rank it first emphasize its balance of power and accessibility for departmental use, while the two that place it second or third stress constraints around enterprise governance, complex permissions, very large data volumes, and intricate business logic. Claude-opus-5, which ranks it third, notes that teams "can get expensive and hit performance limits at high record volumes or complex logic," and deepseek-v4.1-flash cautions it "may hit limits for very complex enterprise-scale apps." No model questions its suitability for operational trackers, approvals, or project workflows—the disagreement centers on how far those use cases can stretch before requiring a more robust platform.

Airtable hits the sweet spot for non-technical teams: a spreadsheet-familiar interface on top of a real relational database, plus built-in automations, forms, and an interface designer for turning data into usable internal apps.

kimi-k3

Airtable is my strongest general recommendation for nontechnical teams because it combines familiar spreadsheet-style data management with custom interfaces, forms, and workflow automation.

gpt-6-astra

Per model summary

  • gemini-3.7-flash
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    10/10

    Airtable is consistently praised for combining an intuitive spreadsheet-like interface with powerful relational database capabilities, Interface Designer, and built-in automations that empower non-technical teams to quickly deploy sophisticated internal tools without code.

  • gpt-6-astra
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    10/10

    Airtable is repeatedly positioned as the strongest general-purpose recommendation that balances a familiar spreadsheet-style interface with custom interfaces and automations for building internal tools, though complex permissions, business logic, and higher usage may require more expensive plans or technical assistance.

  • kimi-k3
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    10/10

    Airtable is consistently described as hitting the optimal balance of accessibility and power, combining a spreadsheet-familiar interface with relational database capabilities, templates, and integrations that enable non-technical teams to ship working internal apps in hours or days without engineering support.

  • mistral-medium-3.5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    10/10

    Airtable is repeatedly characterized by its flexibility and ease of use, combining an intuitive spreadsheet-like interface with database power, automation capabilities, and strong integrations to enable custom workflows and internal tools without code.

  • deepseek-v4.1-flash
    brand score
    76
    first choice
    53
    median rank
    2.0
    present in
    10/10

    Airtable is repeatedly described as the most approachable and fastest way for non-developers to build internal apps and workflows, excelling at departmental use cases and quick iteration, but facing challenges with enterprise-scale governance and very complex processes.

  • claude-opus-5
    brand score
    62
    first choice
    35
    median rank
    3.0
    present in
    10/10

    Airtable is consistently characterized as the easiest entry point for non-technical teams, combining a familiar spreadsheet-database interface with interfaces, forms, and automations that enable rapid deployment of operational workflows, though it faces limitations with very large data volumes, complex logic, and granular permissions at scale.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gemini-3.7-flash7 sites · 30 references
gpt-6-astra2 sites · 20 references
Airtable×10 · 50%

airtable.com

  • ×10/
kimi-k35 sites · 30 references
Airtable×10 · 33%

airtable.com

  • ×10/
mistral-medium-3.53 sites · 30 references
deepseek-v4.1-flash4 sites · 30 references
Airtable×11 · 37%

airtable.com

claude-opus-58 sites · 30 references
Airtable×10 · 33%

airtable.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · No Code Platforms
#2

Microsoft

also named Microsoft Power Apps, Microsoft Power Platform, Power Apps

49
brand score
35first choice
  • gemini-3.7-flash 3.0
12345

Summary

Microsoft ranks as a moderately strong recommendation in the no-code space, earning a Borda score of 49 out of 100, but the models are sharply split on its placement. Five of the six models included Microsoft Power Platform in their rankings, yet their scores span from 94 down to 28, reflecting fundamentally different assessments of how well the platform serves companies seeking to build internal workflows without developer help. The central tension is consistent across nearly every response: Microsoft's deep native integration with the Office 365 and Teams ecosystem, coupled with enterprise-grade governance and security controls, makes it a natural and often cost-effective choice for organizations already standardized on Microsoft tooling. However, the same models repeatedly flag licensing complexity, a steeper learning curve, formula-based customization requirements, and governance overhead that can push citizen developers back toward IT support—qualities that cut against the pure no-code brief.

Microsoft

per-model scores 094 of 100 · mean 49 across 6 models

The sources models cited cluster heavily around analyst research and official documentation. G2×9 and Gartner Magic Quadrant for Enterprise Low-Code Application Platforms×10 together account for the majority of recalled references, emphasizing Microsoft's enterprise credentials and market-leader status. Models also drew extensively on Microsoft Power Platformpowerplatform.microsoft.com×5 and Microsoft Learn pages to validate technical capabilities around Power Apps, Power Automate, and Dataverse. The recurring theme is that Microsoft excels when existing organizational investment, IT governance needs, and Microsoft 365 integration matter most—but that same technical depth can require semi-technical champions or "citizen developer" programs to unlock value, making it less approachable than simpler alternatives for teams truly working without any developer support.

Power Apps is the strongest fit for enterprise internal workflows because it combines no-code building with deep Microsoft 365 integration, Dataverse, and Power Automate, while giving IT governance and security controls.

deepseek-v4.1-flash

Power Platform offers the deepest enterprise capability with strong native Microsoft 365 integration, governance, and security that many companies already license, making it especially natural for Microsoft-centric organizations. However, its steeper learning curve, confusing licensing tiers, and technical complexity often require citizen developer champions or semi-technical staff, which cuts against the pure no-code brief.

kimi-k3

Per model summary

  • deepseek-v4.1-flash
    brand score
    94
    first choice
    88
    median rank
    1.0
    present in
    10/10

    Power Apps is the strongest fit for enterprise internal workflows because it combines no-code building with deep Microsoft 365 integration, Dataverse, and Power Automate, while giving IT governance and security controls. The platform enables citizen developers to build apps while maintaining enterprise compliance, though it can involve licensing complexity and some learning curve.

  • claude-opus-5
    brand score
    78
    first choice
    48
    median rank
    2.0
    present in
    10/10

    Power Apps plus Power Automate is the default choice for Microsoft 365 organizations because of deep native integration with SharePoint, Teams, and Dataverse, along with enterprise-grade governance and compliance controls. The main trade-offs are a steeper learning curve and more complex licensing compared to lighter no-code tools.

  • gemini-3.7-flash
    brand score
    28
    first choice
    18
    median rank
    3.0
    present in
    5/10

    Microsoft Power Platform provides enterprise-grade capabilities with deep native integration across the Microsoft 365 ecosystem, allowing non-developers to build secure applications using familiar organizational data sources. The platform's administrative setup, governance controls, and licensing model can present a steeper learning curve compared to standalone alternatives.

  • gpt-6-astra
    brand score
    38
    first choice
    23
    median rank
    3.0
    present in
    8/10

    Power Apps and Power Automate are compelling for companies already using Microsoft 365 due to extensive integrations and enterprise governance, but formulas, licensing complexity, environment administration, and customization requirements can make truly developer-free implementation harder as applications grow. The platform ranks lower specifically for the requirement to work without developers, though existing Microsoft environment and expertise could improve its position.

  • kimi-k3
    brand score
    56
    first choice
    34
    median rank
    3.0
    present in
    9/10

    Power Platform offers the deepest enterprise capability with strong native Microsoft 365 integration, governance, and security that many companies already license, making it especially natural for Microsoft-centric organizations. However, its steeper learning curve, confusing licensing tiers, and technical complexity often require citizen developer champions or semi-technical staff, which cuts against the pure no-code brief.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

deepseek-v4.1-flash8 sites · 30 references
G2×10 · 33%

g2.com

claude-opus-54 sites · 30 references
gemini-3.7-flash6 sites · 15 references
gpt-6-astra2 sites · 17 references
Microsoft Learn — Power Apps×14 · 82%

learn.microsoft.com

kimi-k311 sites · 27 references
Gartner Magic Quadrant for Enterprise Low-Code Application Platforms×6 · 22%

no URL recalled

  • ×6named without a URL

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · No Code Platforms
#3

Zapier

39
brand score
24first choice
12345

Summary

Zapier scores a Borda of 39.33 across the panel, ranked by five of the six models but with considerable disagreement on where it belongs: per-model scores range from 0 to 74, yielding a standard deviation of 27.92. The strongest supporters place it consistently at rank two, praising its unmatched library of integrations—often cited as 6,000 to 8,000 apps—and its recent expansion into lightweight app-building through Tables, Interfaces, and Canvas. These models, drawing frequently from G2×9 and Zapierzapier.com×17, frame Zapier as the automation benchmark and an increasingly viable option for simple internal tools. At the other end, two models rank it fourth or fifth when they include it at all, arguing that its core strength remains cross-app workflow orchestration rather than full application development, and that task-based pricing can become prohibitive at scale.

Zapier

per-model scores 074 of 100 · mean 39 across 6 models

The central tension in the assessments is whether Zapier qualifies as a true internal-app platform or remains primarily an integration layer. Models that rank it higher acknowledge the addition of database and interface features but still note it "excels more at connecting systems than serving as a full internal app platform." Those placing it lower emphasize that "it is not a full app builder" and works best alongside dedicated platforms like Airtable or Retool. Even enthusiastic rankings carry the caveat that complex data modeling and rich UIs are not its natural territory, a view reinforced by citations to review sites such as Capterra×19 and workflow-automation category pages on G2. The result is a brand recognized universally for automation breadth yet assessed quite differently depending on whether the question prioritizes workflow glue or complete application construction.

Zapier remains the broadest integration layer, with thousands of app connectors plus Tables, Interfaces, and AI agents that now cover simple internal apps as well as automation. It's ideal for stitching SaaS tools together without engineering involvement, but it's weaker than dedicated app builders for data-heavy UIs and can become costly as task volume grows.

claude-opus-5

Zapier is the benchmark for no-code workflow automation, connecting thousands of apps so non-developers can automate cross-tool processes in minutes, and its Tables and Interfaces features now cover lightweight apps too. It ranks mid-list because it is primarily an automation layer between tools rather than a full internal app builder, and costs scale quickly with task volume.

kimi-k3

Per model summary

  • kimi-k3
    brand score
    74
    first choice
    45
    median rank
    2.0
    present in
    10/10

    Zapier is the benchmark for workflow automation across thousands of apps with expansion into lightweight app-building via Tables and Interfaces, though it excels more at connecting systems than serving as a full internal app platform, with per-task pricing that can scale quickly.

  • mistral-medium-3.5
    brand score
    72
    first choice
    43
    median rank
    2.0
    present in
    10/10

    Zapier is positioned as the leader in workflow automation with thousands of integrations for streamlining repetitive tasks, though it lacks full app-building capabilities and is best suited for process automation rather than building complete applications.

  • gemini-3.7-flash
    brand score
    40
    first choice
    25
    median rank
    3.0
    present in
    8/10

    Zapier is the industry standard for connecting disparate applications and automating workflows with the largest integration library, now expanding into app-building with Tables and Interfaces, though its core strength remains event-driven automation rather than complex UI-heavy applications.

  • claude-opus-5
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    10/10

    Zapier is positioned as the leading cross-app automation platform with the largest connector library, now extended with lightweight app-building features (Tables, Interfaces, Canvas), though it remains automation-first rather than a true app platform and task-based pricing can become expensive at scale.

  • deepseek-v4.1-flash
    brand score
    10
    first choice
    9
    median rank
    5.0
    present in
    4/10

    Zapier is recognized as a workflow automation and integration layer for connecting SaaS apps rather than a full app-building platform, working best as a complement to other no-code tools.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

kimi-k36 sites · 30 references
mistral-medium-3.54 sites · 30 references
G2×10 · 33%

g2.com

gemini-3.7-flash6 sites · 21 references
claude-opus-56 sites · 30 references
Zapier×12 · 40%

zapier.com

deepseek-v4.1-flash4 sites · 12 references
G2×5 · 42%

g2.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · No Code Platforms
#4

Retool

32
brand score
26first choice
12345

Summary

Retool registers sharply split opinions across the panel, scoring 31.67 Borda points with a standard deviation of 38.62. Four of the six models ranked it, but their placements diverge dramatically: one model consistently names it first, valuing its pre-built components and database connectivity for internal tools, while others place it fourth or fifth, citing its low-code nature as a poor fit for teams without any technical support. The core tension is whether "no developers" means zero technical literacy or simply no dedicated engineering staff. Models that interpret the brief strictly penalize Retool for expecting SQL or JavaScript familiarity, whereas those allowing for "semi-technical" or operations users with light scripting skills rate it as the category leader. Coverage sits at only 67 percent of the panel, and two models omit it entirely, reinforcing the platform's niche appeal.

Retool

per-model scores 0100 of 100 · mean 32 across 6 models

Source citations cluster around three themes. Reviews and community discussion—G2×9, Hacker News×2, and Gartner Peer Insights—validate Retool's strength in internal tooling but also surface the learning curve. Vendor documentation and announcements from the Retool site itself and TechCrunch illustrate feature depth and growth, while Product Hunt references highlight early adopter enthusiasm. The consensus in the reasonings is that Retool is "arguably the most powerful purpose-built tool for internal apps" yet "leans low-code rather than pure no-code," a duality that drives both its top rank among technical evaluators and its relegation by models prioritizing ease of use for fully non-technical teams.

Retool is the strongest option for building internal tools and workflow apps quickly on top of existing databases and APIs, with a drag-and-drop UI builder plus optional code for power users.

claude-opus-5

Retool is arguably the most powerful internal-tool builder here, with deep data connectivity and a huge component library. The catch is that it leans low-code rather than no-code: real value typically requires some SQL or JavaScript, so it fits a company 'without developers' poorly.

kimi-k3

Per model summary

  • claude-opus-5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    10/10

    Retool is the most capable platform for building internal tools quickly with drag-and-drop components and deep database/API connectivity. It scales from no-code usage to optional JavaScript/SQL for power users, though it performs best when a semi-technical builder is involved.

  • mistral-medium-3.5
    brand score
    68
    first choice
    40
    median rank
    3.0
    present in
    10/10

    Retool is designed specifically for building internal tools quickly with pre-built components and deep database/API integrations. It offers high customization and is developer-friendly, though it may require some technical knowledge or familiarity to unlock its full potential.

  • gemini-3.7-flash
    brand score
    4
    first choice
    3
    median rank
    4.0
    present in
    1/10

    Retool is a leading platform for building operational tools and dashboards connected to databases and APIs. It leans low-code rather than pure no-code and often requires SQL or scripting familiarity for maximum utility, which may challenge fully non-technical teams.

  • kimi-k3
    brand score
    18
    first choice
    14
    median rank
    4.5
    present in
    6/10

    Retool is arguably the most powerful purpose-built tool for internal apps with deep data connectivity and rich components. It ranks lower for strictly non-developer audiences because it leans low-code, typically requiring SQL or JavaScript to extract real value.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-56 sites · 30 references
G2×10 · 33%

g2.com

mistral-medium-3.57 sites · 30 references
G2×8 · 27%

g2.com

gemini-3.7-flash2 sites · 2 references
G2×1 · 50%

g2.com

kimi-k34 sites · 18 references

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · No Code Platforms
#5

Glide

28
brand score
18first choice
12345

Summary

Glide earns a Borda score of 28.33 across the panel, with only three of the six models ranking it at all—a pattern the data labels as "sharply split." Among the models that do include it, opinions range from consistent second-place rankings (gemini-3.7-flash and gpt-6-astra both score it at 68) down to a more cautious fourth or fifth spot (kimi-k3 at 34). The models agree that Glide excels at converting spreadsheets and databases into polished, mobile-friendly internal apps with minimal learning curve, often citing its drag-and-drop interface and pre-built components as ideal for non-technical teams building directories, inventory trackers, and field tools. Sources cluster around platform reviews on G2×9 and Product Hunt×12, alongside the platform's own documentation. The split emerges when models weigh simplicity against power: those ranking Glide higher emphasize speed and accessibility for straightforward use cases, while the more reserved assessment flags limitations in complex workflow orchestration, enterprise-grade permissions, and scaling multi-tiered data relationships.

Glide
12345

The disagreement centers on whether Glide's design-first, rapid-deployment strengths outweigh its constraints in advanced scenarios. Models that place it near the top see it as "exceptionally fast" and requiring "zero coding knowledge," well-suited to departmental tools that prioritize polish over deep logic. The more cautious view acknowledges the same ease of use but notes that "complex logic, permissions, and large-scale data needs can push it past its limits" and that "its automation depth and ability to handle complex logic lag behind the leaders." Three models declined to rank Glide at all, reinforcing a perception that it occupies a narrower niche—brilliant for simple, fast internal apps, but less versatile than platforms built for heavier workflow or governance demands.

Glide excels at turning existing business data from spreadsheets and databases into polished, responsive web apps in minutes. Non-technical staff can rapidly create internal portals, inventory trackers, and team directories with pre-built components.

gemini-3.7-flash

Glide is especially strong for nontechnical teams building polished, mobile-friendly internal apps like directories and inventory tools through an approachable visual builder. However, usage limits, complex workflow orchestration, integration requirements, and enterprise-level customization should be carefully evaluated before company-wide adoption.

gpt-6-astra

Per model summary

  • gemini-3.7-flash
    brand score
    68
    first choice
    42
    median rank
    2.0
    present in
    10/10

    Glide excels at rapidly converting existing business data from spreadsheets and databases into polished, responsive web and mobile apps through an intuitive drag-and-drop interface requiring zero coding knowledge. Its pre-built components enable exceptionally fast deployment of internal tools, though it faces constraints with highly complex enterprise logic and multi-tiered data relationships.

  • gpt-6-astra
    brand score
    68
    first choice
    42
    median rank
    2.0
    present in
    10/10

    Glide is especially strong for nontechnical teams building polished, mobile-friendly internal apps like directories and inventory tools through an approachable visual builder. However, usage limits, complex workflow orchestration, integration requirements, and enterprise-level customization should be carefully evaluated before company-wide adoption.

  • kimi-k3
    brand score
    34
    first choice
    22
    median rank
    4.0
    present in
    9/10

    Glide offers the fastest, most accessible path to turn spreadsheets into polished, mobile-friendly internal apps with a gentle learning curve ideal for simple operational tools. Its design-first approach produces excellent results for straightforward use cases, but complex logic, granular permissions, enterprise governance, and scaling needs expose its limitations compared to more powerful platforms.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gemini-3.7-flash5 sites · 25 references
gpt-6-astra1 site · 20 references
Glide×20 · 100%

glideapps.com

kimi-k36 sites · 27 references
Glide×9 · 33%

glideapps.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · No Code Platforms

How this was measured

The question

  • Unaided brand recommendation question: “What are the best no-code platforms for a company to build internal workflows and apps without developers?
  • The prompt asks each model to return exactly 5 brands, ranked 1 to 5, and for each one a reason for the recommendation and the sources that informed it.

Sampling

  • Models are not deterministic, so the question is asked over and over — 10 answers per model
  • Spellings of the same brand are normalized to the most commonly used form before counting

Measures of Position: Median and Quartiles

  • Median value indicates that in 50% of answers the brand held this position or higher.
  • Quartiles help visualize the spread of rankings per brand. Q1 indicates that 25% of answers had this rank or higher. Q3 means that 75% of answers ranked the brand as X or better.
  • Medians and quartiles are computed only from answers where the brand was present.

Coverage

  • Count of answers in which the given brand was named, as a percentage.

Brand score - Normalized Borda score

  • The brand score is calculated using a normalized Borda score. Borda count is a voting method: each ballot awards points by position instead of naming one winner. Our ranking responses from the LLM are always fixed to 5 answers, so we assign 100 points for rank 1, then 80, 60, 40, 20 — and 0 if the brand is missing.
Brand score=100ni=1n6ri5\text{Brand\ score} = \frac{100}{n} \sum_{i=1}^{n} \frac{6 - r_i}{5}

where

ri={1,,5if the brand appears6if the brand is not mentionedr_i = \begin{cases} 1,\dots,5 & \text{if the brand appears} \\ 6 & \text{if the brand is not mentioned} \end{cases}
  • This way we can calculate a common score for all brands mentioned across model responses. The score reflects both how high the brand was ranked and how frequently it was mentioned.

First choice score - adjusted MRR - top-rank indicator

  • The first choice score is based on Mean Reciprocal Rank (a common search-engine measure), but adapted to measure the position of a specific brand across repeated ranked responses.
  • Each occurrence receives a reciprocal-position score: 100 for 1st, 50 for 2nd, 33.3 for 3rd, 25 for 4th, 20 for 5th, and 0 when the brand is not mentioned. The scores are averaged across all responses.
First choice score=100ni=1nsi\text{First\ choice\ score} = \frac{100}{n} \sum_{i=1}^{n} s_i

where

si={1/riif the brand is present0if the brand is absents_i = \begin{cases} 1/r_i & \text{if the brand is present} \\ 0 & \text{if the brand is absent} \end{cases}
  • No rank below 1st place gets more than 50, so the score is heavily driven by first places.
Your category

Ask your own question of every model

Tell us what to ask and how wide to sample it — we run it and send back the report.

Sample size
Display results
0/1000

Sampling every model takes up to 24 hours once we start.