Your Brand has an AI reputation and it's shaping purchase decisions right now.
Swipe up to continue
9 AI models · 8 answers per model · ranked by score
ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0
Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.
Swipe left for the full report
Median rank each model gave each brand.
Scroll the table sideways to see every model →
| claude-opus-5 | deepseek-v4.1-flash | gemini-3.7-flash | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | gpt-6-astra | kimi-k3 | mistral-medium-3.5 | |
|---|---|---|---|---|---|---|---|---|---|
| HubSpot | 1.0 | 1.0 | 1.0 | 2.0 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 |
| Salesforce | 2.0 | 2.0 | 2.0 | 1.0 | 2.0 | 2.0 | 2.0 | 2.0 | 2.0 |
| Zoho | 4.0 | 3.5 | 3.0 | 4.0 | 3.0 | 3.0 | 3.0 | 3.0 | 3.0 |
| Pipedrive | 3.0 | 4.0 | 4.0 | 5.0 | 4.5 | 4.0 | 4.0 | 4.0 | 4.0 |
| Microsoft Dynamics 365 | 5.0 | 5.0 | — | 3.0 | 4.0 | 5.0 | 5.0 | 5.0 | — |
| Freshworks | — | 5.0 | 5.0 | — | 5.0 | 5.0 | — | — | — |
| Freshsales | — | — | — | — | — | — | — | — | 5.0 |
| monday.com | — | — | — | — | — | 5.0 | — | — | — |
| Zendesk | — | — | 5.0 | — | — | — | — | — | — |
A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.
0 of 9 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.
No brands where the models disagree
from here down, the models agreegemini-3.7-flash 0 → gpt-5.6-luna 60 · 2 of 9 never named it
gpt-5.6-luna 20 → claude-opus-5 53
gpt-5.6-luna 40 → mistral-medium-3.5 60
claude-opus-5 0 → gemini-3.7-flash 18 · 5 of 9 never named it
gpt-5.6-luna 88 → mistral-medium-3.5 100
claude-opus-5 80 → gpt-5.6-luna 93
A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.
Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.
The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.
HubSpot leads with a score of 98 of 100, ranked by 9 of 9 models.
Across the nine models this category resolves into a stable pecking order — HubSpot, then Salesforce, then Zoho, then Pipedrive — with the disagreement concentrated at the two ends rather than the middle.
The top of the ranking is where the models are most alike. HubSpot is the clear consensus pick, ranked first by seven of nine, and the only real dissent is gpt-5.6-luna, which parks it at a median rank of 2.
per-model scores 88–100 of 100 · mean 98 across 9 models
That same model is the mirror image on Salesforce: gpt-5.6-luna is alone in leading with it, scoring it 92.5 where six other models sit flat at 80. So the HubSpot-versus-Salesforce question is effectively decided by one model's ordering, not by any broad divide — the two brands trade first and second place almost entirely inside gpt-5.6-luna's answers.
HubSpot ahead on 8 of 9 models
The genuinely wide splits appear lower down. Microsoft Dynamics 365 carries the largest disagreement in the field: gpt-5.6-luna rates it at borda 60 and a median rank of 3, while claude-opus-5, gpt-6-astra and kimi-k3 land it at rank 5, and two models leave it off some runs entirely. The divide is about weighting, not fact — models agree it is enterprise-grade but disagree on how heavily licensing and implementation complexity should count against a general recommendation.
Microsoft Dynamics 365 · 60 points apart on a 0–100 scale
Below Dynamics, the tail is defined by coverage rather than conviction. Freshworks appears for only four models, and its spread is driven almost entirely by gemini-3.7-flash naming it in most runs while three others cite it rarely. Freshsales, monday.com and Zendesk are each the property of a single model — mistral-medium-3.5, gpt-5.6-terra and gemini-3.7-flash respectively — so their low, "agreed" scores mostly reflect shared silence rather than a shared judgment.
The recalled sources point the same way for everyone and do little to explain the splits. G2www.g2.com/categories/crm×116 dominates the entire category, with Capterra and PCMag close behind, and the brand-specific pages track wherever a model went. Because the heaviest-cited sources are common across models, they are unlikely to be what separates gpt-5.6-luna from the rest; that divergence looks like a difference in how the model weighs enterprise capability against ease of adoption — a hypothesis the reasonings support but the source data does not confirm.
54 sources across 992 references, grouped by site from 231 recalled names. The top 5 carry 62% of them.
g2.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named HubSpot CRM
HubSpot is the panel's clearest consensus pick, carrying a Borda score of 97.78 and ranked by all 9 of 9 models. The agreement is tight: per-model scores span 87.5 to 100 with a standard deviation of 3.99, which the report labels "models agree." Seven of the nine models placed it at a median rank of 1, and only gpt-5.6-luna diverges meaningfully, settling at a median rank of 2 and giving it the lowest score in the set.
per-model scores 88–100 of 100 · mean 98 across 9 models
The reasoning behind the ranking is remarkably uniform. Models repeatedly cite the same three attributes — an approachable, intuitive interface, a genuinely useful free tier, and integrated marketing, sales, and service tools — while consistently attaching the same caveat that costs rise steeply as contacts, seats, and advanced features are added. Several models frame it explicitly as the safe default recommendation when requirements are unknown. The recalled sources cluster under user-review and directory themes, most prominently G2www.g2.com/categories/crm×116 and the vendor's own HubSpot CRMwww.hubspot.com/products/crm×19, supplemented by Capterra, PCMag, and Gartner references.
“HubSpot offers the best balance of usability, a genuinely useful free tier, and a unified suite spanning marketing, sales and service, which makes it the safest general recommendation for small and mid-sized teams.”
“HubSpot is the strongest general recommendation because it combines an approachable interface, a useful free tier, and well-integrated sales, marketing, and service tools.”
Consistently frames HubSpot as the safest general recommendation for small and mid-sized teams, citing its usability, useful free tier, unified marketing/sales/service suite, and best-in-class onboarding, while noting costs escalate as contacts and add-ons grow.
Recommends HubSpot first for most small and mid-sized teams because of its intuitive interface, useful free tier, fast setup, and integrated sales/marketing/service tools, with the caveat that costs and customization limits appear as usage scales.
Emphasizes HubSpot's exceptionally intuitive interface, robust free tier, and seamless all-in-one ecosystem across marketing, sales, and service, presenting it as the best-balanced, scalable choice despite expensive enterprise tiers.
Repeatedly calls HubSpot the strongest general recommendation for its approachable interface, useful free tier, and integrated marketing/sales/service ecosystem, tempered by substantially rising costs as features and contacts grow.
Describes HubSpot as the strongest all-around recommendation for small and midsize teams due to its approachable CRM, useful free tier, and natural integration with marketing and service workflows, with the trade-off of rising costs from added hubs, seats, and automation.
Treats HubSpot as its default recommendation when company size and requirements are unknown, citing its approachable CRM connected to sales, marketing, and service tools and easy start, while warning costs rise substantially with advanced features and expanded usage.
Recommends HubSpot as the best overall balance of usability and capability for most businesses, stressing its intuitive interface, genuinely useful free tier, fast adoption, and integrated hubs, while noting costs climb steeply at higher tiers.
Consistently describes HubSpot as a highly intuitive, user-friendly CRM with a robust free tier and seamless integration of marketing, sales, and service tools, ideal for small to mid-sized and scaling businesses.
Positions HubSpot as an approachable top pick for small and midsize businesses, highlighting its free tier and unified marketing, sales, service, and content tools, while noting costs rise as advanced features and contacts are added.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
capterra.com
g2.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Salesforce scores strongly across the panel, drawing a Borda of 82.22 of 100 and recognition from all 9 of the 9 models. The spread is narrow — per-model borda runs from 80 to 92.5 with a standard deviation of 3.99, which the report labels as agreement. That consensus is thematic as well as numerical: nearly every model frames Salesforce as the enterprise standard, praised for unmatched customization, its AppExchange ecosystem, and depth of automation and reporting, while placing it just behind HubSpot for general buyers because of implementation complexity, administrative overhead, and total cost of ownership. The recalled evidence sits under those themes consistently, with G2×1 appearing across most models, reinforced by Gartner's Magic Quadrant citations and Salesforce's own product pages.
per-model scores 80–93 of 100 · mean 82 across 9 models
The main point of divergence is placement at the top rather than presence: gpt-5.6-luna ranks it first in 5 of 8 runs for a borda of 92.5, whereas claude-opus-5, gemini-3.7-flash, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra, and mistral-medium-3.5 each land at 80, never ranking it first. The first-choice score of 55.56 of 100 reflects that split between models that lead with Salesforce and those that reserve first place for a more turnkey option.
“Salesforce is the strongest overall recommendation for organizations that need a highly capable, scalable CRM with extensive customization, automation, analytics, and integrations.”
“It ranks lower than HubSpot here only because of higher cost, implementation complexity and admin overhead for smaller teams.”
Frames Salesforce as the strongest overall choice for scalable, highly customizable CRM with broad automation, analytics, and integrations. Its breadth and ecosystem suit larger organizations, but implementation complexity and cost can be significant for smaller teams.
Consistently frames Salesforce as the market leader with unmatched customization, AppExchange ecosystem, and enterprise-grade reporting, ideal for large or complex organizations. It ranks below HubSpot mainly due to cost, implementation complexity, and admin overhead that make it overkill for smaller teams.
Repeatedly describes Salesforce as the most powerful and customizable CRM with an unmatched ecosystem for complex enterprise sales. It ranks it lower because of the steep learning curve, high cost, and the need for dedicated admins, making it less turnkey for smaller teams.
Consistently calls Salesforce the industry standard/benchmark for enterprise customization, analytics, and AppExchange integrations. It places it just behind HubSpot due to steep learning curve, complex setup, administrative overhead, and high total cost of ownership.
Uniformly emphasizes Salesforce's exceptional customization, automation, reporting, and ecosystem for complex or enterprise deployments. It ranks below HubSpot for general audiences because implementation, administration, and total cost can be substantial.
Consistently positions Salesforce as excellent for larger or complex organizations needing deep customization, integrations, governance, and a broad partner ecosystem. Its power comes with higher implementation effort, administration, and total cost than simpler options.
Repeatedly recommends Salesforce for organizations needing extensive customization, sophisticated sales workflows, and a large integration ecosystem. It ranks it below HubSpot for general buyers because implementation, administration, and total cost can be excessive for simpler needs.
Consistently describes Salesforce as the industry standard and most powerful, customizable CRM with an unmatched AppExchange ecosystem. It ranks it second because licensing cost, admin overhead, and implementation complexity make it overkill for many smaller teams.
Uniformly frames Salesforce as the enterprise-level industry leader with unparalleled customization, scalability, and ecosystem. It notes a steeper learning curve and complexity that can make it overkill for smaller businesses.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named Zoho CRM
Zoho occupies a firmly mid-tier position in this study, ranked by all 9 of the 9 models and earning a Borda score of
Stat unavailablewith no model placing it first. The spread of per-model scores is narrow — from 40 to 60, a standard deviation of 6.64 — which the report labels as agreement. That consensus rests on a shared reading: models consistently frame Zoho as the strongest value-for-money option, feature-rich and tightly bound to the wider Zoho suite, well-suited to budget-conscious SMBs, while noting an interface and configuration experience that feels less polished than higher-ranked alternatives. The clustering of median ranks around 3 and 4 reflects that this trade-off caps its standing rather than sinking it.
per-model scores 40–60 of 100 · mean 54 across 9 models
The recalled sources sit largely under review-aggregator and vendor themes, with G2×1 and Capterra×4 recurring across nearly every model alongside Zoho's own pages, and PCMag×4 supporting the value-and-features framing. The value case and the polish caveat surface together throughout the reasonings.
“Zoho CRM delivers the strongest feature-per-dollar value, with automation, analytics and a huge suite of integrated business apps at a fraction of competitors' prices.”
“Zoho provides broad CRM functionality and strong value, especially for small and midsize businesses that may also use its wider business-software suite.”
Emphasizes an exceptionally feature-rich, cost-effective CRM with native integration into the broader Zoho ecosystem, offset by an interface that can feel cluttered, fragmented, or disjointed.
Consistently characterizes Zoho as strong value with broad functionality tied to its wider business-software suite for SMBs, while noting the interface and configuration feel less polished than top-ranked options.
Highlights Zoho's strong value and broad CRM capability, especially for cost-conscious teams using its wider suite, while noting the interface and setup feel less polished than leading alternatives.
Ranks Zoho third as an affordable, value-oriented choice for budget-conscious SMBs with a broad business-software ecosystem, noting that configuring across its many options and applications takes more effort than simpler alternatives.
Consistently frames Zoho as the best value with a broad feature set (AI, automation, multichannel) at a fraction of leaders' prices, ranked mid-pack because its interface, polish, and ecosystem trail HubSpot and Salesforce.
Repeatedly describes Zoho as a cost-effective, feature-rich option with automation and AI, well-suited to small businesses and startups, though possibly lacking advanced features for larger enterprises.
Repeatedly positions Zoho as excellent value with a broad, customizable feature set for budget-conscious SMBs, while noting its interface can feel less polished or cluttered and the wider ecosystem sprawling.
Consistently frames Zoho as the best feature-per-dollar value, integrated with the broad Zoho One suite and suited to budget-conscious SMBs and international teams, while noting a less polished UI and inconsistent support relative to HubSpot and Salesforce.
Describes Zoho as a broad, cost-effective, value-oriented CRM with a large application ecosystem, ranked below leaders because its interface and administration experience feel less polished or streamlined.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
capterra.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Pipedrive scores a Borda average of 38.61 of 100 and is ranked by all 9 of the 9 models, a consistent mid-table placement rather than a top pick — none of the 9 models ranked it first. With per-model borda spanning 20 to 52.5 and a standard deviation of 8.51, the report labels this "broad agreement," and the spread reflects fine gradations rather than genuine dispute: claude-opus-5 places it highest at 52.5, while gpt-5.6-luna anchors the bottom at 20.
per-model scores 20–53 of 100 · mean 39 across 9 models
The models converge on a single reading, with the theme steady across all nine: Pipedrive is a sales-first, visually-driven pipeline CRM that small teams adopt quickly with low overhead, held back only by comparatively thin marketing, service, reporting, and enterprise capabilities. That shared framing — strong at its narrow purpose, weaker as an all-in-one platform — explains why it clusters at rank 4 for most models while gpt-6-astra notes it "could be my first choice for a small, sales-focused team." The recalled evidence sits mainly on review aggregators and the vendor's own pages, with G2×1 and Capterra×4 recurring most heavily alongside frequent citations of Pipedrive's own site.
“Pipedrive is a sales-first, pipeline-centric CRM that small sales teams can adopt in days, with clear visual deal stages and low pricing.”
“Pipedrive excels in simplicity and ease of use, focusing on sales pipeline management.”
Pipedrive is consistently framed as a sales-first, visual pipeline CRM that small teams adopt quickly with low overhead and affordable pricing, ranked lower because its marketing, service, and reporting capabilities are comparatively thin.
Pipedrive is portrayed as a focused, easy-to-use pipeline CRM ideal for small sales teams, ranked lower because it lacks the broader marketing, service, and customization depth of all-in-one platforms.
Pipedrive is described as purpose-built around visual pipeline management for deal-driven sales teams with minimal onboarding, ranked lower for lacking native marketing and customer support features found in full-suite platforms.
Pipedrive is framed as a strong fit for sales-led teams wanting a clean visual pipeline and quick adoption without enterprise complexity, ranked lower for limited native marketing, service, and cross-functional capabilities.
Pipedrive is portrayed as excellent for sales teams prioritizing pipeline visibility and deal tracking, ranked fourth for general use because it is less comprehensive for marketing and customer-service needs, though it could top a small sales team's list.
Pipedrive is described as a sales-first CRM built around a clean visual pipeline that reps enjoy using, ranked lower because its marketing, service, and reporting features are thin and often rely on add-ons.
Pipedrive is consistently characterized by simplicity and ease of use for sales-focused teams, with the caveat that it offers fewer marketing and service features than competitors like HubSpot or Salesforce.
Pipedrive is consistently described as easy to adopt and effective for visual, sales-focused pipeline management, ranked lower because its marketing, customer-service, and enterprise capabilities are narrower than broader platforms.
Pipedrive is presented as an intuitive, sales-focused CRM strong on visual pipeline and deal management, ranked lower because it is less comprehensive for marketing, service, analytics, and enterprise needs.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
pipedrive.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named Microsoft, Microsoft Dynamics, Microsoft Dynamics 365 Sales
Microsoft Dynamics 365 scores modestly across the panel, with a Borda average of 20.56 of 100 and a first-choice score of 15.86, ranked by 7 of the 9 models but placed first by none. The models converge on a shared narrative more than the numbers alone suggest: Dynamics 365 is consistently framed as a capable, enterprise-grade CRM whose natural home is organizations already committed to Microsoft 365, Teams, Azure and the Power Platform, but one that ranks lower as a general recommendation because of complex licensing, partner-led implementation, and a steeper adoption curve. Where the models differ is in how much weight that complexity carries: gpt-5.6-luna places it at a median rank 3 with a borda of 60, while claude-opus-5, gpt-6-astra and kimi-k3 all settle at a median rank 5, several of them explicit that the placement reflects ease of adoption rather than capability.
That gap defines the spread the report labels "broad agreement" despite a standard deviation of 16.78. The evidence behind these judgments leans on vendor and review sources — Microsoft Dynamics 365 Saleswww.microsoft.com/en-us/dynamics-365/products/sales×16 recurs across models alongside G2 review pages and Gartner's Magic Quadrant for Sales Force Automation Platforms — supporting a consistent picture of a powerful platform whose value is conditional on ecosystem fit.
per-model scores 0–60 of 100 · mean 21 across 9 models
“It ranks fifth as a default recommendation because configuration and licensing can be complex, though it could rank much higher for a Microsoft-centric enterprise.”
“I rank it last here mainly because it's the least broadly applicable of the five, not because of capability.”
Consistently describes it as a strong, flexible enterprise CRM for companies already invested in Microsoft 365, Teams, Azure and Power Platform, but notes its configuration, licensing, and expertise requirements make it more complex than simpler options.
Repeatedly frames it as a compelling choice for organizations deeply invested in the Microsoft ecosystem, but ranks it lower (fourth or fifth) as a general recommendation due to licensing and implementation complexity, especially for smaller teams.
Consistently frames Dynamics 365 as a powerful enterprise option best suited for organizations already invested in Microsoft 365, Teams, Azure and Power Platform, but ranks it last due to complex licensing, partner-led implementation, and lower ease of adoption—not capability.
Repeatedly positions it as a robust CRM ideal for organizations already invested in the Microsoft ecosystem, while ranking it low generally because of complex, costly licensing and implementation that feels heavier than simpler alternatives.
Consistently presents it as a sensible option for larger organizations invested in Microsoft 365, Teams, Power Platform and Azure, but ranks it lower for general use because implementation and administration are complex and often require specialist or partner support.
Repeatedly ranks it fifth as a general-purpose default due to configuration and licensing complexity, while emphasizing it could rank much higher for Microsoft-centric enterprises where its ecosystem fit is a strong advantage.
Consistently describes it as a powerful, enterprise-grade CRM with deep Microsoft ecosystem integration, but ranks it last because of confusing licensing, high implementation cost, steep learning curve, and limited value outside Microsoft-centric organizations.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
microsoft.com
microsoft.com
no URL recalled
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
where
where
6 AI models · 5 answers per model · ranked by score
ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0
Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.
Swipe left for the full report
Median rank each model gave each brand.
Scroll the table sideways to see every model →
| claude-opus-5 | gemini-3.7-flash | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | kimi-k3 | |
|---|---|---|---|---|---|---|
| Accenture | 2.0 | 3.0 | 1.0 | 1.0 | 1.0 | 1.0 |
| McKinsey & Company | 1.0 | 1.0 | 2.0 | 3.0 | 2.0 | 2.0 |
| Boston Consulting Group | 4.0 | 2.0 | 4.0 | 4.0 | 4.0 | 3.0 |
| Deloitte | 3.0 | 5.0 | 3.0 | 2.0 | 3.0 | 4.0 |
| IBM Consulting | 5.0 | 5.0 | 5.0 | 5.0 | 5.0 | 5.0 |
| Bain & Company | — | 4.0 | 5.0 | — | 5.0 | 4.0 |
A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.
0 of 6 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.
No brands where the models disagree
from here down, the models agreegemini-3.7-flash 28 → gpt-5.6-sol 72
gpt-5.6-sol 52 → gemini-3.7-flash 100
gemini-3.7-flash 60 → gpt-5.6-terra 100
gpt-5.6-terra 40 → gemini-3.7-flash 80
claude-opus-5 0 → gemini-3.7-flash 28 · 2 of 6 never named it
gemini-3.7-flash 4 → gpt-5.6-sol 28
A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.
Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.
The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.
Accenture leads with a score of 91 of 100, ranked by 6 of 6 models.
Across the six models, the top of the field is settled and the disagreements concentrate in the middle. Accenture and McKinsey trade the top two spots, while Boston Consulting Group, Deloitte, IBM Consulting, and Bain sort into a more contested lower tier where individual models diverge noticeably.
The clearest consensus is on the extremes. Accenture holds a category-wide median of 1 and IBM Consulting a rock-solid median of 5 (Q1 5, Q3 5) with every model landing on exactly 5. IBM's uniformity extends to reasoning: all six credit engineering depth, watsonx, and regulated-environment fit, and five of the six cite vendor bias toward IBM's own stack as the reason it sits mid-tier. There is little to separate the models on either brand.
McKinsey is also broadly agreed on in substance — C-suite strategy credibility and QuantumBlack depth recur everywhere — even though its placement varies by a rank or two.
The sharpest divergences are best read model by model:
BCG shows the widest disagreement of any brand: gemini places it 2nd, kimi 3rd, and the remaining four settle at 4. Deloitte spans an even wider model range (2 through 5) despite a tidy 3.5 category median — a case where the aggregate obscures real disagreement.
The source data is suggestive rather than conclusive, and these links are hypotheses.
The most legible pattern is on Accenture: models ranking it 1st lean on first-party material (Accenture service pages, Technology Vision, the $3B investment newsroom item), whereas gemini-3.7-flash — the lowest ranker — is described as drawing almost entirely on external outlets (Gartner, IDC MarketScape, Bloomberg, FT, WSJ) with no first-party Accenture citations. It is plausible that a reliance on outside analyst and press framing produced its more measured, integrator-oriented view, but the data shows correlation, not cause.
A similar hypothesis fits Deloitte: gemini is again noted as leaning more on Gartner, IDC MarketScape, Bloomberg, FT, and Fortune, coinciding with its lowest placement — consistent with a pattern of external sourcing tracking cooler positioning, though the sample is too small to confirm.
For McKinsey and BCG, the recalled sources are heavily first-party (State of AI, QuantumBlack; BCG X, AI Radar) across models regardless of rank, so the sourcing does not obviously explain why claude and gemini rate them higher than the gpt models. Here the placement differences appear to come from how each model weighs strategy versus at-scale delivery rather than from distinct source pools.
Bain and IBM rest on fewer ranked answers (11 and 19) than the leaders (30), so their model-level medians are less stable. All sources are recalled by the models from memory rather than verified citations, which limits how far any source-to-view link can be pushed.
91 sources across 428 references, grouped by site from 256 recalled names. The top 5 carry 44% of them.
mckinsey.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Accenture lands at the top of the field, with a median rank of 1 (Q1 1, Q3 2) across 30 ranked answers from 6 models. Agreement is strong at the high end: four models — gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, and kimi-k3 — place it at a median rank of 1 with no spread (Q1 1, Q3 1), while claude-opus-5 sits slightly lower at median 2 (Q1 1, Q3 2) and gemini-3.7-flash is the clear outlier at median 3 (Q1 3, Q3 3). The recurring theme uniting the top rankings is scale of delivery and execution: strategy-through-implementation breadth, cloud and integration depth, a multi-billion-dollar AI investment, and hyperscaler alliances. Several models frame Accenture as the strongest default for moving beyond pilots to enterprise-wide deployment, though most also qualify that its strength skews toward execution rather than top-table strategy — a point noted explicitly by kimi-k3, claude-opus-5, and gemini-3.7-flash, with gpt-5.6-sol adding caveats around scope, cost, and complexity.
The sources behind these themes divide along two lines. Models placing Accenture at rank 1 lean heavily on first-party material — Accenture's own AI and data-AI service pages, Technology Vision 2024, the Reinvention in the Age of Generative AI insight, annual reports, and the $3 billion AI investment newsroom announcement — supplemented by third-party analysts such as Gartner, Everest Group, IDC, Forrester, and the Stanford AI Index. claude-opus-5 grounds its execution-focused view in earnings-release and investor material on generative AI bookings alongside NVIDIA/Microsoft alliance announcements, Gartner's Magic Quadrant, and HFS/Everest Group assessments. By contrast, gemini-3.7-flash, the lowest ranker, draws almost entirely on external outlets — Gartner, IDC MarketScape, Bloomberg, Financial Times, Forrester, and the Wall Street Journal — with no first-party Accenture citations, which aligns with its more measured positioning of the firm as an integration and infrastructure play rather than a strategy leader.
Accenture is presented as the strongest overall choice, combining strategy, technology implementation, and managed services at scale, particularly for moving beyond pilots to enterprise-wide deployment.
Accenture is framed as the strongest default choice given its broad combination of strategy, integration, cloud, and delivery scale, while noting clients must control scope, cost, and complexity.
Accenture is consistently the strongest all-around choice, pairing board-level strategy with large-scale delivery to move from pilots to enterprise-wide operating-model and workforce change.
Accenture is highlighted for unmatched implementation and delivery scale, multi-billion-dollar AI investment, and key platform alliances, though its brand skews toward execution rather than top-table strategy.
Accenture is characterized as having the largest scaled AI delivery capability — multi-billion-dollar gen-AI bookings, tens of thousands of AI-trained practitioners, and deep hyperscaler partnerships — making it the strongest choice for end-to-end execution rather than strategy alone.
Accenture is described as offering unmatched global scale and systems integration for massive IT and infrastructure overhauls, backed by multi-billion-dollar AI investment, though it leans more toward implementation than pure top-down strategy.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
accenture.com
accenture.com
accenture.com
newsroom.accenture.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named McKinsey
McKinsey & Company holds a strong overall position, with a median rank of 2 (Q1 1, Q3 2) across 30 ranked answers from 6 models, and appears under both "McKinsey & Company" and "McKinsey." The models converge on a consistent underlying rationale: McKinsey is valued for C-suite strategy credibility, use-case prioritization, and operating-model redesign, with its QuantumBlack unit repeatedly cited as adding technical and data-science depth. This theme is supported by heavy recall of McKinsey's own research and capability pages — particularly the State of AI report and QuantumBlack insights — alongside third-party analyst and business-press sources such as Gartner, Forrester, Forbes, Harvard Business Review, IDC MarketScape, and coverage of McKinsey's internal Lilli tool (Reuters, CNBC).
Agreement on the reasoning is high, but placement varies more than the median suggests. Gemini-3.7-flash and claude-opus-5 rank it strongest (both median 1, with gemini at Q1 1/Q3 1 and no noted drawbacks), while gpt-5.6-luna, gpt-5.6-terra, and kimi-k3 land at median 2, and gpt-5.6-sol places it lowest at median 3 (Q1 3, Q3 4). The recurring caveat behind the lower placements is consistent across those models: McKinsey is seen as ranking below Accenture or larger systems integrators for hands-on, at-scale implementation and ongoing operations, with premium pricing and a thinner delivery bench sometimes requiring additional implementation partners.
McKinsey's QuantumBlack unit is framed as combining board-level strategy credibility with a large bench of data scientists, backed by influential research (State of AI), making it best for enterprise-wide value-case definition and operating-model redesign — though at premium cost and often needing implementation partners at scale.
McKinsey is consistently presented as pairing world-class C-suite strategy with deep technical execution via QuantumBlack, positioning it as the premier end-to-end choice for large-scale Fortune 500 transformation and change management, with no noted drawbacks.
McKinsey is described as strong at AI strategy, use-case prioritization, and operating-model change, but consistently ranked below Accenture because large-scale technical implementation and ongoing operations may require systems integrators or other partners.
McKinsey is framed as excellent for defining the AI agenda, prioritizing use cases, and operating-model/executive alignment, with QuantumBlack adding capability, while consistently cautioning clients to validate hands-on engineering and delivery capacity for large implementations.
McKinsey is characterized by unmatched C-suite credibility, strategy rigor, and QuantumBlack's technical depth plus influential State of AI research, but ranked second to Accenture due to premium pricing and a thinner hands-on, at-scale implementation bench.
McKinsey is repeatedly cited for executive alignment, operating-model redesign, and tying AI to measurable business value with QuantumBlack support, but ranked below larger systems integrators for implementation-heavy, at-scale programs that may need additional partners.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
mckinsey.com
mckinsey.com
mckinsey.com
mckinsey.com
mckinsey.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named BCG
Boston Consulting Group lands mid-pack across the six models, with a median rank of 3 (Q1 3, Q3 4) over 30 ranked answers. There is moderate agreement on its placement but a visible spread in enthusiasm: gemini-3.7-flash ranks it highest at a median of 2 (Q1 2, Q3 2), kimi-k3 places it at 3 (Q1 3, Q3 3), and the remaining four models — claude-opus-5, gpt-5.6-luna, gpt-5.6-sol, and gpt-5.6-terra — each land at a median of 4, with gpt-5.6-terra the most consistent at that level (Q1 4, Q3 4). The firm is named as both Boston Consulting Group and BCG.
The recurring theme is a strategy-plus-build profile: models consistently credit BCG for elite corporate strategy, operating-model redesign, and value/ROI framing, paired with technical delivery through BCG X, and they cite the firm's research output, including the 10-20-70 framing and OpenAI/Anthropic partnerships. The countervailing theme, and the main reason it ranks behind McKinsey, Accenture, and Deloitte in several answers, is a perceived thinner global implementation, systems-integration, and managed-services footprint for full Fortune 500-scale rollouts. Sources cluster around BCG's own materials — BCG X, BCG's AI capabilities pages, the AI Radar, and the BCG Henderson Institute — supplemented by MIT Sloan Management Review (including joint MIT Sloan–BCG research), Harvard Business Review, and the Harvard Business School–BCG field experiment on GenAI productivity. Analyst and press validation appears through Gartner, IDC MarketScape, the Forrester Wave (AI Consultancies / AI Services), Forbes, the World Economic Forum, and Financial Times and Reuters coverage of BCG's AI revenue share.
BCG pairs elite corporate strategy with technical build and engineering through BCG X, emphasizing proprietary domain-specific AI solutions, business value/ROI, and responsible AI governance for large enterprises.
BCG blends strategy with BCG X build capability and highly cited AI research (e.g., 10-20-70 framing), but ranks behind McKinsey and Accenture because its large-scale global delivery bench is thinner and less proven at full enterprise-wide deployment.
BCG X provides genuine build capability alongside top-tier strategy, with roughly a fifth of revenue from AI work and notable OpenAI/Anthropic partnerships, but its delivery footprint is smaller than Accenture's or Deloitte's for large-scale integration and long-term run services.
BCG excels at business strategy, innovation, and business-model redesign around AI, but ranks below implementation-led firms because a Fortune 500 transformation may need more systems integration, engineering capacity, and managed services.
BCG combines senior-level strategy and operating-model redesign with product and technical build capabilities, strong for high-value differentiated AI use cases, though its global implementation and managed-services scale trails Accenture's and Deloitte's.
BCG is compelling for value-led AI strategy, portfolio prioritization, and business-model innovation supported by BCG X, but ranks below the top firms because Fortune 500-scale rollouts often need larger global implementation and managed-delivery depth.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
bcg.com
bcg.com
bcg.com
bcg.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Deloitte lands in the middle of the pack overall, with a median rank of 3.5 (Q1 3, Q3 4) across 30 ranked answers from 6 models. Model-level placement spans a moderate range: gpt-5.6-sol ranks it highest at a median of 2 (Q1 2, Q3 2), while claude-opus-5, gpt-5.6-luna, and gpt-5.6-terra cluster at a median of 3, and kimi-k3 and gemini-3.7-flash place it lower at medians of 4 and 5 respectively. The models converge on a shared characterization of breadth even as they diverge on how favorably to weigh it.
The dominant theme is Deloitte's end-to-end coverage tying AI to risk, regulatory, cyber, finance, tax, and workforce processes, making it a strong fit for regulated Fortune 500 enterprises—a framing echoed by gpt-5.6-sol, claude-opus-5, gpt-5.6-luna, and gpt-5.6-terra, several of which add the recurring caveat that delivery quality and AI depth vary by practice, member firm, and geography. The lower-ranking models reframe that same breadth as a limitation: kimi-k3 sees the AI work as more generalist and integration-led and less distinctive than MBB strategy houses or Accenture, while gemini-3.7-flash positions Deloitte as an execution and compliance partner rather than a frontier strategy or deployment leader. Sources underpinning these themes lean heavily on Deloitte's own materials—the State of Generative AI in the Enterprise reports, the AI Institute, the Trustworthy AI framework, Tech Trends, and various AI and Data service pages—supplemented by third-party analyst references including Gartner (Magic Quadrant for Data and Analytics Service Providers and the Market Guide for AI Consulting and System Integration Services), the IDC MarketScape, Forrester Wave, and Everest Group PEAK Matrix, with gemini-3.7-flash notably drawing more on Gartner and IDC MarketScape plus outlets such as Bloomberg, the Financial Times, and Fortune.
Deloitte is emphasized for integrating AI with enterprise processes—risk, cyber, finance, tax, regulatory and workforce transformation—making it strong for complex regulated companies, though delivery quality and strategic distinctiveness vary by team and region.
Deloitte is consistently framed as a broad, end-to-end choice combining strategy, technology, risk and regulatory/AI-governance capability (Deloitte AI Institute, Trustworthy AI) well-suited to regulated Fortune 500 industries, with the recurring caveat that delivery quality varies by practice and geography.
Deloitte is portrayed as offering broad multidisciplinary coverage (strategy, risk, cybersecurity, technology, compliance) valuable for highly regulated Fortune 500 firms, with the consistent caveat that experience and AI depth vary by member firm, geography and delivery team.
Deloitte is consistently positioned as a strong fit where AI must be tightly integrated with risk, regulatory, cyber, finance, tax and workforce processes across complex or regulated enterprises, with a recurring recommendation to vet the specific proposed team since experience varies by practice and geography.
Deloitte is credited with enormous breadth, its AI Institute, generative AI research and strength in regulated sectors tying AI to risk/tax/audit, but repeatedly ranked lower because its AI work is perceived as less distinctive and more generalist/integration-led than the MBB strategy houses or Accenture.
Deloitte is described as offering comprehensive breadth in AI governance, regulatory compliance and risk via its Trustworthy AI framework and AI Institute/Academy, ideal for heavily regulated enterprises, but ranked fifth because it is seen as an execution/compliance partner rather than a frontier strategy or deployment leader.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
www2.deloitte.com
www2.deloitte.com
deloitte.com
deloitte.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named IBM
IBM Consulting occupies a consistent mid-tier position across the models, with a median rank of 5 (Q1 5, Q3 5) over 19 ranked answers from 6 models. The agreement here is unusually tight: every one of the six models—claude-opus-5, gemini-3.7-flash, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, and kimi-k3—landed on a median rank of 5, with no dispersion between the first and third quartiles at either the aggregate or individual level. The firm is referenced under two names, IBM Consulting and IBM.
The thematic consensus is equally strong. Models uniformly credit IBM with technical and engineering depth, hybrid-cloud capability, watsonx assets, data governance, and suitability for regulated and legacy-heavy environments—claude-opus-5 notes competitive pricing, while gpt-5.6-sol and gpt-5.6-terra emphasize fit where IBM or Red Hat is already strategic. The recurring reason for the lower placement is perceived vendor bias toward IBM's own stack and weaker vendor-neutrality, cited by claude-opus-5, gemini-3.7-flash, gpt-5.6-luna, gpt-5.6-sol, and gpt-5.6-terra; claude-opus-5 additionally flags weaker C-suite/board-level strategy influence, and kimi-k3 points to less breadth and large-scale transformation track record than bigger firms. These themes draw on a mix of IBM's own materials—IBM watsonx, IBM Consulting AI Services, the IBM Institute for Business Value, and IBM Annual Report and earnings commentary—alongside third-party analyst sources including Gartner, Forrester, IDC (including the IDC MarketScape), Everest Group PEAK Matrix assessments, and HFS Research.
Consistently frames IBM Consulting as strong on engineering depth, hybrid-cloud, and watsonx assets with competitive pricing, but ranks it lower due to perceived bias toward IBM's own stack and weaker C-suite/board-level strategy influence.
Emphasizes deep hybrid-cloud and technical engineering strength for legacy modernization in regulated environments, while flagging potential vendor bias toward IBM's own platform ecosystem.
Presents IBM as credible for hybrid cloud, data governance, and regulated enterprise environments, but ranks it lower because its fit is strongest for clients already aligned with IBM's ecosystem rather than those seeking vendor-neutral strategy.
Positions IBM as strong for technically complex, hybrid-cloud, regulated, and legacy-heavy transformations, ranking it lower because its advantages depend on architecture aligning with IBM/Red Hat and it is seen as less vendor-neutral.
Describes IBM as credible for hybrid-cloud, data-platform, and legacy modernization needs, especially where IBM is already strategic, but ranks it fifth for being more platform-centric and less vendor-neutral.
Recognizes IBM's watsonx heritage and legacy-modernization strength, but ranks it fifth because its brand is tied to IBM's own stack and it lacks the breadth and large-scale transformation track record of bigger firms.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
ibm.com
ibm.com
ibm.com
ibm.com
ibm.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
where
where
6 AI models · 4 answers per model · ranked by score
ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0
Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.
Swipe left for the full report
Median rank each model gave each brand.
Scroll the table sideways to see every model →
| claude-opus-5 | gemini-3.7-flash | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | kimi-k3 | |
|---|---|---|---|---|---|---|
| Stanford Graduate School of Business | 1.0 | 1.0 | 4.0 | 4.0 | — | 1.0 |
| INSEAD | 3.5 | 2.0 | 1.0 | 1.0 | 1.0 | 2.0 |
| Harvard Business School | 2.0 | 3.0 | — | 4.0 | 5.0 | 3.5 |
| BYU Marriott School of Business | 5.0 | 5.0 | 1.0 | 2.0 | 1.0 | 5.0 |
| The Wharton School | 3.0 | 4.0 | 5.0 | 3.0 | — | 3.0 |
| Texas McCombs | — | — | 2.0 | — | 2.0 | 5.0 |
| Chicago Booth | 4.5 | 5.0 | — | 2.5 | 2.0 | 2.0 |
| Georgia Tech Scheller | — | — | 3.0 | 1.0 | 3.0 | — |
| Darden | — | — | 3.0 | — | 3.5 | — |
| Indiana Kelley | — | — | 4.0 | 3.0 | 5.0 | — |
| Gies College of Business | — | — | — | 1.0 | — | — |
| Massachusetts Institute of Technology | — | — | 4.0 | 5.0 | — | — |
| Rice | — | — | — | — | 4.0 | — |
| Ross | — | — | 4.0 | — | 4.0 | — |
| Questrom School of Business | — | — | — | 2.0 | — | — |
| Tuck | — | — | — | — | 3.0 | — |
| University of Florida | — | — | — | 3.0 | — | — |
| Kellogg School of Management | — | — | — | 5.0 | — | 5.0 |
| UNC Kenan-Flagler | — | — | 4.0 | — | — | — |
| University of Georgia | — | — | — | 4.0 | — | — |
| Carnegie Mellon Tepper | — | — | 5.0 | — | — | — |
| Indian Institute of Management Ahmedabad | 5.0 | — | — | — | — | — |
| UCLA Anderson | — | — | 5.0 | — | — | — |
| University of Washington | — | — | — | 5.0 | — | — |
A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.
6 of 24 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.
gpt-5.6-terra 0 → kimi-k3 100 · 1 of 6 never named it
claude-opus-5 0 → gpt-5.6-luna 80 · 3 of 6 never named it
gpt-5.6-luna 0 → claude-opus-5 75 · 1 of 6 never named it
gemini-3.7-flash 5 → gpt-5.6-luna 75
gpt-5.6-luna 25 → gemini-3.7-flash 80
gpt-5.6-terra 0 → kimi-k3 55 · 1 of 6 never named it
A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.
Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.
The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.
Stanford Graduate School of Business leads with a score of 58 of 100, ranked by 5 of 6 models.
The category splits cleanly along one fault line: whether "ROI" means prestige-and-earnings or cost-and-payback. The two readings produce almost opposite rankings, and the split runs largely between model families rather than across every brand evenly.
Stanford GSB is the clearest example. It carries a median rank of 1 across all models, but that headline hides a divide. claude-opus-5, gemini-3.7-flash, and kimi-k3 each rank it first, anchoring on the highest post-MBA compensation and Silicon Valley access. Both GPT models that ranked it — gpt-5.6-luna and gpt-5.6-sol — place it at median 4, weighting attendance cost and forgone earnings more heavily.
However, its high cost and the earnings forgone during the degree make its expected financial ROI less consistently compelling than lower-cost programs, especially for students without scholarship support. — gpt-5.6-luna
The likely driver is source mix rather than a genuine dispute over outcomes. The models ranking Stanford top lean on employment-report and ranking sources; the GPT models pair those with cost-side pages such as Stanford GSB Cost of Attendance. That said, the source lists alone do not prove causation — this is a plausible reading, not something the data confirms.
BYU Marriott shows the widest divergence in the category, median rank 4 but Q1 1 and Q3 5. The gpt-5.6 family clusters at the top (luna and terra at 1, sol at 2), treating low tuition and solid six-figure placement as decisive. claude-opus-5, gemini-3.7-flash, and kimi-k3 all put it at 5, accepting the same value logic but treating the capped earnings ceiling as disqualifying for a top slot.
BYU Marriott is one of the strongest value choices because its tuition is unusually low for a nationally recognized full-time MBA, while graduates still access consulting, technology, and finance recruiting. — gpt-5.6-luna
This is the cleanest illustration of the two ROI definitions producing opposite ranks from shared facts.
Booth mirrors the split but with the model families swapped. gpt-5.6-terra (2), kimi-k3 (2), and gpt-5.6-sol (2.5) rank it high; claude-opus-5 (4.5) and gemini-3.7-flash (5) rank it low. Here kimi-k3 sits on the payback side, citing a Forbes ROI-based ranking, while claude-opus-5 emphasizes total cost.
Forbes' ROI-based ranking put Booth at the top for five-year MBA gain, reflecting exceptional salary growth relative to cost. — kimi-k3
The Forbes ROI ranking is a plausible source behind kimi-k3's high placement, though the same model puts BYU at 5 — so its logic is not uniformly cost-first, cautioning against reading any model as applying one fixed definition everywhere.
INSEAD is the strongest point of consensus, median 2, with all three GPT models at 1 and gemini and kimi at 2. Every model credits the one-year format for halving opportunity cost. Only claude-opus-5 (3.5) qualifies it, on the compressed timeline limiting internships — a difference of emphasis, not of the core cost case.
Wharton also draws fairly tight agreement (median 3.5), with disagreement confined to cost rather than earning power. Kellogg is unanimous where ranked, both gpt-5.6-sol and kimi-k3 at 5, both citing weaker finance exposure.
HBS lands at median 3 with a real spread: claude-opus-5 at 2, gpt-5.6-terra at 5. No model disputes the compensation or alumni network; the divergence is entirely about whether total cost and two-year opportunity cost offset that.
I rank it fourth because its large total cost and two-year opportunity cost can lengthen payback, particularly for candidates who already have lucrative careers or receive limited aid. — gpt-5.6-sol
A large share of brands — Gies, Questrom, Tuck, University of Florida, Rice, UNC Kenan-Flagler, University of Georgia, Tepper, IIM Ahmedabad, UCLA Anderson, University of Washington — appear in only one model's ranking, so no agreement or disagreement can be measured. Notably, most of the affordability and online-focused picks (Gies, Questrom, Georgia Tech Scheller at rank 1 for sol) come from the GPT family, consistent with its payback-first framing. Whether that reflects the models' reasoning or just which brands each happened to recall is not something the data settles.
The most-recalled sources — Poets&Quants (25×), U.S. News, and the Financial Times Global MBA Ranking — appear across both camps, so they do not explain the divergence on their own. The distinguishing feature is that the payback-leaning rankings more often surface program-specific cost-of-attendance and tuition pages alongside employment reports, and kimi-k3 and gpt-5.6-terra reach for Forbes' ROI-based ranking. The link between those cost-side sources and the lower prestige-school placements is the most consistent pattern in the data, but it remains a hypothesis about what drove each ranking, not a demonstrated cause.
74 sources across 338 references, grouped by site from 177 recalled names. The top 5 carry 46% of them.
usnews.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named Stanford University
Stanford Graduate School of Business lands at a median rank of 1 across all models (Q1 1, Q3 1) over 16 ranked answers from 5 models, though that headline figure masks a clear split. Three models—claude-opus-5, gemini-3.7-flash, and kimi-k3—each place it at a median rank of 1, converging on a single theme: the highest post-MBA compensation of any program, paired with Silicon Valley access to venture capital, technology, private equity, and founder equity, which they argue offsets premium tuition to produce the strongest long-run payback. These models draw on employment-report data (Stanford GSB Employment Report) alongside ranking sources such as the Financial Times Global MBA Ranking, Poets&Quants, Forbes, U.S. News, and Bloomberg Businessweek. Kimi-k3 in particular anchors its case in specific compensation figures.
The disagreement comes from the two GPT models. gpt-5.6-luna places it at a median rank of 4 (Q1 4, Q3 4) over 1 ranked answer, and gpt-5.6-sol at a median rank of 4 (Q1 2.5, Q3 4.5) over 3 ranked answers. Both acknowledge the same upside but weight cost and opportunity cost more heavily, arguing that high attendance costs and uncertain near-term entrepreneurial outcomes make the measurable ROI less predictable than lower-cost programs ranked above it. Their reasoning leans on cost-side sources such as Stanford GSB Cost of Attendance and Stanford GSB Financial Aid alongside the employment reports and Financial Times ranking.
Stanford consistently posts the highest median total compensation of any MBA program, with recent classes exceeding $250,000 thanks to outsized placement in tech, venture capital, and private equity. — kimi-k3
However, its high cost and the earnings forgone during the degree make its expected financial ROI less consistently compelling than lower-cost programs, especially for students without scholarship support. — gpt-5.6-luna
Consistently emphasizes Stanford's highest median post-MBA compensation and Silicon Valley/VC/tech equity upside, arguing the salary premium and network offset its very high tuition to produce top-ranked long-run payback.
Repeatedly points to the world's highest post-graduation compensation and deep Silicon Valley ties (VC, tech, PE, founder equity), which drive long-term wealth creation and quickly offset premium tuition.
Consistently cites the highest post-MBA median total compensation (often above $230K–$250K) and top placement in tech, VC, and PE, which offset high tuition to yield a short payback and the strongest lifetime earnings uplift.
Acknowledges strong compensation potential and entrepreneurial access but stresses that high cost and forgone earnings make its financial ROI less consistently compelling than lower-cost programs.
Highlights extraordinary long-term upside via tech, VC, and entrepreneurship, while noting that high costs and uncertain near-term entrepreneurial outcomes make its measurable ROI less predictable.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
gsb.stanford.edu
rankings.ft.com
no URL recalled
gsb.stanford.edu
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
INSEAD lands at a median rank of 2 (Q1 2, Q3 2) across 13 ranked answers from 6 models, placing it consistently near the top of the field. Agreement is strong at the upper end: gpt-5.6-luna, gpt-5.6-sol, and gpt-5.6-terra all rank it first (each median 1), and gemini-3.7-flash and kimi-k3 place it second (each median 2). The main dispersion comes from claude-opus-5, whose median of 3.5 (Q1 2.75, Q3 4) reflects a more qualified view. The dominant theme across every model is the accelerated 10-month, one-year format, which cuts forgone salary and living costs roughly in half versus two-year U.S. programs and produces one of the fastest payback periods among elite schools. This is reinforced by consistent references to strong global consulting and corporate placement, particularly MBB, across Europe, Asia, and the Middle East.
The recalled sources cluster around the same ranking and placement authorities: the Financial Times Global MBA Ranking appears under nearly every model, alongside INSEAD's own Employment Statistics and Employment Report, with Poets&Quants, Forbes, Bloomberg Businessweek, and The Economist supporting the ROI and placement themes. Claude-opus-5's lower placement stems from the trade-offs it foregrounds — the compressed timeline limiting internships and access to U.S. finance recruiting — rather than any disagreement on the core cost advantage.
INSEAD's one-year format is the core of its ROI case: you pay one year of tuition and forfeit only one year of salary, cutting total cost dramatically versus two-year US programs. — kimi-k3
The 10-month format cuts opportunity cost roughly in half compared with two-year US programs while still delivering elite consulting placement (MBB hires a large share of each class). — claude-opus-5
The roughly 10-month format reduces lost salary and living costs while its international network and mobility benefit candidates in consulting, finance, or multinational careers.
The accelerated 10-month format minimizes tuition and lost earnings while providing global employer access, producing strong near-term ROI for internationally oriented candidates.
The one-year format reduces forgone salary and living costs while its global recruiting network supports post-MBA mobility, making it a strong ROI choice for international careers.
The intensive 10-month format halves opportunity cost and forgone earnings versus two-year programs, producing one of the fastest payback periods, supported by strong global consulting and corporate placement.
The one-year format is the core ROI advantage, halving total cost by sacrificing only one year of tuition and salary, with strong MBB placement across Europe, Asia, and the Middle East yielding payback periods that beat US peers.
The 10-12 month format cuts tuition and forgone salary roughly in half, driving fast payback and top FT/Forbes ROI rankings, aided by strong MBB consulting placement, though the compressed timeline limits internships and US finance recruiting access.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
insead.edu
rankings.ft.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named Harvard, Harvard University
Harvard Business School lands at a median rank of 3 (Q1 2.25, Q3 3.75) across 14 ranked answers from 5 models, placing it firmly in the upper tier but not at the summit. Model-level positioning spreads notably: claude-opus-5 is the most favorable at median rank 2 (Q1 2, Q3 2.25), gemini-3.7-flash and kimi-k3 sit at median 3 and 3.5 respectively, while gpt-5.6-sol places it at 4 and gpt-5.6-terra at 5. The common thread across all five is agreement on the underlying strengths — elite median compensation (cited near $175K base), the most powerful global alumni network, and generous need-based financial aid — recalled through sources such as the Harvard Business School Employment Data pages, the Financial Times Global MBA Ranking, U.S. News Best Business Schools, and Bloomberg Businessweek Best B-Schools.
The disagreement centers not on quality but on the arithmetic of ROI, where high total cost and two-year opportunity cost weigh against near-term payback. Models ranking HBS lower emphasize this cost drag and, in kimi-k3's case, the diluting effect of a large class size, while higher-ranking models offset the sticker price against fellowships and long-run career compounding. That tension explains why several models place it a notch behind Stanford despite comparable earnings.
HBS combines top-tier median compensation (~$175K base plus bonuses) with arguably the most powerful global alumni network, which compounds earnings over a career. — claude-opus-5
I rank it fourth because its large total cost and two-year opportunity cost can lengthen payback, particularly for candidates who already have lucrative careers or receive limited aid. — gpt-5.6-sol
HBS is consistently described as combining top-tier median compensation with the most powerful global alumni network and generous need-based fellowships that lower effective cost, while ranking slightly behind Stanford on pure ROI math due to high full price.
HBS is portrayed as pairing immense global brand prestige and a vast alumni network with generous need-based financial aid, producing top-tier compensation and strong lifetime compounding returns in fields like private equity and consulting.
HBS delivers elite compensation and arguably the strongest lifetime brand equity and network, but ranks behind peers like Stanford because its very high cost and large class size make near-term salary-to-cost payback slightly less efficient.
HBS offers extraordinary brand value and long-term career optionality, but is ranked lower because high total cost and two-year opportunity cost can lengthen payback, especially with limited aid.
HBS provides enormous long-term upside through employer access and alumni network, but high total cost and lost earnings make near-term payback less certain, making it most compelling for those who leverage the network for leadership or venture-building.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
hbs.edu
hbs.edu
hbs.edu
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named BYU Marriott, Brigham Young University, Brigham Young University Marriott School of Business
BYU Marriott School of Business lands at a median rank of 4 (Q1 1, Q3 5) across 13 ranked answers from 6 models, a spread that reflects genuine disagreement rather than consensus. The gpt-5.6 family clusters at the top—gpt-5.6-luna and gpt-5.6-terra both at median rank 1, gpt-5.6-sol at median rank 2—while claude-opus-5, gemini-3.7-flash, and kimi-k3 all settle at median rank 5. The unifying theme is that Marriott is a value or payback pick: very low tuition paired with solid six-figure placements produces a strong salary-to-cost ratio, with the reduced rate for LDS members frequently cited. The divergence in placement stems from how each model weighs that against a caveat all of them acknowledge—a narrower national and global recruiting footprint and a lower top-end earnings ceiling than elite M7 schools.
The models that rank it highest frame the low cost and access to consulting, technology, and finance recruiting as decisive, subject to fit around the faith-based environment and geographic concentration; those recalling institutional sources such as BYU Marriott MBA Tuition and the BYU Marriott MBA Employment Report sit under this theme. The models that place it fifth accept the same ROI logic but treat the capped ceiling as a reason to rank it below prestige programs, leaning on ROI-focused sources like Forbes Best Business Schools and Poets&Quants alongside U.S. News Best Business Schools.
BYU Marriott is one of the strongest value choices because its tuition is unusually low for a nationally recognized full-time MBA, while graduates still access consulting, technology, and finance recruiting. — gpt-5.6-luna
It ranks fifth here only because its ceiling for ultra-high-pa… — kimi-k3
Highlights low tuition alongside solid consulting, tech, and finance recruiting, while repeatedly noting the ROI is most compelling for students comfortable with its faith-based environment and geographic/eligibility constraints.
Stresses an unusually favorable salary-to-cost payback from low tuition and strong outcomes, while emphasizing fit—its LDS affiliation, location, and recruiting base align best with certain candidates.
Points to strong placement and compensation at substantially lower cost, especially for students eligible for the reduced tuition, with the caveat that its culture and network may not suit everyone.
Consistently frames BYU Marriott as the classic value/payback pick: very low tuition (especially for LDS members) paired with strong six-figure placements, offset by a narrower national/global recruiting footprint and lower earnings ceiling than M7 schools.
Emphasizes an exceptionally favorable salary-to-debt ratio, with very low tuition combined with competitive placements at consulting, tech, and accounting firms.
Describes it as a top ROI standout on a salary-to-debt basis due to heavily subsidized tuition and solid six-figure placements, ranked slightly lower only because its recruiting network and top-end salary ceiling are narrower than elite schools.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
marriott.byu.edu
marriott.byu.edu
no URL recalled
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named The Wharton School (University of Pennsylvania), The Wharton School, University of Pennsylvania, University of Pennsylvania, Wharton, University of Pennsylvania - Wharton, University of Pennsylvania – Wharton School
The Wharton School lands at a median rank of 3.5 (Q1 3, Q3 4) across 14 ranked answers from 5 models, placing it firmly in the upper tier without reaching the top spot. Agreement is fairly tight through the middle of the distribution: claude-opus-5, gpt-5.6-sol, and kimi-k3 all settle at a median of 3, while gemini-3.7-flash lands at 4 and gpt-5.6-luna sits lowest at 5. The consistent theme is Wharton's finance pipeline — private equity, investment banking, hedge funds, and consulting — paired with compensation the models describe as on par with HBS and Stanford. That case rests on career-outcome and ranking sources including the Wharton MBA Career Report and Career Statistics, Poets&Quants, the Financial Times Global MBA Ranking, U.S. News, Forbes, and Bloomberg Businessweek.
The disagreement is almost entirely about cost rather than earning power. The models that rank Wharton highest still flag that its full sticker tuition and Philadelphia living costs trim the ROI advantage, while gemini-3.7-flash frames those costs as rapidly amortized by signing bonuses and gpt-5.6-luna pushes it down to fifth on the grounds that the premium is easiest to justify only for those targeting the industries that reward the brand most.
Its large class and global alumni base give broad access to high-paying roles across industries. Costs are among the highest, which trims the ROI advantage relative to lower-tuition options. — claude-opus-5
I rank it fifth on value because its total cost is very high and its premium is easiest to justify for applicants targeting industries and employers that reward the brand particularly strongly. — gpt-5.6-luna
Wharton is framed as delivering top-tier finance placement (PE, banking, hedge funds) with compensation on par with HBS and Stanford, though its high full tuition slightly dilutes ROI relative to lower-cost or better-aid options.
Wharton is portrayed as offering exceptional compensation and access to finance/consulting/leadership roles plus a strong alumni network, with high tuition and costs making ROI dependable mainly for those making high-paying transitions.
Wharton is credited with elite finance/consulting compensation and strong lifetime earnings, but ranked just below Stanford and HBS because its high tuition, larger class size, and dependence on cyclical finance hiring slightly lengthen payback.
Wharton is consistently described as a dominant finance/banking/PE pipeline whose high compensation and signing bonuses rapidly amortize tuition debt, supported by its alumni network and quantitative reputation.
Wharton is seen as an elite brand for finance, consulting, and leadership pay, but ranked lower on value because its very high cost is best justified for applicants targeting industries that reward the brand most.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
usnews.com
no URL recalled
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
where
where
7 AI models · 160 answers · ranked by score
ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0
Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.
Swipe left for the full report
Median rank each model gave each brand.
Scroll the table sideways to see every model →
| claude-opus-5 | gemini-3.7-flash | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | gpt-6-astra | kimi-k3 | |
|---|---|---|---|---|---|---|---|
| Anthropic | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 |
| OpenAI | 2.0 | 2.0 | 2.0 | 2.0 | 2.0 | 2.0 | 2.0 |
| 3.0 | 4.0 | 3.0 | 3.0 | 3.0 | 3.0 | 3.0 | |
| DeepSeek | 4.0 | 3.0 | 4.0 | 4.0 | 4.0 | 4.0 | 4.0 |
| Mistral AI | 5.0 | 5.0 | 5.0 | 5.0 | 5.0 | 5.0 | — |
| Meta | 5.0 | 5.0 | — | — | — | — | 5.0 |
| GitHub | 4.0 | — | 4.0 | — | 4.0 | — | 4.0 |
| xAI | — | — | — | 5.0 | 5.0 | — | — |
| Alibaba | 5.0 | — | — | — | — | 5.0 | 5.0 |
| Qwen | — | — | — | — | — | 5.0 | — |
| Cursor | — | — | 3.0 | — | — | — | — |
| Microsoft | — | — | 4.0 | — | — | — | — |
A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.
0 of 12 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.
No brands where the models disagree
from here down, the models agreegpt-5.6-luna 13 → gemini-3.7-flash 49
gemini-3.7-flash 0 → gpt-5.6-luna 26 · 3 of 7 never named it
claude-opus-5 0 → gpt-5.6-terra 18 · 5 of 7 never named it
gpt-5.6-luna 0 → kimi-k3 16 · 4 of 7 never named it
kimi-k3 0 → gpt-5.6-luna 18 · 1 of 7 never named it
gemini-3.7-flash 46 → kimi-k3 60
A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.
Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.
The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.
Anthropic leads with a score of 100 of 100, ranked by 7 of 7 models.
The category shows a remarkably stable spine. Anthropic, OpenAI, and Google occupy ranks 1, 2, and 3 with near-total agreement across all seven models, and the disputes only begin at rank 4 and below, where open-weight providers and product-layer tools compete for the same crowded fifth slot.
Anthropic takes first place with no dissent whatsoever: a median of 1 across 160 answers, and every individual model also returns a median of 1. OpenAI mirrors this at rank 2 (median 2 for all seven models), and Google holds rank 3 with six models agreeing. The only crack in the top three is gemini-3.7-flash, which places its own maker Google at a median of 4 rather than 3 — the sole model to rate Google below the field.
The framing splits along a benchmark-versus-workflow line, though this is a difference of emphasis rather than ranking. Benchmark-anchored models (claude-opus-5, gemini-3.7-flash, kimi-k3) lean on SWE-bench, Aider, and LMArena, while the gpt-5.6 family and gpt-6-astra frame the same placements through agentic tooling and documentation. The distinction between OpenAI and Anthropic is repeatedly cast as task-dependent rather than absolute.
“They are particularly strong at algorithmic problem-solving and explaining code, though some developers find them slightly less consistent than Claude on large, multi-step engineering tasks.”
The real disagreement concerns DeepSeek. Six models place it at rank 4, but gemini-3.7-flash rates it a full rank higher at 3 — the same model that demoted Google. Its reasoning leans on benchmark parity and self-hosting appeal (SWE-bench, Hugging Face, LMSYS), whereas the six models holding DeepSeek at 4 emphasize ecosystem and enterprise-support gaps drawn from its API docs and GitHub. Whether gemini-3.7-flash's benchmark focus causes both its DeepSeek promotion and its Google demotion is a plausible reading of its source mix, not something the data confirms.
Below that, three fundamentally different kinds of brand pile up at rank 5:
The GitHub case is notable because its rank-4 placement sits above the open-weight labs at 5, yet every model attributes its ceiling to inherited capability rather than its own model.
“GitHub Copilot is the most widely deployed coding assistant, with deep IDE and pull-request integration, enterprise controls and now a model picker that routes to Anthropic, OpenAI and Google models.”
Cursor, Microsoft, and Qwen each rest on a single model's answers, so their ranks reflect one perspective rather than any consensus. Cursor (gpt-5.6-luna, rank 3) and Qwen (gpt-6-astra, rank 5) are internally consistent within their lone raters, but carry no cross-model signal. Both are worth reading as individual judgments, not category positions.
A recurring qualifier across the low-ranked open-weight entries is that the placement would rise sharply if self-hosting were the priority — gpt-6-astra says as much for both Alibaba and Qwen, which suggests the rank-5 clustering reflects an implied "typical user" default more than a capability verdict.
“I rank it fifth for a typical user seeking an immediately useful coding assistant, but it could rank considerably higher if self-hosting and control over deployment are your priorities.”
The source pool is dominated by a handful of leaderboards and benchmarks — Aider (152 times), SWE-bench (147), LMArena (79), and Artificial Analysis (48) — which explains why placements agree so tightly across models drawing on a common evidence base. Vendor documentation appears mainly for the lower-ranked and open-weight brands, consistent with those rankings being argued on deployment and ecosystem grounds rather than measured performance.
65 sources across 2145 references, grouped by site from 496 recalled names. The top 5 carry 45% of them.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Anthropic sits at the front of this category with unusual consistency: a median rank of 1 (Q1 1, Q3 1) across 160 ranked answers from all 7 models, and every one of the seven models individually returns a median rank of 1 as well. There is effectively no disagreement about placement — the divergence is only in emphasis. The benchmark-anchored models (claude-opus-5, gemini-3.7-flash, kimi-k3) lean on measured performance, citing SWE-bench Verified leadership and integration into tools like Cursor and GitHub Copilot, drawing on the SWE-bench leaderboard, Aider LLM leaderboards, and LMArena. The workflow-oriented models (gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra) frame the same ranking through repository-scale reasoning, multi-file edits, and agentic tooling such as Claude Code, sourcing Anthropic documentation and the Claude Code overview.
“Claude models (Sonnet/Opus 4.x series) are widely regarded as the strongest at real-world software engineering tasks, leading benchmarks like SWE-bench Verified and powering tools such as Claude Code, Cursor and GitHub Copilot.”
Underlying the top placement is a shared theme of large-codebase competence — long context, coordinated multi-file changes, careful instruction-following, and clear explanations rather than isolated snippet generation. Both camps land on the same conclusion from different angles, and one model even flags its own conflict of interest while retaining the ranking.
“Claude models are especially strong at understanding large codebases, making coordinated multi-file changes, and explaining code clearly.”
Consistently emphasizes that Claude (Sonnet/Opus 4.x) leads real-world coding benchmarks like SWE-bench Verified and excels at agentic, multi-file refactoring in large codebases. Repeatedly cites Claude Code and integration into tools like Cursor and GitHub Copilot as reinforcing its top position.
Repeatedly frames Anthropic's Claude models (especially Claude 3.5 Sonnet) as the industry gold standard for code generation, refactoring, and architectural reasoning. Consistently highlights large context windows, strong instruction-following, and minimal hallucinations.
Consistently recommends Anthropic for its strength in understanding large codebases, debugging, refactoring, and following detailed instructions to produce maintainable code. Emphasizes careful reasoning, clear explanations, and agentic/repository-level workflows.
Consistently positions Anthropic as the strongest overall coding recommendation, citing repository-scale reasoning, multi-file edits, and clear explanations. Emphasizes long context, agentic tooling like Claude Code, and reliable instruction-following.
Repeatedly recommends Anthropic for demanding, professional coding work, emphasizing repository-scale reasoning, code review, debugging, and following detailed engineering constraints. Highlights long-context capability and agentic workflows over lowest cost.
Consistently names Anthropic as its default choice for complex, general-purpose coding, focused on understanding existing codebases, coordinated multi-file changes, and explaining design decisions. Emphasizes Claude Code and agentic, repository-oriented workflows over snippet generation.
Consistently cites Claude models topping coding benchmarks like SWE-bench Verified and strong developer sentiment, with emphasis on multi-file reasoning and agentic tools like Claude Code. Notes precise instruction-following and integration into tools such as Cursor.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
anthropic.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
OpenAI lands at a consistent second across the field, with a median rank of 2 (Q1 2, Q3 2) over 159 ranked answers, and every one of the 7 models reports the same median of 2. The agreement is unusually tight: each model frames OpenAI as a close runner-up to Anthropic rather than a distant one. The recurring theme is a split between recognized strengths — algorithmic problem solving, debugging, reasoning-heavy tasks, and the broadest tooling ecosystem — and the reason it stops just short of first, namely that Claude is perceived as more reliable on large-repo, multi-file agentic work. claude-opus-5, kimi-k3, and gemini-3.7-flash all lean on ecosystem breadth, citing sources like the Aider LLM leaderboards, LMArena, GitHub Copilot, and SWE-bench, while the gpt-5.6 variants and gpt-6-astra emphasize Codex tooling and API maturity, backed by SWE-bench and OpenAI Codex documentation.
“It is a very close second and often better for algorithmic reasoning and competitive-programming style problems.”
The second theme is variability: several models qualify their ranking by noting that the best OpenAI choice depends on the specific model, task, or workflow, which is the main reason it settles narrowly behind Anthropic rather than tying or leading. This shows up plainly in the gpt-5.6 reasonings, where the ranking is repeatedly described as task-dependent, and in gpt-6-astra's advice to test both providers on one's own repository rather than assume a universal winner.
“They are particularly strong at algorithmic problem-solving and explaining code, though some developers find them slightly less consistent than Claude on large, multi-step engineering tasks.”
GPT-5/o-series and Codex models are praised for algorithmic problem solving, debugging, and competitive-programming tasks with the broadest tooling ecosystem, but consistently placed a very close second to Anthropic due to Claude's edge on large-repo agentic coding.
Emphasizes GPT-4o and o1 reasoning models as industry-leading for algorithmic problem solving and debugging, with unmatched ecosystem integration into tools like GitHub Copilot and Cursor, while noting it is slightly edged out on nuanced multi-file/full-stack tasks.
Describes OpenAI as a highly capable, versatile general-purpose option for code generation, debugging, and tool use, ranked slightly below Anthropic because coding quality varies by model and workflow.
Highlights strong code generation, debugging, and agentic/Codex tooling backed by a mature API ecosystem, ranking it just behind Anthropic since quality and model selection can vary by task.
Frames OpenAI as an excellent, safe general-purpose default with strong reasoning, mature APIs, and broad tooling, ranked just below Anthropic because the best choice depends on the specific environment and workflow.
Positions OpenAI as a close second, strong for code generation, debugging, tests, and Codex-supported workflows with a major ecosystem advantage, while slightly favoring Anthropic for repository-heavy editing and recommending testing both on one's own repo.
Notes GPT/o-series models excel at algorithmic and reasoning-heavy tasks with the most mature ecosystem (GitHub Copilot, API), but rank just behind Anthropic as developers find Claude more reliable on complex, multi-file software engineering.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
openai.com
platform.openai.com
openai.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named Google DeepMind
Google occupies a stable third position in the models' coding recommendations, with a median rank of 3 (Q1 3, Q3 3) across 159 ranked answers from 7 models. Six of the seven models—claude-opus-5, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra, and kimi-k3—all converge on a median of 3, while gemini-3.7-flash sits slightly lower at a median rank of 4 (Q1 3, Q3 4). The models refer to the brand as both Google and Google DeepMind. The recurring theme behind this placement is a distinctive strength in large context windows paired with a perceived gap in day-to-day coding consistency relative to the top two brands.
The dominant point of agreement is that Gemini's very large context window—cited by several models as reaching one to two million tokens—makes it well suited to reasoning across entire repositories and documentation-heavy projects, reinforced by multimodal capabilities and Google Cloud, Android, and developer-ecosystem integration. This is backed by sources including the Google DeepMind Gemini model page (cited 19 times by claude-opus-5), LMArena (24 times by kimi-k3), SWE-bench, the Aider LLM leaderboards, and Gemini API and Code Assist documentation. The consistent caveat, and the reason Google lands third rather than higher, is that its agentic and iterative coding output is seen as less consistent than Anthropic's or OpenAI's, with gemini-3.7-flash also flagging occasional shortfalls in code precision on complex edge cases.
“Gemini 2.5/3 Pro offers a huge context window that is great for reasoning across entire repositories, plus strong performance and generous free tiers via AI Studio and Gemini CLI.”
“They rank third because, while highly capable and improving rapidly, developer consensus generally places their coding output quality just behind Anthropic and OpenAI on day-to-day tasks.”
Consistently highlights Gemini 2.5/3 Pro's very large context window for repository-wide reasoning, strong benchmarks, multimodal/frontend strength, and generous free-tier access via AI Studio and Gemini CLI, while ranking it a close third due to less consistent agentic code editing than Claude or GPT.
Frames Google as a strong option for long-context, multimodal work and for developers in the Google Cloud/Android ecosystem, but consistently ranks it behind Anthropic and OpenAI due to less consistent coding quality across model versions and integrations.
Stresses Gemini's very large context windows for analyzing large repositories and documentation plus Google Cloud integration, while noting coding consistency varies across model tiers and trails the top two choices.
Highlights large context windows, multimodal inputs, and Google Cloud/ecosystem integration as strengths, positioning coding performance as competitive but variable by model version and workflow, generally below the top two for demanding agent tasks.
Consistently recommends Gemini as a strong context-heavy and multimodal option ranked third for general coding, but potentially a first choice for workflows involving large documentation, many files, or Google's developer ecosystem.
Emphasizes Gemini's exceptionally large context window for reasoning across entire repositories and competitive benchmark scores and pricing, while ranking it third because developer mindshare and day-to-day coding consistency slightly trail Anthropic and OpenAI.
Repeatedly emphasizes massive multi-million-token context windows for ingesting entire codebases and documentation, plus multimodal and ecosystem integration, while noting occasional shortfalls in code precision or idiomatic output on complex edge cases relative to top competitors.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
cloud.google.com
ai.google.dev
ai.google.dev
ai.google.dev
deepmind.google
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
DeepSeek settles at a median rank of 4 across all seven models (Q1 4, Q3 4) over 141 ranked answers, marking it as a consistent mid-table recommendation rather than a top pick. Agreement is notably tight: six of the seven models place it at a median of 4, with only gemini-3.7-flash landing higher at a median of 3 (Q1 3, Q3 3). The shared narrative pairs a strong upside — near-frontier coding quality at a fraction of the cost, with open weights for self-hosting — against a recurring limitation around less mature tooling, enterprise support, and reliability on complex agentic tasks. This value-versus-maturity framing runs through claude-opus-5, kimi-k3, and the GPT-family models alike, drawing heavily on sources such as the Aider LLM leaderboards, DeepSeek's API documentation and GitHub repositories, Hugging Face model cards, Artificial Analysis, and SWE-bench.
The cost-and-openness theme is where models converge most tightly, typically citing the V3 and R1 lineage as the anchor for their reasoning.
“DeepSeek's V3 and R1 models deliver coding performance that rivals much more expensive proprietary models, making them exceptional value and a leading open-weight option.”
The offsetting theme — why it rarely climbs above rank 4 — centers on deployment effort, governance, and ecosystem maturity, drawn from the API documentation and GitHub-heavy source pool that the GPT models lean on.
“I rank it fourth because its surrounding tools, enterprise support, and overall developer ecosystem are less mature than those of the leading providers.”
Gemini's higher placement reflects a heavier emphasis on benchmark parity and self-hosting appeal, supported by references to the SWE-bench Leaderboard, Hugging Face, and the LMSYS Chatbot Arena.
Repeatedly presents DeepSeek as an open-weights powerhouse rivaling proprietary models on coding benchmarks at a fraction of the cost, ideal for self-hosting, with a less mature developer tooling ecosystem as its main limitation.
Consistently frames DeepSeek's V3/R1-lineage models as delivering near-frontier coding performance at a fraction of the cost with open weights for self-hosting, while trailing the top proprietary labs on complex agentic tasks and tooling polish.
Emphasizes DeepSeek's capable coding performance and cost efficiency with open-weight availability, ranking it below top providers due to less predictable reliability, tooling consistency, and deployment/ecosystem maturity.
Consistently describes DeepSeek as offering capable coding and reasoning at competitive prices with open-weight/self-hosting options, ranking it lower because support, governance, privacy, and ecosystem maturity require more evaluation.
Frames DeepSeek as a value-oriented option strong on cost efficiency, open weights, and self-hosting flexibility, placing it below leading providers because deployment, support, safety controls, and production tooling need more evaluation.
Positions DeepSeek as attractive for cost efficiency and access to open model weights, ranking it below the leading integrated assistants because deployment choices, self-hosting, and workflow tooling require more setup effort.
Consistently highlights DeepSeek's V3 and R1 models delivering near-frontier coding at very low cost with open weights for self-hosting, ranking it below the top labs due to less mature tooling, reliability, and consistency on complex tasks.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
huggingface.co
github.com
github.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named Mistral
Mistral AI settles at a median rank of 5 (Q1 5, Q3 5) across 57 ranked answers from all 6 models, and the tight quartiles signal strong agreement: every model places it fifth, with only gemini-3.7-flash showing marginally more spread (Q1 4, Q3 5). The consensus is less about weakness than positioning. Models frame Mistral as a specialist and deployment-oriented pick rather than a top capability choice, repeatedly citing its coding-focused models Codestral and Devstral. claude-opus-5 and gemini-3.7-flash lean on the technical strengths — fast, cheap, open-weight models tuned for autocomplete and fill-in-the-middle — while noting they trail frontier models on hard multi-file work. These themes rest on sources such as the Hugging Face Devstral model card, Mistral AI docs, the Codestral announcement, and Artificial Analysis.
“Codestral and Devstral are purpose-built coding models with permissive/open weights and excellent latency for fill-in-the-middle autocomplete, plus EU data-residency appeal.”
The second recurring theme, most prominent in the gpt-5.6 family, is that Mistral appeals when openness, deployment flexibility, and European hosting matter more than maximizing raw coding performance. gpt-5.6-luna, gpt-5.6-sol, and gpt-5.6-terra consistently rank it fifth on the grounds of a less established ecosystem and weaker demonstrated results on demanding repository-scale tasks, drawing on Mistral AI documentation alongside benchmark references like Aider LLM Leaderboards and SWE-bench. gpt-6-astra echoes the completion-workflow angle, treating it as a targeted rather than general-purpose fit.
“Mistral is a good choice when openness, deployment flexibility, and control over infrastructure matter more than achieving the strongest possible coding results.”
Consistently highlights Codestral and Devstral as fast, cheap, open-weight models good for autocomplete, fill-in-the-middle, and EU/on-prem deployment, while noting they lag frontier models on hard multi-file tasks, making Mistral a privacy/cost choice rather than a top capability pick.
Emphasizes Codestral as an efficient, lightweight model optimized for low-latency code completion, fill-in-the-middle, and broad language support, ideal for IDE autocompletion and self-hosting, but ranked fifth because it trails frontier models on complex multi-step reasoning.
Positions Mistral as appealing for open-weight options, deployment flexibility, and European hosting, but ranks it fifth because its coding ecosystem, tooling, and performance on complex multi-file tasks are less consistently strong than leading providers.
Frames Mistral as a good option for deployment flexibility, European hosting, and open-weight/Codestral coding models, but ranks it below others due to a less established ecosystem and weaker demonstrated performance on demanding repository-scale tasks.
Recommends Mistral when European hosting, deployment flexibility, or open-weight options matter, while generally preferring higher-ranked providers first for the hardest end-to-end software-engineering tasks.
Values Mistral for specialized code completion and fill-in-the-middle workflows with deployment flexibility, ranking it fifth for broad complex coding assistance while noting it can fit targeted completion-oriented workloads (and advises checking licensing).
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
mistral.ai
docs.mistral.ai
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
where
where
9 AI models · 42 answers · ranked by score
ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0
Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.
Swipe left for the full report
Median rank each model gave each brand.
Scroll the table sideways to see every model →
| claude-opus-5 | deepseek-v4.1-flash | gemini-3.7-flash | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | gpt-6-astra | kimi-k3 | mistral-medium-3.5 | |
|---|---|---|---|---|---|---|---|---|---|
| Acquired | 2.0 | 1.0 | 2.0 | 1.0 | 1.0 | 2.0 | 3.0 | 1.0 | — |
| How I Built This | 1.0 | 2.0 | 1.0 | 4.0 | 3.0 | 1.0 | 2.0 | 2.0 | 2.0 |
| Masters of Scale | 4.0 | 3.0 | 3.5 | 2.5 | 4.0 | 3.0 | 5.0 | 4.0 | 3.0 |
| HBR IdeaCast | 4.0 | 5.0 | 2.0 | 2.5 | 2.0 | 5.0 | 1.0 | 5.0 | 5.0 |
| The Tim Ferriss Show | 5.0 | 4.0 | — | — | — | 4.0 | — | 4.5 | 1.0 |
| Planet Money | 3.0 | 4.0 | 3.0 | 5.0 | — | 3.5 | 4.0 | — | — |
| How I Built This with Guy Raz | — | 1.0 | — | — | — | — | — | — | 2.0 |
| Harvard Business Review | — | — | 1.5 | 2.5 | — | — | — | — | — |
| Masters of Scale with Reid Hoffman | — | 2.0 | — | — | — | — | — | — | 3.0 |
| The Diary of a CEO | 4.5 | 5.0 | 4.0 | 5.0 | 5.0 | 4.0 | — | 3.0 | — |
| NPR | — | — | 1.5 | 3.0 | — | — | — | — | — |
| The GaryVee Audio Experience | — | 5.0 | — | — | — | — | — | — | 4.0 |
| The Indicator from Planet Money | 3.0 | — | — | — | — | — | 4.0 | — | — |
| The Wall Street Journal | — | — | 3.5 | — | — | — | — | — | — |
| My First Million | — | — | — | 5.0 | — | — | — | 3.0 | — |
| Wondery | — | — | 4.0 | — | — | — | — | — | — |
| Bloomberg | — | — | 4.5 | — | — | — | — | — | — |
| The Journal | — | 3.0 | — | — | — | — | — | — | — |
| Pivot | — | — | — | — | 5.0 | — | — | — | — |
| The $100 MBA | — | — | — | — | — | — | — | — | 5.0 |
| HubSpot | — | — | — | 5.0 | — | — | — | — | — |
| The Journal. | — | — | — | — | — | 5.0 | — | — | — |
| The Prof G Pod | — | — | — | — | 5.0 | — | — | — | — |
| Business Wars | — | — | — | — | — | — | — | 5.0 | — |
| Freakonomics Radio | — | — | 5.0 | — | — | — | — | — | — |
| Masters in Business | — | — | 5.0 | — | — | — | — | — | — |
| Smart Passive Income | — | — | — | — | — | — | — | — | 5.0 |
| The Indicator | — | 5.0 | — | — | — | — | — | — | — |
A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.
5 of 28 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.
mistral-medium-3.5 0 → gpt-5.6-luna 100 · 1 of 9 never named it
gemini-3.7-flash 0 → mistral-medium-3.5 100 · 4 of 9 never named it
deepseek-v4.1-flash 4 → gpt-5.6-sol 85
gpt-5.6-luna 30 → claude-opus-5 100
claude-opus-5 0 → mistral-medium-3.5 48 · 7 of 9 never named it
claude-opus-5 8 → gpt-5.6-luna 65
A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.
Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.
The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.
Acquired leads with a score of 68 of 100, ranked by 8 of 9 models.
The business podcast category divides into a small tier of broadly recognised shows and a long tail of brands that surface in a single model's list. Two titles anchor the top: Acquired (Borda 67.56, ranked by 8 of 9) and How I Built This (66.22, ranked by all 9). Both draw near-universal recognition, and both split the panel not on quality but on placement.
The sharpest fault line runs through Acquired. Models agree on its deeply researched, long-form company histories, but disagree on whether episode length keeps it out of the top slot.
Acquired · 100 points apart on a 0–100 scale
How I Built This divides along a parallel axis — narrative accessibility versus analytical depth. claude-opus-5 placed it first in every run for a perfect 100; gpt-5.6-luna settled it at rank 4 for a Borda of 30, judging it better for inspiration than rigour. This mirrors Acquired's pattern: the models converge on what a show is and diverge on what to weigh it against.
That same tension governs the middle tier. HBR IdeaCast (34.44) is the clearest case: gpt-5.6-sol and gpt-6-astra treat its research-grounded management substance as the strongest all-around pick (Borda 85 and 84), while deepseek-v4.1-flash, mistral-medium-3.5 and gpt-5.6-terra rank it near last (4, 8, 10), citing an academic tone and short format. The disagreement is entirely about weighting rigour against listenability, not about what the show delivers.
per-model scores 4–85 of 100 · mean 34 across 9 models
Two brands are pushed high by a single dissenting model. The Tim Ferriss Show (18.89) owes its standing almost entirely to mistral-medium-3.5, which ranked it first in all five runs; the other four models that named it placed it fourth or fifth on grounds of topical fit — whether a show ranging across health and self-optimisation counts as business at all. mistral-medium-3.5 also lifts The GaryVee Audio Experience and, uniquely, ranks a cluster of solo titles (Smart Passive Income, The $100 MBA, Masters of Scale with Reid Hoffman), suggesting a broader, entrepreneur-facing frame than its peers.
Where the panel does agree without caveat is at the bottom. Masters of Scale (34.67) and Planet Money (18.44) both draw the "broad agreement" label — every model that ranks them places them mid-table, discounting scaling-specific or news-explainer scope against direct operator guidance. Below that, a long tail of single-model brands (Pivot, Bloomberg, Freakonomics Radio, Masters in Business, Business Wars, The Journal, The Prof G Pod) carries the "models agree" label, but this is agreement by absence rather than conviction.
A recurring artefact worth flagging: several brands appear twice under near-duplicate names — How I Built This and How I Built This with Guy Raz, Masters of Scale and …with Reid Hoffman, The Indicator and The Indicator from Planet Money, The Journal and The Journal. The split versions each sit low precisely because coverage fragments across the variants. This is a naming effect in the recalled data, not evidence that models see these as distinct shows.
The sourcing is strikingly uniform across the whole category and does little to explain the splits. Aggregators dominate — Apple Podcasts×50 and Spotify×31 lead by a wide margin — followed by each show's own site and publisher pages. Because nearly every model reaches for the same directory and official-page material, the recalled sources plausibly explain what the models know about each show, but not why they rank the same show so differently. That divergence tracks editorial judgement — rigour versus accessibility, focus versus breadth — more than any evidence base. The one exception where a source hints at a view is claude-opus-5 citing BBC News reporting on health claims behind its cautious read of The Diary of a CEO; that link is suggestive, not established by the data.
55 sources across 473 references, grouped by site from 141 recalled names. The top 5 carry 68% of them.
podcasts.apple.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Acquired scores a Borda average of 67.56 of 100, ranked by 8 of the 9 models, with a first-choice score of 55.19. The panel converges on what the show offers — deeply researched, long-form company histories that function as strategy case studies — but diverges sharply on where to place it, a split the report labels "sharply split" and reflected in a standard deviation of 33.51. That range runs from gpt-5.6-luna's borda of 100, ranking it first in all four runs, down to gemini-3.7-flash's 16 off a single 20% coverage run, with claude-opus-5, gpt-5.6-terra and gpt-6-astra clustering in the middle at second or third.
Acquired · 100 points apart on a 0–100 scale
The recurring theme behind both the praise and the hesitation is episode length: models value the multi-hour depth for founders, investors and strategy-minded listeners while flagging the time commitment as the reason to rank it just below a more accessible pick. The sourcing sits heavily on the brand's own material — Acquiredwww.acquired.fm×19 recurs across most models — supplemented by Apple Podcastspodcasts.apple.com/us/podcast/acquired/id1050462261×10, with kimi-k3 and claude-opus-5 also reaching to press and reference coverage.
“Its episodes function almost like mini-MBA case studies, and its production quality and host chemistry are consistently praised.”
“The depth is a major strength, but its often very long episodes make it a less convenient default recommendation for a general listener.”
Repeatedly highlights deeply researched, long-form company narratives that serve as strategy case studies, valuable for understanding business models and moats, with episode length noted as a limitation for casual listeners.
Consistently frames it as offering exceptionally detailed, well-researched long-form analysis of company strategy and finance, valuable for those seeking rigorous insight over quick motivational content.
Repeatedly stresses exceptionally detailed, well-researched examinations of companies and strategic decisions, with unusually long episodes offset by depth for serious business students.
Consistently describes exceptionally deep, multi-hour, well-researched company breakdowns functioning like mini-MBA case studies, praised for substance though demanding in length for casual listeners.
Consistently emphasizes hosts Ben Gilbert and David Rosenthal's deeply researched, multi-hour company histories and unmatched strategic depth, noting episode length as a drawback and strong credibility among investors and operators.
Points to exhaustive, multi-hour deep dives into company history, strategy, and financials by Gilbert and Rosenthal, positioning it as essential for investors and strategists.
Consistently emphasizes rigorous, long-form analysis of companies, strategy, and market structure, rewarding listeners who want durable insight despite the larger time commitment.
Repeatedly praises detailed company histories and analysis of competitive advantages and strategy, while noting that unusually long episodes make it a less convenient default recommendation.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
How I Built This posts a Borda score of
Stat unavailableand was ranked by all 9 of the 9 models, giving it broad reach even as opinions on its standing diverge. The per-model Borda spans 30 to 100 with a standard deviation of 24.83, which the report labels "models split," and that gap is real: claude-opus-5 placed it first in all five of its runs for a perfect 100, while gpt-5.6-luna settled it at a median rank 4 for a Borda of 30.
How I Built This · 70 points apart on a 0–100 scale
The disagreement traces to a single recurring theme — accessibility and narrative craft versus analytical depth. Models that reward Guy Raz's NPR-produced founder storytelling rank it at or near the top, while those weighing rigorous strategy analysis push it down, several explicitly seating it just below more study-focused shows like HBR IdeaCast or Acquired. The recalled sources cluster tightly under the show's own footprint: NPR appears across nearly every model, alongside Apple Podcasts, Wondery, Spotify, and Wikipedia. A first-choice score of 53.93 confirms it is frequently named near the top yet not uniformly led.
“It's the safest all-around recommendation for someone wanting an engaging business podcast.”
“It is better for entrepreneurial inspiration than for rigorous analysis of markets, operations, or competitive strategy.”
Consistently emphasizes Guy Raz's NPR-produced, narrative-driven founder interviews with high production quality, a large back catalog of major companies, and broad accessibility, making it the safest all-around recommendation.
Focuses on Guy Raz's narrative-driven, deep-dive interviews with founders of world-renowned companies, stressing storytelling depth, resilience, and universal accessibility.
Presents it as a strong all-around choice with accessible, non-technical founder interviews and clear lessons, while noting it is less analytical than more study-focused shows.
Highlights Guy Raz's candid, well-produced founder interviews that balance emotional origin stories with practical entrepreneurship lessons, praising broad accessibility.
Consistently rates it as an accessible, engaging entry point for entrepreneurship, but caveats that retrospective success stories are inspirational rather than a reliable, actionable playbook, ranking it just below HBR IdeaCast.
Emphasizes Guy Raz's engaging, well-produced founder interviews revealing the messy reality of company-building, calling it the most accessible/universal choice though more inspirational than analytical.
Repeatedly cites NPR's inspiring entrepreneur stories with an engaging narrative style that is both educational and useful for aspiring business leaders.
Describes it as making entrepreneurship accessible through polished founder storytelling, strong on inspiration but offering less analytical depth than top-ranked choices.
Consistently frames it as engaging and accessible for founder stories and inspiration, while noting it offers less rigorous analysis of strategy, markets, and operations.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
podcasts.apple.com
podcasts.apple.com
no URL recalled
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
All nine models ranked Masters of Scale, giving it a Borda score of
Stat unavailableand a first-choice score of 21.39 of 100. No model placed it first in any run, and the per-model Borda spread runs from 8 to 65 with a standard deviation of 17.64 — a range the report labels "broad agreement." That spread tracks a consistent split in placement rather than in sentiment: gpt-5.6-luna sits it at a median rank of 2.5 while gpt-6-astra parks it at a median rank of 5 across all five of its runs.
Masters of Scale · 57 points apart on a 0–100 scale
The disagreement is one of degree rather than direction. Models converge on the same picture — Reid Hoffman's polished, thesis-driven interviews with prominent founders and executives, strong on growth and scaling themes — and then diverge on how much that framing counts against it. Several note the format can feel more inspirational, promotional, or theoretical than analytically rigorous, and gpt-6-astra's fifth-place placements rest specifically on its high-growth focus being less applicable to a general audience. Recalled sources cluster tightly around the show's own footprint: Masters of Scalemastersofscale.com×21 recurs across models, supported by directory and platform listings such as Apple Podcasts×50 and producer references to WaitWhat.
“It balances inspiring stories with practical lessons about leadership, scaling, and decision-making.”
“It ranks fifth for a general audience because its emphasis on scaling high-growth companies is less applicable to many everyday business situations.”
Highlights thoughtful founder and executive interviews organized around clear growth principles, with polished production, while noting the inspirational framing can feel more polished than operational or analytical.
Consistently frames it as valuable for scaling and leadership insights from prominent operators and founders, but notes some episodes feel promotional or more interview-driven than analytically rigorous.
Describes polished, thematic conversations with recognizable operators on growth and scaling, useful for leadership insights but less analytically deep or granular than company-specific case-study shows.
Frames it as Reid Hoffman exploring how companies grow from zero to a billion with insights from successful founders, while noting it can feel more theoretical than other shows.
Focuses on Reid Hoffman deconstructing unconventional strategies for growing startups into large enterprises, praising high production and access to top executives as providing practical frameworks.
Emphasizes Reid Hoffman's scaling framework and strong, high-profile guest list with polished production, while noting the scripted, thesis-driven format can feel less authentic.
Points to prominent founders and strong production centered on growth and organizational challenges, but notes the content is thesis-driven and more inspirational and celebratory than critically probing or analytical.
Emphasizes Reid Hoffman's LinkedIn credibility, strong guests, and polished production, but notes the scripted, thesis-driven format can feel promotional of his portfolio and less candid than long-form interviews.
Recommends it for startup growth, leadership, and founder perspectives, but consistently ranks it fifth because its focus on high-growth companies is less applicable to a general or everyday-business audience.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
podcasts.apple.com
no URL recalled
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
HBR IdeaCast lands in the middle of the pack with a Borda score of
Stat unavailableand a first-choice score of 25.93, ranked by all 9 of 9 models but with widely divergent enthusiasm. The panel divides along a consistent fault line: models converge on its credibility and research-grounded management substance, but split sharply on how much that matters against entertainment value. gpt-5.6-sol (borda 85) and gpt-6-astra (borda 84) treat it as the strongest all-around pick for practical management learning, while deepseek-v4.1-flash (borda 4), mistral-medium-3.5 (borda 8) and gpt-5.6-terra (borda 10) rank it near last, citing an academic tone and short, topic-led format that feels less immersive than founder-story shows.
HBR IdeaCast · 81 points apart on a 0–100 scale
The recalled sources cluster tightly around the show's own home and directory listings — the Harvard Business Review — HBR IdeaCasthbr.org/podcasts/ideacast×3 page and Apple Podcastspodcasts.apple.com/us/podcast/hbr-ideacast/id152022135×8 listing recur across most models — which underlines that the disagreement is not about what the podcast is but about how to weigh rigor against listenability.
per-model scores 4–85 of 100 · mean 34 across 9 models
“HBR IdeaCast is the strongest all-around recommendation because it delivers concise, research-informed conversations on leadership, strategy, innovation, and management.”
“It ranks last here mainly because its academic tone and shorter format make it less entertaining and less story-driven than the others, even though the substance is solid.”
Repeatedly names it the strongest all-around recommendation for practical leadership and management learning, especially for people managers, while ranking it below story-driven shows for entertainment.
Emphasizes its Harvard Business Review pedigree, academic rigor, and research-backed, concise episodes offering actionable management insights for corporate professionals.
Consistently rates it a top all-around pick for translating management research into concise, practical conversations with broad usefulness, while noting less depth per topic than narrative shows.
Describes it as a credible, evidence-informed source on management and leadership with expert guests and concise episodes, though less entertaining and variable in quality by topic.
Consistently frames it as a credible, research-grounded management interview show with concise episodes, ideal for managers and professionals, while noting its dry, interview-heavy format lowers its entertainment value.
Views it as credible and research-backed but ranks it lowest for general listeners due to its academic tone and shorter, less engaging format.
Sees it as a dependable, research-informed source on management and leadership, but ranks it lower because its short, topic-led format feels less immersive than founder-story shows.
Regards it as substantively solid and research-grounded management content, but consistently ranks it near last because its academic tone and brevity make it less entertaining and narrative-driven.
Characterizes it as well-researched and credible HBR content that can feel more academic and less actionable or engaging than other options.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
podcasts.apple.com
no URL recalled
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
The Tim Ferriss Show scores a Borda of
Stat unavailableacross the nine models, ranked by five of them but with a standard deviation the report labels "sharply split." That divergence turns on one recurring question: whether the show counts as a business podcast at all. mistral-medium-3.5 ranked it first in all five of its runs, praising the in-depth interviews and actionable insights, while claude-opus-5, kimi-k3, gpt-5.6-terra and deepseek-v4.1-flash all placed it lower on grounds of topical fit rather than quality — the show, in their reading, drifts into health, psychedelics, lifestyle and self-optimization, and the long episode formats dilute the business focus.
per-model scores 0–100 of 100 · mean 19 across 9 models
The sources sit under two themes that mirror this split: the platform and canonical listings that establish reach — Apple Podcasts×50 and Spotify×31 — and the show's own presence via Tim Ferriss' blog and podcast pages, cited across mistral-medium-3.5, kimi-k3 and claude-opus-5. The agreement is on influence and guest quality; the disagreement is entirely on whether that value maps to business specifically.
“Tim Ferriss consistently delivers high-quality, in-depth interviews with world-class performers across various fields, offering actionable insights.”
“Ranked last here purely on topical fit rather than quality.”
Consistently praises Ferriss's high-quality, in-depth interviews with top performers and the actionable insights across diverse topics, presenting it as highly valuable for business and personal growth.
Emphasizes that Ferriss extracts tactics and routines from top performers, but that the long, wide-ranging episodes dilute the business focus, lowering its ranking for pure business listeners.
Highlights deep conversations with founders, investors, and thinkers yielding useful mental models, while noting its scope extends beyond business, making it less focused for purely business-oriented listeners.
Repeatedly credits Ferriss with pioneering the long-form format and landing world-class guests with practical tactics, but ranks it lower because it drifts into lifestyle and self-optimization, with variable episode quality and very long formats.
Consistently frames it as a pioneering long-form interview show with an enormous archive of high-profile guests and tactical insights, but notes its scope drifts beyond business into health, psychedelics, and self-improvement, making it less focused and best browsed selectively.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
no URL recalled
podcasts.apple.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
where
where
4 AI models · 5 answers per model · ranked by score
ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0
Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.
Swipe left for the full report
Median rank each model gave each brand.
| claude-opus-5 | gemini-3.7-flash | gpt-6-astra | kimi-k3 | |
|---|---|---|---|---|
| Qatar Airways | 1.0 | 1.0 | 1.0 | 1.0 |
| Singapore Airlines | 2.0 | 2.0 | 2.0 | 2.0 |
| ANA | 4.0 | 3.0 | 3.0 | 3.0 |
| Emirates | 3.0 | 4.0 | 5.0 | 4.0 |
| Cathay Pacific | — | 5.0 | 4.0 | 5.0 |
| Delta Air Lines | 5.0 | 5.0 | — | — |
| Japan Airlines | 4.0 | — | — | — |
| EVA Air | — | — | 5.0 | — |
| Turkish Airlines | — | — | — | 5.0 |
A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.
0 of 9 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.
No brands where the models disagree
from here down, the models agreegpt-6-astra 12 → claude-opus-5 60
claude-opus-5 0 → gpt-6-astra 40 · 1 of 4 never named it
claude-opus-5 24 → gpt-6-astra 60
gpt-6-astra 0 → claude-opus-5 20 · 2 of 4 never named it
A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.
Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.
The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.
Qatar Airways leads with a score of 100 of 100, ranked by 4 of 4 models.
The category shows strong consensus at the top and widening disagreement further down. All four models place Qatar Airways at a median rank of 1 and Singapore Airlines at 2, with no model dissenting on either. The splits emerge among the middle and lower tiers, and much of the divergence is about how far a carrier trails rather than the direction of the judgment.
Where the models agree
Qatar and Singapore are locked. Every model lands on 1 and 2 respectively, and the reasoning converges too: the Qatar Qsuite is treated as the benchmark cabin, and Singapore's second place is consistently attributed to seat privacy and diagonal sleeping position rather than any soft-product weakness. The recalled sources reinforce the agreement — Skytrax World Airline Awards, The Points Guy, and One Mile at a Time appear across all four models for both carriers.
The one wrinkle at the top is framing rather than placement: gpt-6-astra shares Qatar's first rank but qualifies it, flagging that the Qsuite is not fleet-wide.
Qsuite is not available on every aircraft, so confirm the scheduled seat configuration before booking.
— gpt-6-astra
Where the models split
The clearest divergence is on Emirates. claude-opus-5 rates it highest at a median of 3, gemini-3.7-flash and kimi-k3 sit at 4, and gpt-6-astra is the low outlier at 5. All four cite the same tension — strong A380 and ground experience against the older 2-3-2 777 seating — but weigh the seating variability very differently. gpt-6-astra lets that flaw set the ceiling, while claude-opus-5 treats the new 777 suites as a mitigating factor.
It ranks fifth among these options because the seating experience varies substantially: some older Boeing 777 cabins lack direct aisle access for every passenger and offer less privacy than the leading alternatives.
— gpt-6-astra
ANA shows a milder split: three models place it at 3, but claude-opus-5 alone sits at 4. Notably, claude-opus-5 rates Emirates above ANA (3 vs 4), while gpt-6-astra reverses that ordering (ANA 3, Emirates 5). The underlying complaint for ANA is uniform — The Room is exceptional but flight-dependent — so the disagreement is about how heavily to penalize limited availability versus Emirates' seat inconsistency.
Cathay Pacific splits by degree among the three models that ranked it: gpt-6-astra at 4, gemini-3.7-flash and kimi-k3 at 5. kimi-k3 pushes it lowest on post-pandemic service and catering recovery, a concern the other models do not foreground.
Single-model tails
Several carriers were ranked by only one model, so they carry no cross-model signal:
These reflect differences in which brands each model chose to recall as much as differences in evaluation. That claude-opus-5 surfaced Japan Airlines while kimi-k3 surfaced Turkish Airlines, for instance, may be a coverage difference rather than a disagreement about relative quality — but the data does not let us test that, since the other models did not rank these carriers at all.
On the sources
The recalled source base is highly uniform across the category: Skytrax World Airline Awards (57×), The Points Guy (45×), and One Mile at a Time (39×) dominate every brand. This makes it difficult to attribute any specific model's split to a distinctive source. One plausible pattern — and it is a hypothesis, not something the data confirms — is that the US-inflected placements lean on region-specific sources: Delta's ranking draws on J.D. Power North America Airline Satisfaction Study, which appears nowhere else in the category. Whether that source drives claude-opus-5 and gemini-3.7-flash to surface Delta at all, versus merely accompanying a judgment reached otherwise, cannot be determined from these numbers.
35 sources across 293 references, grouped by site from 53 recalled names. The top 5 carry 77% of them.
worldairlineawards.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Qatar Airways occupies an unambiguous top position in this category, holding a median rank of 1 (Q1 1, Q3 1) across all 20 ranked answers from 4 models, with each individual model also landing on a median rank of 1. Agreement is complete: no model placed it below first, and the reasoning behind that consensus is consistent, centering almost entirely on the Qsuite business-class product. Across claude-opus-5, gemini-3.7-flash, and kimi-k3, the Qsuite is repeatedly described as the benchmark for the cabin — fully enclosed suites with doors, double-bed and quad configurations, and dine-on-demand service — supplemented by the Al Mourjan lounge at Doha's Hamad International and repeated Skytrax World's Best Business Class wins. The recalled sources under these themes converge as well, with Skytrax World Airline Awards, The Points Guy, and One Mile at a Time appearing across all four models, alongside Business Traveller and Qatar's own site.
Its Qsuite business class offers fully enclosed suites with doors and quad configurations, widely regarded as the best business class product in the world.
— claude-opus-5
The one point of nuance comes from gpt-6-astra, which shares the top ranking but qualifies its endorsement by flagging that the flagship product is not universally deployed, advising travelers to verify the aircraft and seat map before booking.
Qsuite is not available on every aircraft, so confirm the scheduled seat configuration before booking.
— gpt-6-astra
Consistently centers on the Qsuite as the benchmark business-class product (enclosed suites with doors, quad configurations), alongside top-rated dining, the Doha/Al Mourjan lounge, global connectivity, and repeated Skytrax World's Best Business Class wins.
Repeatedly frames Qatar Airways as setting the industry benchmark via its Qsuite (privacy doors, double-bed configurations), combined with dine-on-demand fine dining and world-class lounge/ground service at Doha's Hamad International.
Recommends Qatar Airways as the top choice for long-haul business class based on Qsuite's privacy, flexible layouts, and dining, while consistently cautioning that Qsuite is not on every aircraft and advising travelers to check the seat map.
Describes the Qsuite as the benchmark for business class (enclosed suites, doors, double-bed/quad options, dine-on-demand), reinforced by consistent fleet-wide quality, the Doha hub/lounge, and Skytrax awards.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
worldairlineawards.com
worldairlineawards.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Singapore Airlines occupies a remarkably stable position across all four models, holding a median rank of 2 (Q1 2, Q3 2) over 20 ranked answers. Every model — claude-opus-5, gemini-3.7-flash, gpt-6-astra, and kimi-k3 — landed on the same median of 2, indicating broad agreement that the carrier is a near-top choice but not the single strongest recommendation. The shared themes are consistent: polished, world-renowned cabin service, exceptionally wide business-class seats, and the Book the Cook dining program, with several models also crediting the Changi hub experience. These recurring praises draw on sources such as the Skytrax World Airline Awards, One Mile at a Time, The Points Guy, and Condé Nast Traveler.
The disagreement is less about placement than about the reason it sits at second rather than first. Both gpt-6-astra and kimi-k3 explicitly frame it as trailing Qatar's Qsuite, citing seat privacy and sleeping configuration as the deciding factors, while claude-opus-5 and kimi-k3 emphasize the manual flip and diagonal sleeping position of the seat as its recurring drawback. Those hardware caveats, drawn from reviewer sources including One Mile at a Time and The Points Guy, temper otherwise strong soft-product assessments.
Extremely wide business class seats, excellent Book the Cook dining and famously polished cabin crew make it a benchmark carrier.
— claude-opus-5
It loses the top spot mainly because most of its business seats lack privacy doors and require sleeping diagonally, making it slightly less private than Qatar's Qsuite.
— kimi-k3
Consistently praised for polished cabin service, very wide business seats, and Book the Cook dining, while repeatedly noting the drawback that some seats require manual flipping to lie flat and older configurations lack direct aisle access.
Emphasizes world-renowned cabin crew hospitality, wide lie-flat seating, the Book the Cook dining program, and consistency, along with the Changi hub experience.
Regards it as an excellent all-round choice for polished service, dining, and spacious seats, but consistently ranks it just behind Qatar because seat privacy and sleeping comfort vary by aircraft and trail the Qsuite.
Highlights exceptionally wide seats, polished service, and Book the Cook dining, while noting that seats require sleeping diagonally and often lack privacy doors, keeping it slightly behind Qatar's Qsuite.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
worldairlineawards.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named ANA (All Nippon Airways), All Nippon Airways, ANA All Nippon Airways
ANA lands consistently in the upper-middle of the field, with a median rank of 3 across 18 ranked answers from 4 models (Q1 3, Q3 4). Three of the four models — gemini-3.7-flash, gpt-6-astra, and kimi-k3 — each place it at a median of 3, while claude-opus-5 sits slightly lower at a median of 4 (Q1 4, Q3 4). The models agree closely on the underlying reasoning: the standout attraction is The Room, described as one of the widest, door-equipped business suites in commercial aviation, paired with attentive Japanese hospitality and strong dining. The recurring caveat that keeps ANA from ranking higher is limited availability — the best seat appears only on select aircraft and routes, making the experience flight-dependent — alongside a network described as more Japan- or Asia-centric.
That tension between an exceptional hard product and inconsistent access runs through every model's assessment. gpt-6-astra ties its third-place ranking directly to that limitation, while kimi-k3 and claude-opus-5 point to both fleet inconsistency and a thinner global network. The recalled sources cluster around a small set of aviation-review outlets, most frequently Skytrax World Airline Awards, One Mile at a Time, and The Points Guy, which appear under both the praise for The Room and the notes on route availability.
ANA pairs attentive service and excellent Japanese dining with an outstanding seat on aircraft equipped with The Room. It ranks third because that exceptionally spacious, enclosed configuration is available only on selected aircraft, making the recommendation more flight-dependent.
— gpt-6-astra
It ranks third because the product is only on select aircraft and routes, and its global network is thinner than the Gulf and Singapore carriers.
— kimi-k3
Repeatedly emphasizes 'The Room' as one of the widest, most private business suites on select 777-300ERs, paired with Japanese hospitality and refined dining, with the main caveat being limited availability of the top product.
Focuses on attentive service, excellent Japanese dining, and the standout spacious seat on aircraft with The Room, while ranking ANA third due to that cabin's limited availability making the experience flight-dependent.
Consistently praises 'The Room' on 777s as one of the widest, door-equipped seats with strong Japanese service and catering, but notes fleet inconsistency, limited route availability, and a thinner network outside Asia-Pacific.
Consistently highlights 'The Room' as arguably the most spacious business seat with doors, plus excellent Japanese service and catering, while noting its limitation to select aircraft/routes and a more Japan/Asia-centric network.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Emirates settles in the upper-middle of the field, with an overall median rank of 4 (Q1 3, Q3 4) across 18 ranked answers from 4 models. Agreement on placement is fairly tight but not uniform: claude-opus-5 puts it highest at a median rank of 3 (Q1 3, Q3 3), gemini-3.7-flash and kimi-k3 both land at a median of 4, while gpt-6-astra is the outlier with a median of 5 (Q1 5, Q3 5). The narrative behind the numbers is remarkably consistent across all four — the brand is credited for its A380 onboard bar and lounge, ICE entertainment, chauffeur service, Dubai ground experience and vast network, then held back by one recurring flaw.
That flaw is the older 2-3-2 seating on much of the 777 fleet, which lacks direct aisle access and drives the perception of product inconsistency. This tension is what separates the more generous placements from gpt-6-astra's fifth-place ranking, where the seating variability weighs heaviest.
The A380 business cabin's 1-2-1 layout is good but older 777 cabins in 2-3-2 are a weak point, though the new 777 suites are much improved.
— claude-opus-5
It ranks fifth among these options because the seating experience varies substantially: some older Boeing 777 cabins lack direct aisle access for every passenger and offer less privacy than the leading alternatives.
— gpt-6-astra
The recalled sources cluster around industry review and awards bodies that underpin both the praise and the caveat: Skytrax World Airline Awards appears under every model, alongside One Mile at a Time and The Points Guy, with claude-opus-5 and gpt-6-astra also citing Emirates directly, and gemini-3.7-flash and kimi-k3 leaning on Business Traveller.
Emirates is praised for its A380 onboard bar/lounge, strong entertainment, excellent Dubai ground experience and vast network, but consistently criticized for the older 2-3-2 777 business cabins lacking direct aisle access, which newer retrofits are improving.
Emirates is highlighted for its glamorous A380 onboard bar/lounge, ICE entertainment, chauffeur service and extensive Dubai network, with the recurring caveat that inconsistent 2-3-2 seating on older 777s keeps it from ranking higher.
Emirates is noted for its glamorous A380 onboard bar/lounge, strong soft product, chauffeur service and Dubai network, but repeatedly flagged for inconsistency due to older 2-3-2 777 seats without direct aisle access.
Emirates is valued for its extensive network, entertainment and A380 onboard lounge, but ranks fifth because seating varies substantially by aircraft, with older 777 cabins offering less privacy and no direct aisle access for every seat.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Cathay Pacific occupies a mid-table position across the three models, landing at a median rank of 4 (Q1 4, Q3 5) over 9 ranked answers. The models agree on the underlying strengths — comfortable reverse-herringbone seating, the newer Aria Suite, and highly regarded Hong Kong lounges such as The Pier — but diverge on placement. gpt-6-astra holds it steadiest at a median of 4 across 5 answers, while gemini-3.7-flash and kimi-k3 both place it at 5. The gap is one of degree rather than direction: all three cite the same qualities, but weigh the shortcomings differently.
For gpt-6-astra, the ceiling is set by cabin generation, noting that older business-class cabins offer less privacy than leading enclosed suites and that the Aria Suite is not fleet-wide. kimi-k3 pushes it lower on service and recovery concerns, pointing to an aging hard product and inconsistent post-pandemic performance. The recalled sources cluster around Skytrax World Airline Awards, The Points Guy, and One Mile at a Time, with Cathay Pacific's own site and outlets like Business Traveller and Executive Traveller feeding the seat, lounge, and dining themes.
Its newer Aria Suite strengthens the offering, but varying cabin generations keep it below the top three for an unspecified itinerary.
— gpt-6-astra
It ranks last here because the hard product is aging, and service and catering have been inconsistent during its post-pandemic rebuild compared with the carriers above.
— kimi-k3
Consistently frames Cathay Pacific as a well-rounded long-haul option with comfortable seating and excellent Hong Kong lounges, ranking it fourth because older cabin generations offer less privacy than leading suites and the newer Aria Suite is not fleet-wide.
Emphasizes the well-designed reverse-herringbone seats and new Aria Suite, along with top-rated Hong Kong lounges like The Pier and quality bedding and dining.
Notes comfortable reverse-herringbone seats and excellent Hong Kong lounges like The Pier, but ranks it last/fifth citing an aging hard product, inconsistent post-pandemic service, and slower recovery compared with rivals.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
where
where
6 AI models · 5 answers per model · ranked by score
ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0
Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.
Swipe left for the full report
Median rank each model gave each brand.
Scroll the table sideways to see every model →
| claude-opus-5 | deepseek-v4.1-flash | gemini-3.7-flash | gpt-6-astra | kimi-k3 | mistral-medium-3.5 | |
|---|---|---|---|---|---|---|
| Tesla | 2.0 | 1.0 | 1.0 | 3.0 | 1.0 | 1.0 |
| Hyundai | 1.0 | 2.0 | 2.0 | 1.0 | 2.0 | 4.0 |
| Kia | 3.0 | 3.0 | 3.0 | 2.0 | 3.0 | — |
| Rivian | 5.0 | 5.0 | 4.0 | — | 4.0 | 3.0 |
| BMW | 4.0 | — | 4.0 | 4.0 | 4.0 | — |
| Ford | 5.0 | 4.0 | 5.0 | 5.0 | 4.5 | 4.0 |
| BYD | — | — | — | 5.0 | 5.0 | 2.0 |
| Nissan | — | — | — | — | — | 5.0 |
| Chevrolet | — | 5.0 | — | — | 5.0 | — |
| Volvo | — | — | — | 5.0 | — | — |
| Lucid | — | — | — | — | 5.0 | — |
A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.
5 of 11 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.
mistral-medium-3.5 8 → gpt-6-astra 100
claude-opus-5 0 → mistral-medium-3.5 80 · 3 of 6 never named it
mistral-medium-3.5 0 → gpt-6-astra 80 · 1 of 6 never named it
gpt-6-astra 0 → mistral-medium-3.5 60 · 1 of 6 never named it
deepseek-v4.1-flash 0 → gemini-3.7-flash 44 · 2 of 6 never named it
gpt-6-astra 60 → mistral-medium-3.5 100
A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.
Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.
The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.
Tesla leads with a score of 89 of 100, ranked by 6 of 6 models.
Tesla anchors the category, ranked by every model and scoring 88.67, but the panel divides on how much weight to give its well-documented caveats rather than on the caveats themselves. Four models put it first almost always; claude-opus-5 and gpt-6-astra pull it down, with gpt-6-astra sitting furthest away at a median rank 3, placing it behind Hyundai and Kia on interface and build-quality grounds.
per-model scores 60–100 of 100 · mean 89 across 6 models
The sharpest genuine split in the field is Hyundai, whose per-model scores run the full width from 8 to 100. Claude-opus-5 and gpt-6-astra rank it first outright, three models settle it second, and mistral-medium-3.5 names it just once at rank 4. The high scorers converge on the E-GMP platform and value story, so the divide is one of enthusiasm rather than of substance.
Hyundai · 92 points apart on a 0–100 scale
Kia tracks Hyundai closely as its E-GMP near-peer, but the divergence there is coverage, not sentiment — gemini-3.7-flash names it in only 40% of runs while others hold it steadily at rank 3, and no model ranks it first. Rivian and BMW sit in similar territory: both are read consistently on the merits, with the spread driven by how far each model lets caveats (young-company risk for Rivian, price for BMW) pull the brand down, and by silent models scoring zero.
The most instructive disagreement is BYD, where mistral-medium-3.5's rank-2 conviction collides with two models placing it fifth. All three agree on the battery and value strengths; they split on whether limited Western availability should cap the ranking — a geography question, not an engineering one.
The lower tier — Ford, Nissan, Chevrolet, Volvo, Lucid — shows "agreement" that is mostly omission. Ford is the exception, named by all six but never first, its rank-4-to-5 clustering a true consensus. The others register with one or two models each, so their low standard deviations reflect near-universal absence rather than a shared verdict.
The recalled source base is uniform enough that it explains little of the divergence: Car and Driverwww.caranddriver.com×32 and Edmundswww.edmunds.com×29 dominate nearly every brand, with Consumer Reportswww.consumerreports.org×23 recurring on reliability points. Because the same outlets sit behind both high and low placements, any link between a specific source and a model's harsher or softer read is a hypothesis the data does not confirm — the models appear to weight shared evidence differently rather than to draw on different evidence.
37 sources across 431 references, grouped by site from 141 recalled names. The top 5 carry 66% of them.
caranddriver.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Tesla scores 88.67 of 100 on Borda and was ranked by all 6 of the 6 models, placing it among the strongest recommendations in the category. The panel reaches broad agreement on its merits: the recurring theme across every model is the Supercharger network as the anchor of an EV ownership ecosystem, paired with energy efficiency, range, and over-the-air software updates. Where models diverge, they do so on how much weight to give the widely noted caveats — inconsistent build quality, minimalist touchscreen-heavy controls, and uneven service — rather than on whether those caveats exist.
per-model scores 60–100 of 100 · mean 89 across 6 models
That split is visible in the per-model placements. Four models (deepseek-v4.1-flash, gemini-3.7-flash, kimi-k3, mistral-medium-3.5) rank it first most or all of the time, treating build quality as a minor caveat within a seamless ecosystem. Claude-opus-5 settles it at a median rank 2, framing it as the top pick specifically for charging convenience and long-distance travel, while gpt-6-astra sits furthest away at a median rank 3, consistently ranking it below Hyundai and Kia because the interface and fit-and-finish trade-offs make it, in its words, a less universal recommendation. The recalled sources cluster consistently under these themes — Consumer Reportswww.consumerreports.org×23 and Edmunds recur behind the reliability and ownership caveats, while EPA fueleconomy.gov and InsideEVs underpin the efficiency and range claims.
“Tesla remains the benchmark for EVs due to its Supercharger network, long range, and advanced software.”
“I rank it below Hyundai and Kia because its touchscreen-heavy controls, reported build-quality inconsistencies, and uneven service experiences are meaningful trade-offs.”
Repeatedly positions Tesla as the EV benchmark for most buyers based on its Supercharger network, range, performance, and over-the-air updates, while noting mixed build quality and service or reliability concerns.
Consistently calls Tesla the industry benchmark for its Supercharger network, energy efficiency, software integration, and OTA updates, treating variable build quality as a minor caveat within a seamless ownership ecosystem.
Repeatedly describes Tesla as the most complete and easiest EV to own thanks to efficiency, Supercharger access, software/OTA updates, competitive pricing, and resale value, while citing inconsistent build quality and service.
Consistently frames Tesla as the leader in EV technology, battery innovation, and charging infrastructure via the Supercharger network, emphasizing performance, range, and software with no drawbacks noted.
Consistently frames Tesla as the top pick for charging convenience and long-distance travel due to its mature Supercharger network, efficiency, range, and software/OTA updates, while flagging inconsistent build quality, minimalist interiors, and variable service.
Highlights Tesla's efficiency, integrated route planning, and convenient charging for road trips, but consistently ranks it below Hyundai and Kia due to touchscreen-heavy controls, fit-and-finish variability, and uneven service.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
no URL recalled
consumerreports.org
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Hyundai lands at a Borda score of
Stat unavailableand was ranked by all 6 models, yet the panel is far from unanimous. Per-model borda spans 8 to 100 with a standard deviation of 31.28, which the report labels "sharply split." Two models (claude-opus-5 and gpt-6-astra) placed it first in all five of their runs, three others (deepseek-v4.1-flash, gemini-3.7-flash, kimi-k3) settled it at a median rank of 2, while mistral-medium-3.5 named it only once at rank 4 for a coverage of 20% and a borda of 8. The recurring theme across the high scorers is Hyundai's E-GMP platform: 800V ultra-fast charging, competitive pricing, long range, and a generous warranty, framed as the best all-around value alternative to Tesla.
Hyundai · 92 points apart on a 0–100 scale
The sources under these themes cluster on mainstream automotive review outlets — Car and Driver, Edmunds, MotorTrend, IIHS, and Consumer Reports recur across models — with award citations such as World Car Awards backing the "well-rounded" framing. The caveats that hold Hyundai just behind the top are consistent too: less mature software, historically weaker charging network access (partly offset by NACS adoption), and ICCU recalls that several models flag as worth verifying before purchase. Mistral's lone, lower placement rests on a similar read but positions Hyundai as reliable rather than cutting-edge.
“It's the best all-around balance of value, tech, and dependability for most buyers.”
“While not as cutting-edge as Tesla or BYD, it provides reliable and well-rounded options.”
Consistently highlights the E-GMP platform's 800V ultra-fast charging, strong efficiency/range, competitive pricing, and excellent warranty, framing Ioniq 5/6 as the best all-around value, with NACS access and awards noted; ICCU recalls and software glitches mentioned as caveats.
Frames Hyundai as the best all-around recommendation for balancing efficiency, comfort, value, and fast charging, while repeatedly cautioning that charging performance varies by model and buyers should verify recall completion.
Emphasizes the E-GMP platform's 800V ultra-fast charging, long range, competitive pricing, and generous warranty, positioning Hyundai as a well-rounded value choice; a smaller charging network is noted as the main improving downside.
Repeatedly centers on the advanced 800V E-GMP platform enabling ultra-fast charging at mainstream prices, paired with distinctive styling, generous warranty, and strong overall value.
Focuses on the E-GMP 800V platform's very fast charging, distinctive design, long battery warranty, and price advantage, ranking Hyundai just behind Tesla due to less mature software and charging network access.
Describes Hyundai as offering a diverse EV lineup with solid range and competitive pricing, reliable and well-rounded but not as cutting-edge as Tesla or BYD.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Kia lands at
Stat unavailableon Borda, ranked by 5 of the 6 models, but the panel does not settle on a single reading. Per-model borda spans 0–80 with a standard deviation of 27.23, which the report labels "models split." The divide is largely one of coverage rather than sentiment: gpt-6-astra places it consistently at median rank 2 for a borda of 80, while gemini-3.7-flash ranks it in only 2 of 5 runs, holding coverage to 40% and borda to 24. The three remaining models cluster tightly at median rank 3, scoring between 60 and 64.
per-model scores 0–80 of 100 · mean 48 across 6 models
The recurring theme is that Kia is a near-peer to Hyundai on the shared 800V E-GMP platform, valued for the three-row EV9, fast charging, and warranty coverage, but held just behind on pricing at higher trims, inconsistent dealer experience, and availability. Those points recur across the models' recalled sources, with Car and Driver×10 and Edmunds×6 appearing most often, alongside award citations and reliability references such as J.D. Power, IIHS, and Consumer Reports. First-choice score sits at 27.78, and no model ranked Kia first in any run, consistent with the repeated framing of it as a strong second or co-equal rather than the lead pick.
“Kia is a close second, offering compelling electric family vehicles with practical interiors, strong warranties, and excellent charging speeds on its newer dedicated EV platforms.”
“Essentially a co-equal to Hyundai, ranked just behind on value per dollar.”
Kia is consistently a close second to Hyundai, strongest for families needing space and three-row seating, but placed behind because larger, better-equipped models become expensive and value depends on needs.
Kia is framed as a co-equal to Hyundai on the shared 800V E-GMP platform, praised for the three-row EV9, fast charging, and strong warranty, but ranked just behind Hyundai due to rising prices on top trims and inconsistent dealer experience.
Kia is positioned as sharing Hyundai's E-GMP platform while adding bolder styling and value, with the three-row EV9 as a standout, though dealer markups and availability are noted drawbacks.
Kia is highlighted for exceptional value, fast charging, and family-friendly practicality, with models consistently earning awards for ergonomics, range, and price-to-performance.
Kia shares Hyundai's excellent E-GMP platform with strong styling and the standout three-row EV9, but ranks just below Hyundai due to higher pricing on top trims and inconsistent dealer/availability.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Rivian lands at a modest
Stat unavailableand is named by 5 of the 6 models, yet the panel is far from settled on where it belongs. Per-model borda runs from 0 to 60 with a standard deviation of 20.58, which the report labels "models split" — mistral-medium-3.5 ranked it median 3 across all five runs, while claude-opus-5 and deepseek-v4.1-flash placed it at median rank 5 where they named it at all. The gap reflects a shared read applied with different weight rather than genuine disagreement about the product: every model describes the R1T and R1S as capable, adventure-focused vehicles with strong software, then discounts them for premium pricing, a thin service network, and the financial and reliability uncertainty of a young automaker.
per-model scores 0–60 of 100 · mean 25 across 6 models
The evidence sits under a consistent set of automotive review sources, with MotorTrendwww.motortrend.com×21 and Edmundswww.edmunds.com×29 recurring across models, alongside Consumer Reports where reliability risk is raised. What separates the rankings is how far each model lets the caveats pull the brand down: for the more cautious models the young-company risk and thin support network reduce it to a niche recommendation, while mistral-medium-3.5 treats the same limitations as constraints on breadth rather than on quality. No model ranked it first in any run.
“The R1T and R1S are arguably the best-driving adventure EVs available, with excellent software, clever packaging, and some of the highest owner-satisfaction scores in the industry.”
“Rivian excels in adventure-focused EVs with impressive off-road capabilities and a strong brand identity.”
Rivian excels in adventure-focused EVs with off-road capability and premium build quality, though a limited model lineup, higher prices, and still-growing production scale restrict broader appeal.
Rivian excels in premium, adventure-focused trucks and SUVs with off-road capability, refined interiors, and strong software, but higher prices and a still-expanding service and charging network limit its appeal for everyday drivers.
The R1T and R1S are described as outstanding adventure EVs with clever design and high owner satisfaction, but high prices, a thin service network, and financial/reliability risks for a young company make it a compelling yet riskier choice.
Rivian's R1T and R1S are praised as capable, well-engineered adventure vehicles with strong software and loyal owners, but the recommendation is qualified by premium pricing, thin service networks, and young-automaker financial and reliability uncertainty.
The R1T and R1S are praised for off-road capability and performance, but high prices and a still-growing service network make Rivian a niche choice for adventure enthusiasts rather than most buyers.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
BMW lands at a Borda score of 22 of 100, held down by being ranked by only 4 of the 6 models rather than by any hostility from those that did name it. Among the four that ranked it, placement is strikingly stable: gemini-3.7-flash gave it 44, claude-opus-5 and gpt-6-astra both 40, with each of those three covering it in every run and consistently settling on a median rank of 4. The gap that the report labels "models split" comes almost entirely from the two silent models scoring 0 and from kimi-k3's thin 20% coverage yielding a borda of 8.
per-model scores 0–44 of 100 · mean 22 across 6 models
The themes are consistent across the models that engaged with it: refined driving dynamics, quiet and well-built premium cabins, and above-average reliability, offset uniformly by higher purchase prices, costly options, and 400V-class charging that trails the 800V Korean rivals. That is why the brand clusters at fourth rather than higher — the models frame the ranking as a value trade-off, not an EV weakness. The recalled sources reflect this enthusiast-and-ownership lens, with Car and Driver×10 and Edmunds×6 recurring alongside Consumer Reports×7 on the reliability point.
“BMW is my strongest premium pick here for refined cabins, ride quality, and engaging driving dynamics across its electric range.”
“They rank lower mainly on value: prices are high, options add up quickly, and charging speeds are good rather than class-leading.”
Consistently frames BMW's i4/iX/i5 as refined, well-built premium EVs with above-average reliability and traditional luxury feel, while noting higher costs and slower charging than 800V Korean rivals.
Repeatedly emphasizes that BMW translates its signature driving dynamics and luxury craftsmanship into EVs with quiet, refined cabins, but cites higher prices and lower efficiency/slower charging as drawbacks.
Consistently positions BMW as a premium pick for refinement, cabin quality, and engaging driving dynamics, ranking it fourth mainly due to high purchase prices and costly options rather than any EV weakness.
Describes BMW's i4 and iX as refined, well-built premium EVs with respectable range, ranking lower on value due to high prices, costly options, and merely good charging speeds.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
caranddriver.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
where
where
5 AI models · 5 answers per model · ranked by score
ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0
Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.
Swipe left for the full report
Median rank each model gave each brand.
Scroll the table sideways to see every model →
| claude-opus-5 | gemini-3.7-flash | gpt-6-astra | kimi-k3 | mistral-medium-3.5 | |
|---|---|---|---|---|---|
| Cartier | 2.0 | 2.0 | 2.0 | 2.0 | 3.0 |
| Omega | 3.0 | 1.0 | 4.0 | 1.0 | 2.0 |
| Seiko | 2.0 | 3.0 | 3.0 | 4.0 | 4.0 |
| Rolex | 5.0 | 5.0 | 5.0 | 5.0 | 1.0 |
| NOMOS Glashütte | — | — | 2.0 | — | — |
| Tudor | — | 3.0 | — | 3.0 | — |
| Tissot | 3.0 | 5.0 | — | — | — |
| Casio | — | — | — | 5.0 | 5.0 |
| Apple | — | — | — | 5.0 | — |
| Hublot | — | 5.0 | — | — | — |
| Timex | — | — | 5.0 | — | — |
A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.
5 of 11 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.
gpt-6-astra 16 → mistral-medium-3.5 100
gpt-6-astra 0 → claude-opus-5 60 · 3 of 5 never named it
gpt-6-astra 44 → gemini-3.7-flash 100
claude-opus-5 0 → kimi-k3 48 · 3 of 5 never named it
kimi-k3 36 → claude-opus-5 80
mistral-medium-3.5 60 → gpt-6-astra 88
A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.
Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.
The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.
Cartier leads with a score of 78 of 100, ranked by 5 of 5 models.
The models agree most closely on Cartier, converging at a median rank of 2 with tight quartile boundaries. The consensus hinges on its design-first heritage and romantic elegance: every model frames Cartier as signaling taste over wealth, with the Tank and Santos repeatedly praised for sliding under a cuff without shouting. The sources behind this alignment—Hodinkee, GQ, Vogue, Esquire—are lifestyle and watch publications that emphasize cultural credibility and aesthetic literacy. No model strays far from this narrative, though mistral-medium-3.5 edges it to rank 3, likely reflecting a minor preference for other options rather than a substantive disagreement.
Omega reveals the category's sharpest divergence. Two models (gemini-3.7-flash and kimi-k3) place it at rank 1, praising its balance of prestige and approachability. Two others (mistral-medium-3.5 and claude-opus-5) rank it at 2 or 3, noting its versatility but occasional tilt toward "trying a bit hard." Gpt-6-astra consistently ranks it at 4, arguing that its sporty visibility and recognizable luxury feel more conspicuous than the understated charm the date context demands. The split suggests different thresholds for what counts as "statement-making," with gemini and kimi reading Omega as confident sophistication while gpt-6-astra sees it as a step too assertive. Hodinkee, Fratello Watches, and GQ anchor all perspectives, meaning the divergence likely stems from differing model priors about appropriate visibility rather than conflicting source material.
Rolex sits at a median rank of 5 overall, but mistral-medium-3.5 ranks it at 1 while the other four models all place it at 5. The majority view, supported by Hodinkee, GQ, r/Watches, and Robb Report, holds that Rolex's overwhelming brand recognition risks overshadowing personality and steering attention toward wealth. Mistral-medium-3.5 draws on Forbes, WatchTime Magazine, and the Rolex official website to frame the same visibility as confidence and universally understood elegance. This is the only brand where a model completely inverts the dominant interpretation, suggesting mistral may weight luxury marketing narratives more heavily than social signaling concerns.
NOMOS Glashütte stands out as the brand championed most exclusively by a single model. Gpt-6-astra ranks it at 2 across five separate assessments, citing its clean minimalist design and ability to signal personal taste without drawing excessive attention. No other model mentions NOMOS at all. The sources are entirely official NOMOS materials, raising the hypothesis that gpt-6-astra's training or retrieval corpus may include stronger representation of this independent German brand, while other models default to more widely discussed names. The brand's absence elsewhere in the rankings underscores how differently models weight niche design-forward marques versus mainstream luxury.
Seiko earns broad middle-tier consensus at a median rank of 3, with claude-opus-5 slightly higher at 2 and kimi-k3 and mistral-medium-3.5 at 4. The models converge on Seiko as understated, value-driven, and signaling horological curiosity over status, drawing on Hodinkee, Worn & Wound, r/Watches, and Teddy Baldassarre. The tension centers on occasion-appropriateness: while all acknowledge Seiko's charm, those ranking it lower flag that many models skew too casual or tool-like for dressier dates. The brand's wide catalog means specific model choice matters more than the name itself, yet the sources cited—enthusiast communities and watch blogs—are consistent across models, suggesting the ranking spread reflects different weightings of formality over authenticity rather than conflicting source narratives.
Tudor receives tight agreement at rank 3 from the two models that include it (gemini-3.7-flash and kimi-k3), both emphasizing Rolex-adjacent quality in a lower-key package. The models frame Tudor as ideal for those who prioritize substance over status, drawing on Hodinkee, aBlogtoWatch, and Monochrome Watches. The shared concern is mainstream recognition: Tudor resonates with enthusiasts but may go unnoticed by others, and its tool-watch aesthetic can feel utilitarian for formal dates. The narrow spread suggests strong alignment on Tudor's positioning, with the main variation being whether a given scenario skews casual enough to favor its grounded confidence.
Tissot occupies a median rank of 3.5, with claude-opus-5 at 3 and gemini-3.7-flash at 5. Both models praise the PRX as polished and versatile, supported by Fratello Watches, Worn & Wound, GQ, and Hodinkee. The disagreement centers on distinctiveness: claude views Tissot as a "low-risk, high-polish option," while gemini faults it for lacking "the distinct character or conversational intrigue of higher-ranked heritage brands." The PRX's ubiquity is flagged by both—claude notes it "says less about you personally" due to widespread popularity, while gemini sees it as "more like a sensible default." This suggests models diverge on whether safety or individuality matters more for first dates, even when drawing on overlapping sources.
Casio lands uniformly at rank 5 from both models that assess it, reflecting agreement that its utilitarian aesthetic undercuts sophistication. Kimi-k3 and mistral-medium-3.5 draw on TechRadar, the Casio Official Website, and Gear Patrol to acknowledge reliability and occasional conversational charm (vintage digital models), but both conclude that typical resin designs read as too casual. The consensus is tight, with the only variation being kimi's acknowledgment that specific lines like the Edifice might work for relaxed settings. The shared sources and reasoning suggest this is one of the category's clearest points of agreement.
The remaining brands—Apple, Hublot, Timex—appear only once each, all at rank 5. Apple is flagged by kimi-k3 as too distracting and insufficiently formal, supported by Apple's own site and The Verge. Hublot, assessed by gemini-3.7-flash, is deemed too bold and polarizing, with Time and Tide Watches and Watchuseek cited as community perspectives on its loudness. Timex, ranked by gpt-6-astra, is praised as unpretentious but noted as less distinctive than preferred alternatives, with Timex's official site as the sole source. These single-model assessments reflect idiosyncratic inclusion rather than category-wide themes.
The most frequently recalled sources—Hodinkee (53 citations), GQ (17), Fratello Watches (15), Worn & Wound (12), and r/Watches (11)—dominate across models, anchoring the consensus that first-date watches should balance quality with restraint. The hypothesis that these outlets shape model views more than official brand sites is supported by the fact that even when models cite manufacturer websites, they pair them with editorial perspectives that emphasize social appropriateness over technical specs. The exception is gpt-6-astra's exclusive reliance on NOMOS's official materials, suggesting that where a brand lacks strong third-party coverage in a model's training data, official sources may fill the gap and produce outlier rankings.
55 sources across 313 references, grouped by site from 99 recalled names. The top 5 carry 44% of them.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Cartier achieves a median rank of 2 across all models, with strong consensus around that position—the first and third quartiles sit at 2 and 3, respectively. Four of the five models place it at rank 2, while one (mistral-medium-3.5) ranks it at 3. The models converge on a consistent narrative: Cartier signals design literacy and romantic elegance rather than technical prowess or wealth display. Its iconic models—particularly the Tank and Santos—are repeatedly framed as jewelry-adjacent styling choices that read as thoughtful and refined. The brand is seen as recognizable enough to non-enthusiasts that it can start conversations without feeling like ostentation, though several models note its luxury associations and dressier character may skew slightly formal for very casual first encounters.
The sources supporting these themes are predominantly watch and lifestyle publications including Hodinkee, GQ, Vogue, Esquire, and A Blog to Watch, along with direct references to Cartier's own website. These outlets anchor the perception of Cartier as a design-first heritage brand with cultural credibility extending beyond collector circles. The emphasis on romantic Parisian elegance, timeless style, and aesthetic over specification is consistent across models, with only minor variation in how strongly each flags the potential formality trade-off.
“A Tank or Santos reads as taste rather than money — slim, elegant, and it slides under a shirt cuff without shouting.”
“Cartier offers timeless elegance and an effortlessly romantic design aesthetic suited for a date setting.”
Consistently emphasizes that Cartier reads as taste and design rather than wealth-flexing, noting it's elegant, recognizable to non-enthusiasts, and slides discreetly under a cuff while carrying design credibility from cultural icons.
Repeatedly highlights Cartier's romantic elegance, artistic design pedigree, and timeless sophistication that signals refined personal style and cultural appreciation rather than technical specifications or wealth display.
Focuses on Cartier's distinctive, jewelry-like design aesthetic as thoughtful styling rather than technical showmanship, though notes the luxury association can feel formal or conspicuous for casual settings.
Emphasizes Cartier's romantic, Parisian elegance and design-first appeal that reads as refined and thoughtful rather than status-seeking, though consistently notes its dressier character may be slightly formal for very casual dates.
Repeatedly describes Cartier as refined, understated luxury with sleek elegance, highlighting its sophistication and heritage in fine jewelry and watchmaking as ideal for classy, upscale dates.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Omega lands consistently near the top of model recommendations, with a median rank of 2 overall. Two models rank it at the top (median rank 1), two place it slightly lower at rank 2 and 3, and one model consistently positions it at rank 4. The core theme uniting most perspectives is that Omega occupies a strategic middle ground: prestigious enough to signal taste and success, yet sufficiently understated to avoid appearing ostentatious or status-driven on a first date. Models emphasize its versatility across casual and formal settings, its rich heritage through associations with space exploration and James Bond, and its ability to serve as a natural conversation starter without demanding attention. Sources supporting this view cluster around enthusiast publications like Hodinkee, Fratello Watches, and GQ, alongside official Omega materials.
The primary point of divergence appears in how models weigh visibility against subtlety. While gemini-3.7-flash and kimi-k3 prize Omega's balance as ideal for communicating refined taste without trying too hard, gpt-6-astra consistently ranks it lower, citing its sporty aesthetic and visual prominence as more conspicuous than preferred for an understated first-date impression. Claude-opus-5 occupies the middle, acknowledging Omega's quality and story but occasionally noting it can register as "slightly more statement than charm" or risk tipping into "trying a bit hard" territory depending on context. The tension centers on whether Omega's recognizable luxury presence reads as confident sophistication or as a step beyond what some models consider optimally low-key.
“Omega hits the sweet spot for a first date: genuinely respected by watch people yet not a status symbol that screams money, and models like the Aqua Terra or Speedmaster work with both a blazer and a t-shirt.”
“Omega strikes the ideal balance between prestige, timeless design, and understated sophistication. It communicates refined taste without appearing ostentatious or trying too hard on an initial meeting.”
Omega repeatedly represents an optimal balance between prestige and restraint, communicating refined taste, success, and sophistication without crossing into ostentation or appearing to try too hard.
Omega reliably hits the sweet spot of signaling genuine taste and success without ostentation, with its Speedmaster and Seamaster heritage providing natural conversation starters that read as confident rather than flashy.
Omega consistently blends luxury with sportiness and approachability, while its associations with space exploration and James Bond add elements of intrigue, adventure, and charm that support both refinement and conversational interest.
Omega consistently occupies a strategic middle ground—respected by enthusiasts yet not overtly status-driven, versatile enough for varied settings, and offering natural conversation hooks through its Moonwatch and Bond heritage without demanding attention.
Omega is recognized as offering good versatility and conversation potential through its technical heritage, but its sporty aesthetic and visual prominence consistently register as more conspicuous than this model's preference for understated first-date choices.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Seiko lands at a median rank of 3 across models, with fairly strong consensus (Q1 3, Q3 4), though claude-opus-5 rates it slightly higher (median rank 2) while kimi-k3 and mistral-medium-3.5 place it at rank 4. The models converge on a common set of themes: Seiko signals understated confidence, genuine appreciation for craftsmanship over branding, and accessible quality that avoids pretense or status-seeking. It earns consistent praise for models like the Presage and Alpinist that deliver visual refinement without requiring luxury spending, making them safe, approachable choices that keep attention on the wearer rather than the watch. The brand draws support primarily from enthusiast sources including Hodinkee, Worn & Wound, r/Watches, Fratello Watches, and Teddy Baldassarre, alongside official Seiko materials.
The main tension in the rankings centers on versatility versus occasion-appropriateness. Models note that while Seiko offers solid value and thoughtful design, many pieces skew too casual or tool-like for dressier date settings, and the brand lacks the romantic recognition or elevated finish of Swiss alternatives. The broad range within Seiko's catalog also means the specific model chosen matters more than the brand name itself. Still, the consensus holds that Seiko communicates substance, humility, and horological curiosity—traits the models view as charming rather than flashy for a first-date context.
“A Seiko (think Presage, Alpinist, or a slim SUR-series dress piece) looks genuinely good without announcing a price tag, which is exactly the tone you want on a first date.”
“Seiko is the enthusiast's choice — wearing one suggests you care about craftsmanship and value rather than logos, which can come across as genuinely charming.”
Seiko delivers understated charm and horological credibility without signaling wealth or status, keeping attention on the wearer rather than the watch. Models like the Presage and Alpinist punch above their price point while avoiding the risk of seeming pretentious.
Seiko represents grounded confidence and appreciation for craftsmanship without relying on luxury branding or pretense. It communicates practical elegance and authenticity, though it lacks the elevated polish that creates a special-occasion feeling.
Seiko provides attractive, understated options that work for relaxed dates without requiring luxury spending, though its broad range makes the specific model choice more critical than the brand name. It allows intentional style on an accessible budget.
Seiko signals genuine enthusiasm for craftsmanship and value over logos, which reads as charming and unpretentious to those who appreciate substance. It ranks lower because many models skew too casual or tool-like for dressier date settings and it lacks the romantic recognition of Swiss brands.
Seiko balances quality craftsmanship, reliability, and affordability while offering versatile styling, particularly in lines like Presage and Cocktail Time. Though not prestigious, it demonstrates practical good taste and substance over flash.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Rolex sits at a median rank of 5 across the models, revealing a sharp divide in perception. One model consistently places it first, drawing on sources including Forbes, WatchTime Magazine, and the Rolex official website to argue that the brand signals timeless elegance, confidence, and universally recognized quality. The remaining four models all assign it a median rank of 5, citing sources such as Hodinkee, GQ, r/Watches, and Robb Report to emphasize that while Rolex offers exceptional craftsmanship, its overwhelming brand recognition risks overshadowing personality and steering attention toward wealth rather than genuine connection. The disagreement centers on whether iconic status is an asset or a liability in first-date contexts.
The majority theme across models is that Rolex's visibility as a luxury symbol creates conversational friction before rapport is established. Three models warn it may read as conspicuous consumption or status-flexing, and two note that its price is immediately recognizable to most observers, inviting assumptions about the wearer's priorities. Even the dissenting model acknowledges the brand's strong luxury association, though it frames this as confidence rather than distraction. The split reflects tension between Rolex's undisputed prestige and the social risk of leading with a watch that, as multiple models note, speaks before the wearer does.
“Rolex is superbly made, but it's the one brand almost everyone can price at a glance, so it tends to speak before you do.”
“While Rolex is the most universally recognized luxury watchmaker, its overt prestige can inadvertently make you seem status-focused or ostentatious on a first outing.”
Rolex represents timeless elegance and sophistication that signals confidence, success, and quality while making a strong impression that is universally recognized and respected.
Rolex is excellent in quality but too recognizable as a wealth signal on first dates, risking distraction from genuine connection by broadcasting status before conversation begins.
Despite exceptional craftsmanship, Rolex's overwhelming brand recognition risks projecting pretension or status-seeking that can overshadow personality and distract from authentic first-date conversation.
Rolex offers quality and classic design, but its strong luxury association draws unwanted attention to price and status rather than keeping focus on the conversation during a first meeting.
Rolex is iconic and well-made, but on first dates it carries high risk of appearing as conspicuous consumption or wealth-flexing that overshadows personality before conversation starts.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
NOMOS Glashütte receives consistently strong support from the model that ranked it, landing at a median rank of 2 across five separate assessments. The quartile range from 1 to 2 shows remarkable agreement, with the brand sometimes claiming the top position and otherwise placing second. This narrow spread reflects a coherent perception of NOMOS as a first-date choice that balances aesthetic distinction with restraint. The model draws on five separate references to the official NOMOS Glashütte website to support its evaluations.
The recurring themes center on the brand's clean, minimalist design language and its ability to signal personal taste without drawing excessive attention. The model values NOMOS for typography, understated details, and occasional playful color touches that convey thoughtfulness and design awareness rather than status signaling. When NOMOS ranks second rather than first, the reasoning typically involves comparison to more romantic or versatile options, with the model noting that the minimalist aesthetic may feel slightly more casual or design-focused than alternatives, and potentially less suited to rugged or ornate personal styles.
“My first choice would be NOMOS: its clean, understated designs suggest personal taste without making the watch the center of attention. That balance feels especially right for a first date, whether the setting is coffee or dinner.”
“NOMOS would be my choice for understated style: clean typography and restrained designs suggest attention to detail rather than an obvious status display.”
This model consistently emphasizes NOMOS's clean, understated design and restrained aesthetic that suggests thoughtfulness and personal taste without demanding attention. It values the brand's ability to balance distinctive style with a quiet, minimalist character appropriate for first dates.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
where
where
6 AI models · 4 answers per model · ranked by score
ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0
Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.
Swipe left for the full report
Median rank each model gave each brand.
Scroll the table sideways to see every model →
| claude-opus-5 | deepseek-v4.1-flash | gemini-3.7-flash | gpt-6-astra | kimi-k3 | mistral-medium-3.5 | |
|---|---|---|---|---|---|---|
| Starling Bank | 2.0 | 1.0 | 1.0 | 1.0 | 2.0 | 2.0 |
| Revolut | 3.0 | 3.0 | 2.5 | 3.5 | 3.0 | 1.0 |
| Tide | 1.0 | 4.0 | 3.5 | 3.5 | 1.5 | 3.0 |
| Monzo | 4.5 | 2.0 | 3.5 | 2.0 | 4.0 | 4.0 |
| Wise | 5.0 | 5.0 | 5.0 | 5.0 | 4.0 | — |
| Anna Money | — | 5.0 | — | — | 5.0 | 5.0 |
| Mercury | 4.0 | — | — | — | — | — |
| Mettle | — | — | — | — | 5.0 | — |
A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.
2 of 8 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.
deepseek-v4.1-flash 45 → claude-opus-5 100
claude-opus-5 30 → gpt-6-astra 80
gpt-6-astra 50 → mistral-medium-3.5 100
claude-opus-5 75 → gpt-6-astra 100
mistral-medium-3.5 0 → kimi-k3 20 · 1 of 6 never named it
claude-opus-5 0 → mistral-medium-3.5 20 · 3 of 6 never named it
A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.
Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.
The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.
Starling Bank leads with a score of 90 of 100, ranked by 6 of 6 models.
Starling Bank is the clearest point of consensus in this category, ranked by all six models and placed first by three of them. The split is narrow and structural: deepseek-v4.1-flash, gemini-3.7-flash and gpt-6-astra rank it first every time, while claude-opus-5, kimi-k3 and mistral-medium-3.5 hold it at rank 2 on the same reasoning — a full UK licence and FSCS protection, offset by weaker multi-currency and FX support. The disagreement is about ceiling, not floor.
per-model scores 75–100 of 100 · mean 90 across 6 models
Revolut and Tide are where the models diverge most sharply, and both divides turn on the same fault line: whether an internationally strong but non-fully-licensed provider belongs at the top. mistral-medium-3.5 ranks Revolut first in every run, while gpt-6-astra sits it lowest at a median rank of 3.5, docking it for plan fees on domestic use. Tide shows the widest spread of any brand (standard deviation 19.15), with claude-opus-5 placing it first and deepseek-v4.1-flash placing it fourth or lower — the same ClearBank e-money structure read as reassurance by one and as a gap by the other.
Tide · 55 points apart on a 0–100 scale
Monzo produces the most model-specific reversal. gpt-6-astra treats it as a close second (borda 80) suited to founder-led startups, while claude-opus-5 ranks it near-last (borda 30), treating its thinner multi-currency and integration set as disqualifying for scaling teams. No model ranked it first. This is the clearest case where two models applied the same facts to opposite conclusions about who Monzo is for.
Monzo ahead on 2 of 6 models
Where the panel converges is at the bottom. Wise draws unusual agreement (spread 0–20): five models rank it, none place it first, and all frame it as a treasury complement rather than a primary account, citing its lack of FSCS protection. Anna Money, Mercury and Mettle barely register, each pulled up only by isolated models — Mercury solely by claude-opus-5 on US-incorporation grounds, Mettle solely by kimi-k3.
The source patterns are broadly shared rather than distinguishing. Trustpilot×10 and brand-owned pages appear behind nearly every model's reasoning, which makes it hard to attribute the ranking splits to differing evidence. A plausible but unverified reading is that the licence-versus-features divide reflects how each model weighed FSCS status against international capability, not which sources it happened to recall — the recalled sources overlap too heavily to explain the gaps on their own.
43 sources across 313 references, grouped by site from 141 recalled names. The top 5 carry 43% of them.
trustpilot.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named Starling
Starling Bank is one of the strongest performers in this category, scoring
Stat unavailableand ranked by all six models. The agreement is broad: per-model borda spans 75–100, and the split runs cleanly between two camps. Three models — deepseek-v4.1-flash, gemini-3.7-flash and gpt-6-astra — placed it first in every run, drawn to a consistent set of themes: a full UK banking licence with FSCS protection, a fee-free business account, accounting integrations (Xero, QuickBooks) and strong UK customer support. These points are anchored in recalled sources such as Trustpilot×10, Which?×6 and the Starling Bank business account page.
per-model scores 75–100 of 100 · mean 90 across 6 models
The remaining three — claude-opus-5, kimi-k3 and mistral-medium-3.5 — held it at rank 2, applying the same strengths but noting weaker multi-currency and international/FX support and less startup-specific tooling relative to rivals such as Tide, Revolut and Wise. That reservation, rather than any doubt about safety or day-to-day fitness, is what separates a first-place vote from a second.
“Starling is a fully licensed UK bank with FSCS protection, free everyday business banking, and consistently top-rated customer service in the CMA service quality surveys.”
“It ranks second rather than first because its multi-currency and international payment capabilities are weaker than Revolut's or Wise's.”
Highlights Starling as a fully licensed, FSCS-protected UK bank with a free business account, strong app, and accounting integrations (Xero, QuickBooks), while noting weaker international/FX support.
Repeatedly stresses full UK regulatory protection (FSCS up to £85,000), zero monthly fees, accounting software integrations, and award-winning UK customer support as making it the top choice for startups.
Frames Starling as the first choice for pound-focused UK startups given its no-fee business account, accounting integrations, and FSCS protection, while suggesting international-heavy businesses may need a specialist provider.
Emphasizes Starling's full UK banking licence with FSCS protection and free, top-rated business account, while consistently noting its weaker multi-currency support and less startup-specific/VC-oriented features versus rivals.
Points to Starling as a fully licensed UK bank with FSCS protection and a free, well-reviewed business account, but ranks it just behind rivals due to less startup-specific tooling and weaker international/multi-currency capabilities.
Consistently describes Starling as an FCA-regulated, fee-free business account with fast payments, strong API/integrations, and an emphasis on transparency and customer support for startups.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
starlingbank.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named Revolut Business
Revolut scores a Borda average of
Stat unavailableand is ranked by all 6 of the 6 models, placing it firmly among the recommended options. The report labels the per-model spread of 50–100 "broad agreement," though the models do diverge on how high to place it: mistral-medium-3.5 ranked it first in all 4 of its runs (borda 100), while gpt-6-astra sat it lowest with a median rank of 3.5 (borda 50).
Revolut · 50 points apart on a 0–100 scale
The theme uniting the models is Revolut's strength for internationally oriented startups — multi-currency accounts, competitive FX, team cards and API-driven expense management recur across every reasoning. The reservations are equally consistent: several models flag the absence of full FSCS deposit protection, the still-maturing UK banking licence in its mobilisation phase, and recurring reports of account freezes and slow support. gpt-6-astra adds that plan fees and usage allowances make it less compelling for a mainly domestic startup. These themes lean on recalled sources including Trustpilot×10 for support and reliability signals, Siftedsifted.eu×7 and TechCrunch×3 for sector context, with Financial Times coverage and the FCA Register cited on the licence question.
“Revolut is the strongest option for startups with international operations, offering multi-currency accounts, near-interbank FX, bulk payments, expense management and a solid API.”
“For a mainly domestic startup, plan fees and usage allowances make it less compelling than the higher-ranked options; check the UK account's current legal provider and deposit-protection terms before holding substantial cash”
Revolut is presented as a top choice for UK startups, emphasizing multi-currency support, competitive FX, expense management, and accounting-tool integrations suited to scaling businesses.
Revolut excels for startups with international/cross-border operations via multi-currency accounts and competitive FX, but is held back by the lack of full UK FSCS deposit protection and support/account-freeze issues.
Revolut is positioned as the strongest option for internationally operating startups (multi-currency, interbank FX, strong API), with its UK banking licence in mobilisation improving credibility, but weighed down by expensive tiers, weak support, and account freezes.
Revolut is favored for multi-currency accounts and international payments, but repeatedly flagged for lacking FSCS protection (e-money status) and inconsistent customer support.
Revolut is the strongest choice for internationally ambitious startups (multi-currency, FX, team cards, API), but its recently granted, restricted UK banking licence in mobilisation plus account freezes and slow support add operational risk.
Revolut is attractive for startups paying overseas suppliers or managing multiple currencies, but less compelling for domestic-focused startups due to plan fees and allowances, with repeated advice to verify the legal entity and deposit protection.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Tide scores 65 of 100 on Borda and is ranked by all 6 of the 6 models, but the recommendation is far from unanimous. Per-model scores span 45 to 100 with a standard deviation of 19.15, which the report labels "models split." That gap traces a clean divide between claude-opus-5, which ranked Tide first in all four of its runs (borda 100), and the cluster of gpt-6-astra (50) and deepseek-v4.1-flash (45), which consistently placed it fourth or lower.
Tide · 55 points apart on a 0–100 scale
The theme uniting every model is that Tide is purpose-built for UK small businesses and startups — fast onboarding, invoicing, expense management, company formation and accounting integrations recur across all six. Where they diverge is on how much its e-money status weighs against it: claude-opus-5 and kimi-k3 treat the ClearBank partnership and FSCS protection as reassurance, while deepseek-v4.1-flash reads the same structure as a lack of protection and gpt-6-astra and gemini-3.7-flash dock it for transaction fees and paid extras as activity scales. Recalled sources concentrate on Trustpilotuk.trustpilot.com/review/tide.co×6 and Tide's own site, with comparison write-ups from Startups.co.uk×4 and MoneySavingExpert supporting the feature and pricing claims.
“Tide is purpose-built for UK small businesses and startups, with free company incorporation bundles, fast onboarding, invoicing, expense cards and accounting integrations.”
“It sits at rank four because per-transaction fees and a leaner feature set make it less attractive once a startup begins to scale.”
Tide is consistently described as purpose-built for UK small businesses and startups, with free company incorporation, fast onboarding, invoicing/expense tools, accounting integrations, and the largest UK SME challenger market share. The recurring caveat is that it is an e-money account rather than a full bank (FSCS protection via ClearBank).
Tide is repeatedly described as purpose-built for UK SMEs and startups with fast onboarding, invoicing, expense management and accounting integrations, noting funds held via ClearBank give FSCS protection, though it is an e-money platform and fees limit it as a startup scales.
Tide is consistently characterized as tailored for small businesses and startups with quick setup and invoicing tools, valued for simplicity and affordability but noted as lacking some advanced or scalable features of competitors.
Tide is consistently framed as tailored specifically for UK small businesses and sole traders with rapid onboarding and built-in bookkeeping/invoicing tools, but ranked lower due to transaction fees and its status as an e-money institution relying on partner banks.
Tide is consistently positioned as strong for business administration (invoicing, bookkeeping, company formation) but ranked below Starling and Monzo because of transaction charges and paid extras, with emphasis that it is a financial platform rather than a bank itself.
Tide is portrayed as a fast, low-cost account popular with small businesses offering invoicing and expense features, but repeatedly flagged as an e-money institution without FSCS protection, best suited to very early-stage needs rather than a scaling primary account.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
no URL recalled
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named Monzo Business
Monzo lands mid-table with a Borda score of 51.67 of 100, ranked by all 6 of the 6 models but with notable disagreement on how highly to place it. The per-model span runs from 30 to 80 with a standard deviation of 18.41, which the report labels "models split."
per-model scores 30–80 of 100 · mean 52 across 6 models
The divide tracks a consistent theme: every model credits Monzo as a fully licensed, FSCS-protected UK bank with a polished app and useful touches like tax pots and invoicing, but they weigh its narrower business feature set differently. gpt-6-astra positions it as a close second to Starling and scored it 80, while claude-opus-5 sees the thinner multi-currency, FX and integration offering as disqualifying for scaling startups and scored it 30. Between them sit deepseek-v4.1-flash at 70 and the more skeptical kimi-k3 (35) and mistral-medium-3.5 (40), the latter pair emphasising that Monzo suits solo founders and micro-startups more than growing teams. No model ranked it first in any run. The recalled sources cluster around review aggregators and Monzo's own pages, with Trustpilotuk.trustpilot.com/review/monzo.com×5 and Which?www.which.co.uk×1 among those cited under these themes.
“Monzo is a close second for a small, founder-led startup because of its straightforward app, spending visibility and useful business administration features.”
“It's ranked last here because its feature set is thinner for scaling startups — weaker multi-currency/FX, fewer accounting and API integrations, and the Pro tier costs extra for capabilities rivals include free.”
Monzo is portrayed as a fully licensed, FSCS-protected bank with an intuitive app, tax pots and easy setup ideal for small founder-led teams, but with fewer features and integrations than Starling or Tide and paid tiers for fuller functionality.
Monzo is consistently positioned as a close second to Starling, praised for its intuitive app and spending visibility with FSCS protection, but with useful business features locked behind paid plans, prompting a cost comparison against Starling.
Monzo is framed as a fully licensed, FSCS-protected UK bank with a polished, award-winning app and features like tax pots and invoicing, but ranked behind Starling because multi-user access and advanced tools sit behind paid subscription tiers.
Monzo is described as a licensed UK bank with excellent UX and features like tax pots that suit solo founders and early-stage teams, but with a thinner, less mature business feature set than Tide or Starling that limits its fit for scaling or international startups.
Monzo is characterized as user-friendly and well-integrated, ideal for sole traders and micro-startups with basic needs, but lacking advanced business features like multi-currency support compared to Revolut or Starling.
Monzo is consistently described as a fully licensed, FSCS-protected UK bank with an excellent app and features like tax pots and invoicing, but with a narrower business feature set (limited multi-currency/FX, fewer integrations) that suits UK-only solo founders and small teams over scaling startups.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named Wise Business
Wise scores 16.67 of 100 on Borda, ranked by 5 of the 6 models but never placed first by any of them. The agreement here is unusually tight: per-model Borda spans 0–20 with a standard deviation of 7.45, which the report labels "models agree." Every model that ranked it landed on the same conclusion, differing only in how far down the list they placed it — kimi-k3 gave it a median rank of 4, while claude-opus-5, deepseek-v4.1-flash, gemini-3.7-flash and gpt-6-astra all settled on a median rank of 5.
per-model scores 0–20 of 100 · mean 17 across 6 models
The theme behind that consensus is consistent across all five: Wise is valued for low-cost international payments, mid-market FX rates and multi-currency holding, but is treated as a complement rather than a primary account because it is an e-money or payments institution lacking FSCS protection, credit and overdrafts. Recalled sources cluster around user-review and comparison material, with Trustpilot×10 appearing most often across models, alongside the Wise Businesswise.com/gb/business×5 official site. The distinction models draw is not about quality but role — a specialist treasury tool paired with a UK business bank rather than the bank itself.
“Best used as a complement rather than the primary account.”
“For most UK startups, it serves best as a secondary treasury tool alongside a primary UK clearing bank.”
Wise is described as outstanding for low-cost international payments and multi-currency accounts, but ranked last as a primary bank because it is an e-money institution without FSCS protection or credit products, best used as a complement.
Wise is positioned as an excellent supplementary account for cross-border payments and multi-currency holding, but ranked lowest as a primary option because it is not a bank—lacking FSCS protection, credit, and overdrafts.
Wise is presented as a strong companion account for international transfers and multi-currency holding, but not a full UK bank—no FSCS protection or lending—so it is best used as a secondary tool rather than a primary account.
Wise is highlighted for mid-market FX rates and multi-currency international payments, but as an electronic money/payment institution rather than a full bank it lacks credit facilities and FSCS coverage, making it a complement to a primary domestic bank.
Wise is valued for international revenue collection and overseas payments with transparent pricing, but ranked fifth as a primary account because it is a payments provider with safeguarded rather than FSCS-protected balances, best paired with a UK business bank.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
wise.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
where
where
6 AI models · 10 answers per model · ranked by score
ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0
Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.
Swipe left for the full report
Median rank each model gave each brand.
Scroll the table sideways to see every model →
| claude-opus-5 | deepseek-v4.1-flash | gemini-3.7-flash | gpt-6-astra | kimi-k3 | mistral-medium-3.5 | |
|---|---|---|---|---|---|---|
| Airtable | 3.0 | 2.0 | 1.0 | 1.0 | 1.0 | 1.0 |
| Microsoft | 2.0 | 1.0 | 3.0 | 3.0 | 3.0 | — |
| Zapier | 4.0 | 5.0 | 3.0 | — | 2.0 | 2.0 |
| Retool | 1.0 | — | 4.0 | — | 4.5 | 3.0 |
| Glide | — | — | 2.0 | 2.0 | 4.0 | — |
| AppSheet | — | 2.5 | 5.0 | 2.5 | 5.0 | 4.0 |
| Quickbase | — | 4.0 | 4.0 | 5.0 | 5.0 | — |
| Make | — | 5.0 | 4.0 | — | 5.0 | — |
| Softr | — | — | 5.0 | 4.0 | 4.0 | — |
| Kissflow | — | 5.0 | — | 4.0 | — | — |
| Bubble | — | — | 5.0 | — | — | 5.0 |
| Appsmith | 5.0 | — | — | — | — | — |
| Smartsheet | — | 5.0 | — | 5.0 | — | — |
| — | — | — | 3.0 | — | — | |
| Monday.com | — | 4.0 | — | — | — | 4.5 |
| Zoho Creator | — | 5.0 | — | — | — | — |
A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.
5 of 16 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.
deepseek-v4.1-flash 0 → claude-opus-5 100 · 2 of 6 never named it
mistral-medium-3.5 0 → deepseek-v4.1-flash 94 · 1 of 6 never named it
claude-opus-5 0 → gpt-6-astra 68 · 3 of 6 never named it
gpt-6-astra 0 → kimi-k3 74 · 1 of 6 never named it
claude-opus-5 0 → deepseek-v4.1-flash 52 · 1 of 6 never named it
claude-opus-5 62 → mistral-medium-3.5 100
A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.
Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.
The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.
Airtable leads with a score of 90 of 100, ranked by 6 of 6 models.
The models divide sharply over which platforms best serve companies building internal workflows without developers, with disagreement centered on whether "no developers" means zero technical literacy or simply no dedicated engineering staff. That distinction drives the biggest splits in the data.
Airtable earns the clearest consensus, appearing in every model's top three and securing a Borda score of 89.67. Four models award it a perfect 100 and place it first, citing its spreadsheet-familiar interface, shallow learning curve, and ability to ship operational trackers and approval workflows in hours. Yet two models rank it second or third, stressing that it "can get expensive and hit performance limits at high record volumes or complex logic." The agreement is broad but not absolute—every model accepts Airtable for departmental use, but assessments diverge on how far those use cases stretch before requiring something heavier.
per-model scores 62–100 of 100 · mean 90 across 6 models
Microsoft provokes the sharpest split in the category. Deepseek-v4.1-flash ranks Power Platform first with a Borda of 94, emphasizing native Office 365 integration, enterprise governance, and Dataverse. At the other end, gemini-3.7-flash scores it at only 28, and three other models place it third. The divide turns on whether existing Microsoft investment and IT oversight are assets or barriers: models citing Gartner Magic Quadrant for Enterprise Low-Code Application Platforms×10 frame governance and security as strengths for large organizations, while others flag licensing complexity and formula-based customization as pushing citizen developers back toward IT support. Kimi-k3 sums up the tension, noting Power Platform's "steeper learning curve, confusing licensing tiers, and technical complexity often require citizen developer champions or semi-technical staff, which cuts against the pure no-code brief."
Microsoft · 94 points apart on a 0–100 scale
Retool triggers a similar but even more dramatic split. Claude-opus-5 ranks it first with a Borda of 100, calling it "the strongest option for building internal tools and workflow apps quickly on top of existing databases and APIs." Two other models omit it entirely, yielding a standard deviation above 38 points. The core disagreement is explicit: models valuing SQL and JavaScript access as power-user features rank Retool highly, while those interpreting the question as requiring zero scripting ability dismiss it. Kimi-k3 acknowledges it as "arguably the most powerful internal-tool builder here" yet notes "real value typically requires some SQL or JavaScript, so it fits a company 'without developers' poorly."
Retool ahead on 2 of 6 models
Zapier divides the panel more subtly. Two models place it second, praising its 6,000-plus integrations and the addition of Tables, Interfaces, and Canvas for lightweight apps. Two others rank it fourth or fifth, arguing it "excels more at connecting systems than serving as a full internal app platform" and warning that task-based pricing "can become costly as task volume grows." The split reflects whether automation breadth compensates for weaker data modeling and UI capabilities. Sources cluster around G2www.g2.com/products/zapier/reviews×17 and the Zapier site itself, reinforcing both its leadership in workflow glue and its secondary status as an app builder.
Glide, AppSheet, and Quickbase show narrower but meaningful disagreement. Glide earns second-place ranks from gemini-3.7-flash and gpt-6-astra, which highlight its "exceptionally fast" spreadsheet-to-app conversion and mobile polish, but kimi-k3 places it fourth, noting "complex logic, permissions, and large-scale data needs can push it past its limits." Three models omit Glide entirely. AppSheet's deepseek score of 52 contrasts with kimi's 2, a split likely tied to Google Workspace penetration and tolerance for expression syntax. Quickbase appears in four models but never first, with all citing strong governance and relational data yet flagging dated UX and enterprise pricing as barriers for smaller or casual teams.
The long tail of brands—Make, Softr, Kissflow, Bubble, Appsmith, Smartsheet, Monday.com, and Zoho Creator—shows broad agreement on irrelevance or narrow fit. Make, Softr, and Bubble each appear in three models or fewer, and none breaks a Borda of 6. Models converge on reasons: Make is "powerful visual automation" but "steeper learning curve"; Softr is "one of the fastest ways to turn Airtable…into client portals" yet "relies heavily on external databases"; Bubble offers "unmatched visual design flexibility" but "may be overkill for simple workflow automation." These platforms occupy specialist lanes—complex branching logic, portal frontends, full-stack custom apps—that models judge peripheral to the internal-workflow mandate.
Across the category, G2www.g2.com/categories/no-code-development-platforms×46 is the most frequently recalled source, appearing 46 times, followed by product-specific G2 pages for Airtable, Zapier, and Microsoft. Gartner and Forrester analyst research cluster around Microsoft and Quickbase, reinforcing enterprise positioning, while Product Hunt and Hacker News references appear primarily for developer-adjacent tools like Retool and Appsmith. The hypothesis that sources drive model splits is plausible but not demonstrated: deepseek's preference for Microsoft and claude's for Retool both draw on similar G2 and vendor-documentation pools, suggesting the divergence reflects weighting of technical complexity over raw capability rather than distinct source sets.
79 sources across 828 references, grouped by site from 254 recalled names. The top 5 carry 53% of them.
g2.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Airtable commands the strongest consensus in the category, earning a Borda score of 89.67 and ranking first among four of the six models. Every model placed it in its top three, yet the span of scores is wide: four models awarded it a perfect 100, while one scored it at 62, yielding broad agreement overall with a standard deviation above 15 points. The core narrative is consistent across all six: Airtable delivers an intuitive spreadsheet-database hybrid that non-technical teams can extend into custom interfaces, automations, and operational workflows without developer involvement. Models cite its shallow learning curve, rich template library, and broad integrations as key reasons business users can ship real internal tools in hours or days. Sources cluster around review platforms—G2www.g2.com/products/airtable/reviews×28, Capterra×19, and TechCrunch×17—alongside Airtable's own documentation and product pages, reinforcing a picture drawn from both user feedback and vendor materials.
Where models diverge is on the severity of Airtable's limits at scale. The four that rank it first emphasize its balance of power and accessibility for departmental use, while the two that place it second or third stress constraints around enterprise governance, complex permissions, very large data volumes, and intricate business logic. Claude-opus-5, which ranks it third, notes that teams "can get expensive and hit performance limits at high record volumes or complex logic," and deepseek-v4.1-flash cautions it "may hit limits for very complex enterprise-scale apps." No model questions its suitability for operational trackers, approvals, or project workflows—the disagreement centers on how far those use cases can stretch before requiring a more robust platform.
“Airtable hits the sweet spot for non-technical teams: a spreadsheet-familiar interface on top of a real relational database, plus built-in automations, forms, and an interface designer for turning data into usable internal apps.”
“Airtable is my strongest general recommendation for nontechnical teams because it combines familiar spreadsheet-style data management with custom interfaces, forms, and workflow automation.”
Airtable is consistently praised for combining an intuitive spreadsheet-like interface with powerful relational database capabilities, Interface Designer, and built-in automations that empower non-technical teams to quickly deploy sophisticated internal tools without code.
Airtable is repeatedly positioned as the strongest general-purpose recommendation that balances a familiar spreadsheet-style interface with custom interfaces and automations for building internal tools, though complex permissions, business logic, and higher usage may require more expensive plans or technical assistance.
Airtable is consistently described as hitting the optimal balance of accessibility and power, combining a spreadsheet-familiar interface with relational database capabilities, templates, and integrations that enable non-technical teams to ship working internal apps in hours or days without engineering support.
Airtable is repeatedly characterized by its flexibility and ease of use, combining an intuitive spreadsheet-like interface with database power, automation capabilities, and strong integrations to enable custom workflows and internal tools without code.
Airtable is repeatedly described as the most approachable and fastest way for non-developers to build internal apps and workflows, excelling at departmental use cases and quick iteration, but facing challenges with enterprise-scale governance and very complex processes.
Airtable is consistently characterized as the easiest entry point for non-technical teams, combining a familiar spreadsheet-database interface with interfaces, forms, and automations that enable rapid deployment of operational workflows, though it faces limitations with very large data volumes, complex logic, and granular permissions at scale.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
capterra.com
airtable.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named Microsoft Power Apps, Microsoft Power Platform, Power Apps
Microsoft ranks as a moderately strong recommendation in the no-code space, earning a Borda score of 49 out of 100, but the models are sharply split on its placement. Five of the six models included Microsoft Power Platform in their rankings, yet their scores span from 94 down to 28, reflecting fundamentally different assessments of how well the platform serves companies seeking to build internal workflows without developer help. The central tension is consistent across nearly every response: Microsoft's deep native integration with the Office 365 and Teams ecosystem, coupled with enterprise-grade governance and security controls, makes it a natural and often cost-effective choice for organizations already standardized on Microsoft tooling. However, the same models repeatedly flag licensing complexity, a steeper learning curve, formula-based customization requirements, and governance overhead that can push citizen developers back toward IT support—qualities that cut against the pure no-code brief.
per-model scores 0–94 of 100 · mean 49 across 6 models
The sources models cited cluster heavily around analyst research and official documentation. G2×9 and Gartner Magic Quadrant for Enterprise Low-Code Application Platforms×10 together account for the majority of recalled references, emphasizing Microsoft's enterprise credentials and market-leader status. Models also drew extensively on Microsoft Power Platformpowerplatform.microsoft.com×5 and Microsoft Learn pages to validate technical capabilities around Power Apps, Power Automate, and Dataverse. The recurring theme is that Microsoft excels when existing organizational investment, IT governance needs, and Microsoft 365 integration matter most—but that same technical depth can require semi-technical champions or "citizen developer" programs to unlock value, making it less approachable than simpler alternatives for teams truly working without any developer support.
“Power Apps is the strongest fit for enterprise internal workflows because it combines no-code building with deep Microsoft 365 integration, Dataverse, and Power Automate, while giving IT governance and security controls.”
“Power Platform offers the deepest enterprise capability with strong native Microsoft 365 integration, governance, and security that many companies already license, making it especially natural for Microsoft-centric organizations. However, its steeper learning curve, confusing licensing tiers, and technical complexity often require citizen developer champions or semi-technical staff, which cuts against the pure no-code brief.”
Power Apps is the strongest fit for enterprise internal workflows because it combines no-code building with deep Microsoft 365 integration, Dataverse, and Power Automate, while giving IT governance and security controls. The platform enables citizen developers to build apps while maintaining enterprise compliance, though it can involve licensing complexity and some learning curve.
Power Apps plus Power Automate is the default choice for Microsoft 365 organizations because of deep native integration with SharePoint, Teams, and Dataverse, along with enterprise-grade governance and compliance controls. The main trade-offs are a steeper learning curve and more complex licensing compared to lighter no-code tools.
Microsoft Power Platform provides enterprise-grade capabilities with deep native integration across the Microsoft 365 ecosystem, allowing non-developers to build secure applications using familiar organizational data sources. The platform's administrative setup, governance controls, and licensing model can present a steeper learning curve compared to standalone alternatives.
Power Apps and Power Automate are compelling for companies already using Microsoft 365 due to extensive integrations and enterprise governance, but formulas, licensing complexity, environment administration, and customization requirements can make truly developer-free implementation harder as applications grow. The platform ranks lower specifically for the requirement to work without developers, though existing Microsoft environment and expertise could improve its position.
Power Platform offers the deepest enterprise capability with strong native Microsoft 365 integration, governance, and security that many companies already license, making it especially natural for Microsoft-centric organizations. However, its steeper learning curve, confusing licensing tiers, and technical complexity often require citizen developer champions or semi-technical staff, which cuts against the pure no-code brief.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
g2.com
learn.microsoft.com
no URL recalled
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Zapier scores a Borda of 39.33 across the panel, ranked by five of the six models but with considerable disagreement on where it belongs: per-model scores range from 0 to 74, yielding a standard deviation of 27.92. The strongest supporters place it consistently at rank two, praising its unmatched library of integrations—often cited as 6,000 to 8,000 apps—and its recent expansion into lightweight app-building through Tables, Interfaces, and Canvas. These models, drawing frequently from G2×9 and Zapierzapier.com×17, frame Zapier as the automation benchmark and an increasingly viable option for simple internal tools. At the other end, two models rank it fourth or fifth when they include it at all, arguing that its core strength remains cross-app workflow orchestration rather than full application development, and that task-based pricing can become prohibitive at scale.
per-model scores 0–74 of 100 · mean 39 across 6 models
The central tension in the assessments is whether Zapier qualifies as a true internal-app platform or remains primarily an integration layer. Models that rank it higher acknowledge the addition of database and interface features but still note it "excels more at connecting systems than serving as a full internal app platform." Those placing it lower emphasize that "it is not a full app builder" and works best alongside dedicated platforms like Airtable or Retool. Even enthusiastic rankings carry the caveat that complex data modeling and rich UIs are not its natural territory, a view reinforced by citations to review sites such as Capterra×19 and workflow-automation category pages on G2. The result is a brand recognized universally for automation breadth yet assessed quite differently depending on whether the question prioritizes workflow glue or complete application construction.
“Zapier remains the broadest integration layer, with thousands of app connectors plus Tables, Interfaces, and AI agents that now cover simple internal apps as well as automation. It's ideal for stitching SaaS tools together without engineering involvement, but it's weaker than dedicated app builders for data-heavy UIs and can become costly as task volume grows.”
“Zapier is the benchmark for no-code workflow automation, connecting thousands of apps so non-developers can automate cross-tool processes in minutes, and its Tables and Interfaces features now cover lightweight apps too. It ranks mid-list because it is primarily an automation layer between tools rather than a full internal app builder, and costs scale quickly with task volume.”
Zapier is the benchmark for workflow automation across thousands of apps with expansion into lightweight app-building via Tables and Interfaces, though it excels more at connecting systems than serving as a full internal app platform, with per-task pricing that can scale quickly.
Zapier is positioned as the leader in workflow automation with thousands of integrations for streamlining repetitive tasks, though it lacks full app-building capabilities and is best suited for process automation rather than building complete applications.
Zapier is the industry standard for connecting disparate applications and automating workflows with the largest integration library, now expanding into app-building with Tables and Interfaces, though its core strength remains event-driven automation rather than complex UI-heavy applications.
Zapier is positioned as the leading cross-app automation platform with the largest connector library, now extended with lightweight app-building features (Tables, Interfaces, Canvas), though it remains automation-first rather than a true app platform and task-based pricing can become expensive at scale.
Zapier is recognized as a workflow automation and integration layer for connecting SaaS apps rather than a full app-building platform, working best as a complement to other no-code tools.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
g2.com
zapier.com
g2.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Retool registers sharply split opinions across the panel, scoring 31.67 Borda points with a standard deviation of 38.62. Four of the six models ranked it, but their placements diverge dramatically: one model consistently names it first, valuing its pre-built components and database connectivity for internal tools, while others place it fourth or fifth, citing its low-code nature as a poor fit for teams without any technical support. The core tension is whether "no developers" means zero technical literacy or simply no dedicated engineering staff. Models that interpret the brief strictly penalize Retool for expecting SQL or JavaScript familiarity, whereas those allowing for "semi-technical" or operations users with light scripting skills rate it as the category leader. Coverage sits at only 67 percent of the panel, and two models omit it entirely, reinforcing the platform's niche appeal.
per-model scores 0–100 of 100 · mean 32 across 6 models
Source citations cluster around three themes. Reviews and community discussion—G2×9, Hacker News×2, and Gartner Peer Insights—validate Retool's strength in internal tooling but also surface the learning curve. Vendor documentation and announcements from the Retool site itself and TechCrunch illustrate feature depth and growth, while Product Hunt references highlight early adopter enthusiasm. The consensus in the reasonings is that Retool is "arguably the most powerful purpose-built tool for internal apps" yet "leans low-code rather than pure no-code," a duality that drives both its top rank among technical evaluators and its relegation by models prioritizing ease of use for fully non-technical teams.
“Retool is the strongest option for building internal tools and workflow apps quickly on top of existing databases and APIs, with a drag-and-drop UI builder plus optional code for power users.”
“Retool is arguably the most powerful internal-tool builder here, with deep data connectivity and a huge component library. The catch is that it leans low-code rather than no-code: real value typically requires some SQL or JavaScript, so it fits a company 'without developers' poorly.”
Retool is the most capable platform for building internal tools quickly with drag-and-drop components and deep database/API connectivity. It scales from no-code usage to optional JavaScript/SQL for power users, though it performs best when a semi-technical builder is involved.
Retool is designed specifically for building internal tools quickly with pre-built components and deep database/API integrations. It offers high customization and is developer-friendly, though it may require some technical knowledge or familiarity to unlock its full potential.
Retool is a leading platform for building operational tools and dashboards connected to databases and APIs. It leans low-code rather than pure no-code and often requires SQL or scripting familiarity for maximum utility, which may challenge fully non-technical teams.
Retool is arguably the most powerful purpose-built tool for internal apps with deep data connectivity and rich components. It ranks lower for strictly non-developer audiences because it leans low-code, typically requiring SQL or JavaScript to extract real value.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
g2.com
g2.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Glide earns a Borda score of 28.33 across the panel, with only three of the six models ranking it at all—a pattern the data labels as "sharply split." Among the models that do include it, opinions range from consistent second-place rankings (gemini-3.7-flash and gpt-6-astra both score it at 68) down to a more cautious fourth or fifth spot (kimi-k3 at 34). The models agree that Glide excels at converting spreadsheets and databases into polished, mobile-friendly internal apps with minimal learning curve, often citing its drag-and-drop interface and pre-built components as ideal for non-technical teams building directories, inventory trackers, and field tools. Sources cluster around platform reviews on G2×9 and Product Hunt×12, alongside the platform's own documentation. The split emerges when models weigh simplicity against power: those ranking Glide higher emphasize speed and accessibility for straightforward use cases, while the more reserved assessment flags limitations in complex workflow orchestration, enterprise-grade permissions, and scaling multi-tiered data relationships.
The disagreement centers on whether Glide's design-first, rapid-deployment strengths outweigh its constraints in advanced scenarios. Models that place it near the top see it as "exceptionally fast" and requiring "zero coding knowledge," well-suited to departmental tools that prioritize polish over deep logic. The more cautious view acknowledges the same ease of use but notes that "complex logic, permissions, and large-scale data needs can push it past its limits" and that "its automation depth and ability to handle complex logic lag behind the leaders." Three models declined to rank Glide at all, reinforcing a perception that it occupies a narrower niche—brilliant for simple, fast internal apps, but less versatile than platforms built for heavier workflow or governance demands.
“Glide excels at turning existing business data from spreadsheets and databases into polished, responsive web apps in minutes. Non-technical staff can rapidly create internal portals, inventory trackers, and team directories with pre-built components.”
“Glide is especially strong for nontechnical teams building polished, mobile-friendly internal apps like directories and inventory tools through an approachable visual builder. However, usage limits, complex workflow orchestration, integration requirements, and enterprise-level customization should be carefully evaluated before company-wide adoption.”
Glide excels at rapidly converting existing business data from spreadsheets and databases into polished, responsive web and mobile apps through an intuitive drag-and-drop interface requiring zero coding knowledge. Its pre-built components enable exceptionally fast deployment of internal tools, though it faces constraints with highly complex enterprise logic and multi-tiered data relationships.
Glide is especially strong for nontechnical teams building polished, mobile-friendly internal apps like directories and inventory tools through an approachable visual builder. However, usage limits, complex workflow orchestration, integration requirements, and enterprise-level customization should be carefully evaluated before company-wide adoption.
Glide offers the fastest, most accessible path to turn spreadsheets into polished, mobile-friendly internal apps with a gentle learning curve ideal for simple operational tools. Its design-first approach produces excellent results for straightforward use cases, but complex logic, granular permissions, enterprise governance, and scaling needs expose its limitations compared to more powerful platforms.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
where
where
Tell us what to ask and how wide to sample it — we run it and send back the report.