Your Brand has an AI reputation and it's shaping purchase decisions right now.

Swipe up to continue

Category · Consulting companies for AI transformation

Which consulting firm should a Fortune 500 company hire for an AI transformation?

6 AI models · 5 answers per model · ranked by score

ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0

  1. #1Accenture
    91
    brand score
    82first choice
    12345
  2. #2McKinsey & Company
    79
    brand score
    60first choice
  3. #3Boston Consulting Group
    53
    brand score
    32first choice
  4. #4Deloitte
    51
    brand score
    32first choice
  5. #5IBM Consulting
    15
    brand score
    13first choice
Show 1 more
  1. #6Bain & Company
    11
    brand score
    8first choice

Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.

Swipe left for the full report

Category · Consulting companies for AI transformation

Model by model

Median rank each model gave each brand.

Scroll the table sideways to see every model →

claude-opus-5gemini-3.7-flashgpt-5.6-lunagpt-5.6-solgpt-5.6-terrakimi-k3
Accenture2.03.01.01.01.01.0
McKinsey & Company1.01.02.03.02.02.0
Boston Consulting Group4.02.04.04.04.03.0
Deloitte3.05.03.02.03.04.0
IBM Consulting5.05.05.05.05.05.0
Bain & Company4.05.05.04.0
1st
2
3
4
5th

A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.

Category · Consulting companies for AI transformation

Where the models disagree

0 of 6 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.

  1. #4Deloitte

    gemini-3.7-flash 28 gpt-5.6-sol 72

  2. #2McKinsey & Company

    gpt-5.6-sol 52 gemini-3.7-flash 100

  3. #1Accenture

    gemini-3.7-flash 60 gpt-5.6-terra 100

  4. #3Boston Consulting Group

    gpt-5.6-terra 40 gemini-3.7-flash 80

  5. #6Bain & Company

    claude-opus-5 0 gemini-3.7-flash 28 · 2 of 6 never named it

  6. #5IBM Consulting

    gemini-3.7-flash 4 gpt-5.6-sol 28

A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.

Category · Consulting companies for AI transformation

Named, and named first

Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.

the default answera narrow favouritelisted, rarely ledthe long tail123456
  1. 1Accenture100/70
  2. 2McKinsey & Company100/30
  3. 3Boston Consulting Group100/0
  4. 4Deloitte100/0
  5. 5IBM Consulting63/0
  6. 6Bain & Company37/0

The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.

Category · Consulting companies for AI transformation

What the models actually said

Models
6
Answers
30

Accenture leads with a score of 91 of 100, ranked by 6 of 6 models.

Across the six models, the top of the field is settled and the disagreements concentrate in the middle. Accenture and McKinsey trade the top two spots, while Boston Consulting Group, Deloitte, IBM Consulting, and Bain sort into a more contested lower tier where individual models diverge noticeably.

Where the models agree

The clearest consensus is on the extremes. Accenture holds a category-wide median of 1 and IBM Consulting a rock-solid median of 5 (Q1 5, Q3 5) with every model landing on exactly 5. IBM's uniformity extends to reasoning: all six credit engineering depth, watsonx, and regulated-environment fit, and five of the six cite vendor bias toward IBM's own stack as the reason it sits mid-tier. There is little to separate the models on either brand.

McKinsey is also broadly agreed on in substance — C-suite strategy credibility and QuantumBlack depth recur everywhere — even though its placement varies by a rank or two.

The main splits

The sharpest divergences are best read model by model:

  • gemini-3.7-flash is the standout dissenter. It ranks Accenture lowest of any model (median 3 vs. 1 for four peers) and BCG highest (median 2), while pushing Deloitte to the bottom (median 5). Its framing of Accenture as an "integration and infrastructure play rather than a strategy leader" is consistent with this ordering, elevating the strategy-first houses over the large integrator.
  • gpt-5.6-sol is the mirror image on Deloitte, ranking it highest of all models (median 2) but placing McKinsey lowest (median 3, Q1 3, Q3 4). It reads Deloitte's regulated-enterprise breadth favorably while adding the most caveats to both Accenture (scope, cost, complexity) and McKinsey (delivery bench).
  • claude-opus-5 ranks McKinsey top (median 1) and Accenture second, the reverse of the four gpt/kimi models, framing Accenture as execution-led rather than a strategy leader.
  • kimi-k3 is the most skeptical of Deloitte among the mid-rankers (median 4), recasting its breadth as "generalist and integration-led."

BCG shows the widest disagreement of any brand: gemini places it 2nd, kimi 3rd, and the remaining four settle at 4. Deloitte spans an even wider model range (2 through 5) despite a tidy 3.5 category median — a case where the aggregate obscures real disagreement.

Plausible source patterns behind the splits

The source data is suggestive rather than conclusive, and these links are hypotheses.

The most legible pattern is on Accenture: models ranking it 1st lean on first-party material (Accenture service pages, Technology Vision, the $3B investment newsroom item), whereas gemini-3.7-flash — the lowest ranker — is described as drawing almost entirely on external outlets (Gartner, IDC MarketScape, Bloomberg, FT, WSJ) with no first-party Accenture citations. It is plausible that a reliance on outside analyst and press framing produced its more measured, integrator-oriented view, but the data shows correlation, not cause.

A similar hypothesis fits Deloitte: gemini is again noted as leaning more on Gartner, IDC MarketScape, Bloomberg, FT, and Fortune, coinciding with its lowest placement — consistent with a pattern of external sourcing tracking cooler positioning, though the sample is too small to confirm.

For McKinsey and BCG, the recalled sources are heavily first-party (State of AI, QuantumBlack; BCG X, AI Radar) across models regardless of rank, so the sourcing does not obviously explain why claude and gemini rate them higher than the gpt models. Here the placement differences appear to come from how each model weighs strategy versus at-scale delivery rather than from distinct source pools.

Caveat

Bain and IBM rest on fewer ranked answers (11 and 19) than the leaders (30), so their model-level medians are less stable. All sources are recalled by the models from memory rather than verified citations, which limits how far any source-to-view link can be pushed.

Category · Consulting companies for AI transformation

What shaped the answers

91 sources across 428 references, grouped by site from 256 recalled names. The top 5 carry 44% of them.

McKinsey State of AI×49 · 11%

mckinsey.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Consulting companies for AI transformation
#1

Accenture

91
brand score
82first choice
12345

Summary

Accenture lands at the top of the field, with a median rank of 1 (Q1 1, Q3 2) across 30 ranked answers from 6 models. Agreement is strong at the high end: four models — gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, and kimi-k3 — place it at a median rank of 1 with no spread (Q1 1, Q3 1), while claude-opus-5 sits slightly lower at median 2 (Q1 1, Q3 2) and gemini-3.7-flash is the clear outlier at median 3 (Q1 3, Q3 3). The recurring theme uniting the top rankings is scale of delivery and execution: strategy-through-implementation breadth, cloud and integration depth, a multi-billion-dollar AI investment, and hyperscaler alliances. Several models frame Accenture as the strongest default for moving beyond pilots to enterprise-wide deployment, though most also qualify that its strength skews toward execution rather than top-table strategy — a point noted explicitly by kimi-k3, claude-opus-5, and gemini-3.7-flash, with gpt-5.6-sol adding caveats around scope, cost, and complexity.

The sources behind these themes divide along two lines. Models placing Accenture at rank 1 lean heavily on first-party material — Accenture's own AI and data-AI service pages, Technology Vision 2024, the Reinvention in the Age of Generative AI insight, annual reports, and the $3 billion AI investment newsroom announcement — supplemented by third-party analysts such as Gartner, Everest Group, IDC, Forrester, and the Stanford AI Index. claude-opus-5 grounds its execution-focused view in earnings-release and investor material on generative AI bookings alongside NVIDIA/Microsoft alliance announcements, Gartner's Magic Quadrant, and HFS/Everest Group assessments. By contrast, gemini-3.7-flash, the lowest ranker, draws almost entirely on external outlets — Gartner, IDC MarketScape, Bloomberg, Financial Times, Forrester, and the Wall Street Journal — with no first-party Accenture citations, which aligns with its more measured positioning of the firm as an integration and infrastructure play rather than a strategy leader.

Per model summary

  • gpt-5.6-luna
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Accenture is presented as the strongest overall choice, combining strategy, technology implementation, and managed services at scale, particularly for moving beyond pilots to enterprise-wide deployment.

  • gpt-5.6-sol
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Accenture is framed as the strongest default choice given its broad combination of strategy, integration, cloud, and delivery scale, while noting clients must control scope, cost, and complexity.

  • gpt-5.6-terra
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    Accenture is consistently the strongest all-around choice, pairing board-level strategy with large-scale delivery to move from pilots to enterprise-wide operating-model and workforce change.

  • kimi-k3
    brand score
    96
    first choice
    90
    median rank
    1.0
    present in
    5/5

    Accenture is highlighted for unmatched implementation and delivery scale, multi-billion-dollar AI investment, and key platform alliances, though its brand skews toward execution rather than top-table strategy.

  • claude-opus-5
    brand score
    88
    first choice
    70
    median rank
    2.0
    present in
    5/5

    Accenture is characterized as having the largest scaled AI delivery capability — multi-billion-dollar gen-AI bookings, tens of thousands of AI-trained practitioners, and deep hyperscaler partnerships — making it the strongest choice for end-to-end execution rather than strategy alone.

  • gemini-3.7-flash
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    5/5

    Accenture is described as offering unmatched global scale and systems integration for massive IT and infrastructure overhauls, backed by multi-billion-dollar AI investment, though it leans more toward implementation than pure top-down strategy.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gpt-5.6-luna6 sites · 14 references
Accenture Technology Vision×7 · 50%

accenture.com

gpt-5.6-sol8 sites · 15 references
Accenture Annual Report×8 · 53%

accenture.com

gpt-5.6-terra5 sites · 12 references
Accenture — AI services×7 · 58%

accenture.com

kimi-k38 sites · 15 references
Accenture×4 · 27%

accenture.com

claude-opus-57 sites · 15 references
Accenture + NVIDIA / Microsoft alliance announcements×5 · 33%

newsroom.accenture.com

gemini-3.7-flash7 sites · 15 references
Gartner×4 · 27%

gartner.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Consulting companies for AI transformation
#2

McKinsey & Company

also named McKinsey

79
brand score
60first choice
12345

Summary

McKinsey & Company holds a strong overall position, with a median rank of 2 (Q1 1, Q3 2) across 30 ranked answers from 6 models, and appears under both "McKinsey & Company" and "McKinsey." The models converge on a consistent underlying rationale: McKinsey is valued for C-suite strategy credibility, use-case prioritization, and operating-model redesign, with its QuantumBlack unit repeatedly cited as adding technical and data-science depth. This theme is supported by heavy recall of McKinsey's own research and capability pages — particularly the State of AI report and QuantumBlack insights — alongside third-party analyst and business-press sources such as Gartner, Forrester, Forbes, Harvard Business Review, IDC MarketScape, and coverage of McKinsey's internal Lilli tool (Reuters, CNBC).

Agreement on the reasoning is high, but placement varies more than the median suggests. Gemini-3.7-flash and claude-opus-5 rank it strongest (both median 1, with gemini at Q1 1/Q3 1 and no noted drawbacks), while gpt-5.6-luna, gpt-5.6-terra, and kimi-k3 land at median 2, and gpt-5.6-sol places it lowest at median 3 (Q1 3, Q3 4). The recurring caveat behind the lower placements is consistent across those models: McKinsey is seen as ranking below Accenture or larger systems integrators for hands-on, at-scale implementation and ongoing operations, with premium pricing and a thinner delivery bench sometimes requiring additional implementation partners.

Per model summary

  • claude-opus-5
    brand score
    92
    first choice
    80
    median rank
    1.0
    present in
    5/5

    McKinsey's QuantumBlack unit is framed as combining board-level strategy credibility with a large bench of data scientists, backed by influential research (State of AI), making it best for enterprise-wide value-case definition and operating-model redesign — though at premium cost and often needing implementation partners at scale.

  • gemini-3.7-flash
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    5/5

    McKinsey is consistently presented as pairing world-class C-suite strategy with deep technical execution via QuantumBlack, positioning it as the premier end-to-end choice for large-scale Fortune 500 transformation and change management, with no noted drawbacks.

  • gpt-5.6-luna
    brand score
    76
    first choice
    47
    median rank
    2.0
    present in
    5/5

    McKinsey is described as strong at AI strategy, use-case prioritization, and operating-model change, but consistently ranked below Accenture because large-scale technical implementation and ongoing operations may require systems integrators or other partners.

  • gpt-5.6-terra
    brand score
    72
    first choice
    43
    median rank
    2.0
    present in
    5/5

    McKinsey is framed as excellent for defining the AI agenda, prioritizing use cases, and operating-model/executive alignment, with QuantumBlack adding capability, while consistently cautioning clients to validate hands-on engineering and delivery capacity for large implementations.

  • kimi-k3
    brand score
    84
    first choice
    60
    median rank
    2.0
    present in
    5/5

    McKinsey is characterized by unmatched C-suite credibility, strategy rigor, and QuantumBlack's technical depth plus influential State of AI research, but ranked second to Accenture due to premium pricing and a thinner hands-on, at-scale implementation bench.

  • gpt-5.6-sol
    brand score
    52
    first choice
    32
    median rank
    3.0
    present in
    5/5

    McKinsey is repeatedly cited for executive alignment, operating-model redesign, and tying AI to measurable business value with QuantumBlack support, but ranked below larger systems integrators for implementation-heavy, at-scale programs that may need additional partners.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-58 sites · 15 references
McKinsey State of AI report×8 · 53%

mckinsey.com

gemini-3.7-flash5 sites · 15 references
Forbes×5 · 33%

forbes.com

gpt-5.6-luna3 sites · 14 references
gpt-5.6-terra2 sites · 12 references
McKinsey — The State of AI×11 · 92%

mckinsey.com

kimi-k36 sites · 15 references
McKinsey & Company×8 · 53%

mckinsey.com

gpt-5.6-sol6 sites · 15 references

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Consulting companies for AI transformation
#3

Boston Consulting Group

also named BCG

53
brand score
32first choice
12345

Summary

Boston Consulting Group lands mid-pack across the six models, with a median rank of 3 (Q1 3, Q3 4) over 30 ranked answers. There is moderate agreement on its placement but a visible spread in enthusiasm: gemini-3.7-flash ranks it highest at a median of 2 (Q1 2, Q3 2), kimi-k3 places it at 3 (Q1 3, Q3 3), and the remaining four models — claude-opus-5, gpt-5.6-luna, gpt-5.6-sol, and gpt-5.6-terra — each land at a median of 4, with gpt-5.6-terra the most consistent at that level (Q1 4, Q3 4). The firm is named as both Boston Consulting Group and BCG.

The recurring theme is a strategy-plus-build profile: models consistently credit BCG for elite corporate strategy, operating-model redesign, and value/ROI framing, paired with technical delivery through BCG X, and they cite the firm's research output, including the 10-20-70 framing and OpenAI/Anthropic partnerships. The countervailing theme, and the main reason it ranks behind McKinsey, Accenture, and Deloitte in several answers, is a perceived thinner global implementation, systems-integration, and managed-services footprint for full Fortune 500-scale rollouts. Sources cluster around BCG's own materials — BCG X, BCG's AI capabilities pages, the AI Radar, and the BCG Henderson Institute — supplemented by MIT Sloan Management Review (including joint MIT Sloan–BCG research), Harvard Business Review, and the Harvard Business School–BCG field experiment on GenAI productivity. Analyst and press validation appears through Gartner, IDC MarketScape, the Forrester Wave (AI Consultancies / AI Services), Forbes, the World Economic Forum, and Financial Times and Reuters coverage of BCG's AI revenue share.

Per model summary

  • gemini-3.7-flash
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    5/5

    BCG pairs elite corporate strategy with technical build and engineering through BCG X, emphasizing proprietary domain-specific AI solutions, business value/ROI, and responsible AI governance for large enterprises.

  • kimi-k3
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    5/5

    BCG blends strategy with BCG X build capability and highly cited AI research (e.g., 10-20-70 framing), but ranks behind McKinsey and Accenture because its large-scale global delivery bench is thinner and less proven at full enterprise-wide deployment.

  • claude-opus-5
    brand score
    44
    first choice
    27
    median rank
    4.0
    present in
    5/5

    BCG X provides genuine build capability alongside top-tier strategy, with roughly a fifth of revenue from AI work and notable OpenAI/Anthropic partnerships, but its delivery footprint is smaller than Accenture's or Deloitte's for large-scale integration and long-term run services.

  • gpt-5.6-luna
    brand score
    48
    first choice
    28
    median rank
    4.0
    present in
    5/5

    BCG excels at business strategy, innovation, and business-model redesign around AI, but ranks below implementation-led firms because a Fortune 500 transformation may need more systems integration, engineering capacity, and managed services.

  • gpt-5.6-sol
    brand score
    48
    first choice
    28
    median rank
    4.0
    present in
    5/5

    BCG combines senior-level strategy and operating-model redesign with product and technical build capabilities, strong for high-value differentiated AI use cases, though its global implementation and managed-services scale trails Accenture's and Deloitte's.

  • gpt-5.6-terra
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    5/5

    BCG is compelling for value-led AI strategy, portfolio prioritization, and business-model innovation supported by BCG X, but ranks below the top firms because Fortune 500-scale rollouts often need larger global implementation and managed-delivery depth.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gemini-3.7-flash10 sites · 15 references
Gartner×3 · 20%

gartner.com

kimi-k37 sites · 15 references
Boston Consulting Group×7 · 47%

bcg.com

claude-opus-58 sites · 15 references
BCG X×8 · 53%

bcg.com

gpt-5.6-luna4 sites · 14 references
BCG Artificial Intelligence×10 · 71%

bcg.com

gpt-5.6-sol7 sites · 15 references
BCG Artificial Intelligence×8 · 53%

bcg.com

gpt-5.6-terra1 site · 12 references
BCG — Artificial Intelligence×12 · 100%

bcg.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Consulting companies for AI transformation
#4

Deloitte

51
brand score
32first choice
12345

Summary

Deloitte lands in the middle of the pack overall, with a median rank of 3.5 (Q1 3, Q3 4) across 30 ranked answers from 6 models. Model-level placement spans a moderate range: gpt-5.6-sol ranks it highest at a median of 2 (Q1 2, Q3 2), while claude-opus-5, gpt-5.6-luna, and gpt-5.6-terra cluster at a median of 3, and kimi-k3 and gemini-3.7-flash place it lower at medians of 4 and 5 respectively. The models converge on a shared characterization of breadth even as they diverge on how favorably to weigh it.

The dominant theme is Deloitte's end-to-end coverage tying AI to risk, regulatory, cyber, finance, tax, and workforce processes, making it a strong fit for regulated Fortune 500 enterprises—a framing echoed by gpt-5.6-sol, claude-opus-5, gpt-5.6-luna, and gpt-5.6-terra, several of which add the recurring caveat that delivery quality and AI depth vary by practice, member firm, and geography. The lower-ranking models reframe that same breadth as a limitation: kimi-k3 sees the AI work as more generalist and integration-led and less distinctive than MBB strategy houses or Accenture, while gemini-3.7-flash positions Deloitte as an execution and compliance partner rather than a frontier strategy or deployment leader. Sources underpinning these themes lean heavily on Deloitte's own materials—the State of Generative AI in the Enterprise reports, the AI Institute, the Trustworthy AI framework, Tech Trends, and various AI and Data service pages—supplemented by third-party analyst references including Gartner (Magic Quadrant for Data and Analytics Service Providers and the Market Guide for AI Consulting and System Integration Services), the IDC MarketScape, Forrester Wave, and Everest Group PEAK Matrix, with gemini-3.7-flash notably drawing more on Gartner and IDC MarketScape plus outlets such as Bloomberg, the Financial Times, and Fortune.

Per model summary

  • gpt-5.6-sol
    brand score
    72
    first choice
    45
    median rank
    2.0
    present in
    5/5

    Deloitte is emphasized for integrating AI with enterprise processes—risk, cyber, finance, tax, regulatory and workforce transformation—making it strong for complex regulated companies, though delivery quality and strategic distinctiveness vary by team and region.

  • claude-opus-5
    brand score
    52
    first choice
    30
    median rank
    3.0
    present in
    5/5

    Deloitte is consistently framed as a broad, end-to-end choice combining strategy, technology, risk and regulatory/AI-governance capability (Deloitte AI Institute, Trustworthy AI) well-suited to regulated Fortune 500 industries, with the recurring caveat that delivery quality varies by practice and geography.

  • gpt-5.6-luna
    brand score
    56
    first choice
    33
    median rank
    3.0
    present in
    5/5

    Deloitte is portrayed as offering broad multidisciplinary coverage (strategy, risk, cybersecurity, technology, compliance) valuable for highly regulated Fortune 500 firms, with the consistent caveat that experience and AI depth vary by member firm, geography and delivery team.

  • gpt-5.6-terra
    brand score
    68
    first choice
    40
    median rank
    3.0
    present in
    5/5

    Deloitte is consistently positioned as a strong fit where AI must be tightly integrated with risk, regulatory, cyber, finance, tax and workforce processes across complex or regulated enterprises, with a recurring recommendation to vet the specific proposed team since experience varies by practice and geography.

  • kimi-k3
    brand score
    32
    first choice
    23
    median rank
    4.0
    present in
    5/5

    Deloitte is credited with enormous breadth, its AI Institute, generative AI research and strength in regulated sectors tying AI to risk/tax/audit, but repeatedly ranked lower because its AI work is perceived as less distinctive and more generalist/integration-led than the MBB strategy houses or Accenture.

  • gemini-3.7-flash
    brand score
    28
    first choice
    22
    median rank
    5.0
    present in
    5/5

    Deloitte is described as offering comprehensive breadth in AI governance, regulatory compliance and risk via its Trustworthy AI framework and AI Institute/Academy, ideal for heavily regulated enterprises, but ranked fifth because it is seen as an execution/compliance partner rather than a frontier strategy or deployment leader.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gpt-5.6-sol8 sites · 15 references
Deloitte AI Institute×3 · 20%

www2.deloitte.com

claude-opus-57 sites · 15 references
Deloitte AI Institute / State of Generative AI in the Enterprise×7 · 47%

www2.deloitte.com

gpt-5.6-luna3 sites · 14 references
gpt-5.6-terra3 sites · 12 references
kimi-k38 sites · 15 references
Deloitte×4 · 27%

deloitte.com

gemini-3.7-flash8 sites · 15 references
Gartner×5 · 33%

gartner.com

  • ×4/
  • ×1named without a URL

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Consulting companies for AI transformation
#5

IBM Consulting

also named IBM

15
brand score
13first choice
12345

Summary

IBM Consulting occupies a consistent mid-tier position across the models, with a median rank of 5 (Q1 5, Q3 5) over 19 ranked answers from 6 models. The agreement here is unusually tight: every one of the six models—claude-opus-5, gemini-3.7-flash, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, and kimi-k3—landed on a median rank of 5, with no dispersion between the first and third quartiles at either the aggregate or individual level. The firm is referenced under two names, IBM Consulting and IBM.

The thematic consensus is equally strong. Models uniformly credit IBM with technical and engineering depth, hybrid-cloud capability, watsonx assets, data governance, and suitability for regulated and legacy-heavy environments—claude-opus-5 notes competitive pricing, while gpt-5.6-sol and gpt-5.6-terra emphasize fit where IBM or Red Hat is already strategic. The recurring reason for the lower placement is perceived vendor bias toward IBM's own stack and weaker vendor-neutrality, cited by claude-opus-5, gemini-3.7-flash, gpt-5.6-luna, gpt-5.6-sol, and gpt-5.6-terra; claude-opus-5 additionally flags weaker C-suite/board-level strategy influence, and kimi-k3 points to less breadth and large-scale transformation track record than bigger firms. These themes draw on a mix of IBM's own materials—IBM watsonx, IBM Consulting AI Services, the IBM Institute for Business Value, and IBM Annual Report and earnings commentary—alongside third-party analyst sources including Gartner, Forrester, IDC (including the IDC MarketScape), Everest Group PEAK Matrix assessments, and HFS Research.

Per model summary

  • claude-opus-5
    brand score
    24
    first choice
    21
    median rank
    5.0
    present in
    5/5

    Consistently frames IBM Consulting as strong on engineering depth, hybrid-cloud, and watsonx assets with competitive pricing, but ranks it lower due to perceived bias toward IBM's own stack and weaker C-suite/board-level strategy influence.

  • gemini-3.7-flash
    brand score
    4
    first choice
    4
    median rank
    5.0
    present in
    1/5

    Emphasizes deep hybrid-cloud and technical engineering strength for legacy modernization in regulated environments, while flagging potential vendor bias toward IBM's own platform ecosystem.

  • gpt-5.6-luna
    brand score
    12
    first choice
    12
    median rank
    5.0
    present in
    3/5

    Presents IBM as credible for hybrid cloud, data governance, and regulated enterprise environments, but ranks it lower because its fit is strongest for clients already aligned with IBM's ecosystem rather than those seeking vendor-neutral strategy.

  • gpt-5.6-sol
    brand score
    28
    first choice
    23
    median rank
    5.0
    present in
    5/5

    Positions IBM as strong for technically complex, hybrid-cloud, regulated, and legacy-heavy transformations, ranking it lower because its advantages depend on architecture aligning with IBM/Red Hat and it is seen as less vendor-neutral.

  • gpt-5.6-terra
    brand score
    12
    first choice
    12
    median rank
    5.0
    present in
    3/5

    Describes IBM as credible for hybrid-cloud, data-platform, and legacy modernization needs, especially where IBM is already strategic, but ranks it fifth for being more platform-centric and less vendor-neutral.

  • kimi-k3
    brand score
    8
    first choice
    8
    median rank
    5.0
    present in
    2/5

    Recognizes IBM's watsonx heritage and legacy-modernization strength, but ranks it fifth because its brand is tied to IBM's own stack and it lacks the breadth and large-scale transformation track record of bigger firms.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-57 sites · 15 references
IBM earnings commentary on generative AI book of business×9 · 60%

ibm.com

gemini-3.7-flash3 sites · 3 references
Forrester×1 · 33%

forrester.com

gpt-5.6-luna2 sites · 8 references
IBM Consulting AI Services×7 · 88%

ibm.com

gpt-5.6-sol7 sites · 15 references
IBM Consulting Artificial Intelligence×9 · 60%

ibm.com

gpt-5.6-terra2 sites · 7 references
IBM Consulting — Artificial intelligence×6 · 86%

ibm.com

kimi-k32 sites · 6 references
IBM Consulting×5 · 83%

ibm.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · Consulting companies for AI transformation

How this was measured

The question

  • Unaided brand recommendation question: “Which consulting firm should a Fortune 500 company hire for an AI transformation?
  • The prompt asks each model to return exactly 5 brands, ranked 1 to 5, and for each one a reason for the recommendation and the sources that informed it.

Sampling

  • Models are not deterministic, so the question is asked over and over — 5 answers per model
  • Spellings of the same brand are normalized to the most commonly used form before counting

Measures of Position: Median and Quartiles

  • Median value indicates that in 50% of answers the brand held this position or higher.
  • Quartiles help visualize the spread of rankings per brand. Q1 indicates that 25% of answers had this rank or higher. Q3 means that 75% of answers ranked the brand as X or better.
  • Medians and quartiles are computed only from answers where the brand was present.

Coverage

  • Count of answers in which the given brand was named, as a percentage.

Brand score - Normalized Borda score

  • The brand score is calculated using a normalized Borda score. Borda count is a voting method: each ballot awards points by position instead of naming one winner. Our ranking responses from the LLM are always fixed to 5 answers, so we assign 100 points for rank 1, then 80, 60, 40, 20 — and 0 if the brand is missing.
Brand score=100ni=1n6ri5\text{Brand\ score} = \frac{100}{n} \sum_{i=1}^{n} \frac{6 - r_i}{5}

where

ri={1,,5if the brand appears6if the brand is not mentionedr_i = \begin{cases} 1,\dots,5 & \text{if the brand appears} \\ 6 & \text{if the brand is not mentioned} \end{cases}
  • This way we can calculate a common score for all brands mentioned across model responses. The score reflects both how high the brand was ranked and how frequently it was mentioned.

First choice score - adjusted MRR - top-rank indicator

  • The first choice score is based on Mean Reciprocal Rank (a common search-engine measure), but adapted to measure the position of a specific brand across repeated ranked responses.
  • Each occurrence receives a reciprocal-position score: 100 for 1st, 50 for 2nd, 33.3 for 3rd, 25 for 4th, 20 for 5th, and 0 when the brand is not mentioned. The scores are averaged across all responses.
First choice score=100ni=1nsi\text{First\ choice\ score} = \frac{100}{n} \sum_{i=1}^{n} s_i

where

si={1/riif the brand is present0if the brand is absents_i = \begin{cases} 1/r_i & \text{if the brand is present} \\ 0 & \text{if the brand is absent} \end{cases}
  • No rank below 1st place gets more than 50, so the score is heavily driven by first places.
Your category

Ask your own question of every model

Tell us what to ask and how wide to sample it — we run it and send back the report.

Sample size
Display results
0/1000

Sampling every model takes up to 24 hours once we start.