Your Brand has an AI reputation and it's shaping purchase decisions right now.
Swipe up to continue
7 AI models · 160 answers · ranked by score
ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0
Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.
Swipe left for the full report
Median rank each model gave each brand.
Scroll the table sideways to see every model →
| claude-opus-5 | gemini-3.7-flash | gpt-5.6-luna | gpt-5.6-sol | gpt-5.6-terra | gpt-6-astra | kimi-k3 | |
|---|---|---|---|---|---|---|---|
| Anthropic | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 |
| OpenAI | 2.0 | 2.0 | 2.0 | 2.0 | 2.0 | 2.0 | 2.0 |
| 3.0 | 4.0 | 3.0 | 3.0 | 3.0 | 3.0 | 3.0 | |
| DeepSeek | 4.0 | 3.0 | 4.0 | 4.0 | 4.0 | 4.0 | 4.0 |
| Mistral AI | 5.0 | 5.0 | 5.0 | 5.0 | 5.0 | 5.0 | — |
| Meta | 5.0 | 5.0 | — | — | — | — | 5.0 |
| GitHub | 4.0 | — | 4.0 | — | 4.0 | — | 4.0 |
| xAI | — | — | — | 5.0 | 5.0 | — | — |
| Alibaba | 5.0 | — | — | — | — | 5.0 | 5.0 |
| Qwen | — | — | — | — | — | 5.0 | — |
| Cursor | — | — | 3.0 | — | — | — | — |
| Microsoft | — | — | 4.0 | — | — | — | — |
A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.
0 of 12 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.
No brands where the models disagree
from here down, the models agreegpt-5.6-luna 13 → gemini-3.7-flash 49
gemini-3.7-flash 0 → gpt-5.6-luna 26 · 3 of 7 never named it
claude-opus-5 0 → gpt-5.6-terra 18 · 5 of 7 never named it
gpt-5.6-luna 0 → kimi-k3 16 · 4 of 7 never named it
kimi-k3 0 → gpt-5.6-luna 18 · 1 of 7 never named it
gemini-3.7-flash 46 → kimi-k3 60
A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.
Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.
The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.
Anthropic leads with a score of 100 of 100, ranked by 7 of 7 models.
The category shows a remarkably stable spine. Anthropic, OpenAI, and Google occupy ranks 1, 2, and 3 with near-total agreement across all seven models, and the disputes only begin at rank 4 and below, where open-weight providers and product-layer tools compete for the same crowded fifth slot.
Anthropic takes first place with no dissent whatsoever: a median of 1 across 160 answers, and every individual model also returns a median of 1. OpenAI mirrors this at rank 2 (median 2 for all seven models), and Google holds rank 3 with six models agreeing. The only crack in the top three is gemini-3.7-flash, which places its own maker Google at a median of 4 rather than 3 — the sole model to rate Google below the field.
The framing splits along a benchmark-versus-workflow line, though this is a difference of emphasis rather than ranking. Benchmark-anchored models (claude-opus-5, gemini-3.7-flash, kimi-k3) lean on SWE-bench, Aider, and LMArena, while the gpt-5.6 family and gpt-6-astra frame the same placements through agentic tooling and documentation. The distinction between OpenAI and Anthropic is repeatedly cast as task-dependent rather than absolute.
“They are particularly strong at algorithmic problem-solving and explaining code, though some developers find them slightly less consistent than Claude on large, multi-step engineering tasks.”
The real disagreement concerns DeepSeek. Six models place it at rank 4, but gemini-3.7-flash rates it a full rank higher at 3 — the same model that demoted Google. Its reasoning leans on benchmark parity and self-hosting appeal (SWE-bench, Hugging Face, LMSYS), whereas the six models holding DeepSeek at 4 emphasize ecosystem and enterprise-support gaps drawn from its API docs and GitHub. Whether gemini-3.7-flash's benchmark focus causes both its DeepSeek promotion and its Google demotion is a plausible reading of its source mix, not something the data confirms.
Below that, three fundamentally different kinds of brand pile up at rank 5:
The GitHub case is notable because its rank-4 placement sits above the open-weight labs at 5, yet every model attributes its ceiling to inherited capability rather than its own model.
“GitHub Copilot is the most widely deployed coding assistant, with deep IDE and pull-request integration, enterprise controls and now a model picker that routes to Anthropic, OpenAI and Google models.”
Cursor, Microsoft, and Qwen each rest on a single model's answers, so their ranks reflect one perspective rather than any consensus. Cursor (gpt-5.6-luna, rank 3) and Qwen (gpt-6-astra, rank 5) are internally consistent within their lone raters, but carry no cross-model signal. Both are worth reading as individual judgments, not category positions.
A recurring qualifier across the low-ranked open-weight entries is that the placement would rise sharply if self-hosting were the priority — gpt-6-astra says as much for both Alibaba and Qwen, which suggests the rank-5 clustering reflects an implied "typical user" default more than a capability verdict.
“I rank it fifth for a typical user seeking an immediately useful coding assistant, but it could rank considerably higher if self-hosting and control over deployment are your priorities.”
The source pool is dominated by a handful of leaderboards and benchmarks — Aider (152 times), SWE-bench (147), LMArena (79), and Artificial Analysis (48) — which explains why placements agree so tightly across models drawing on a common evidence base. Vendor documentation appears mainly for the lower-ranked and open-weight brands, consistent with those rankings being argued on deployment and ecosystem grounds rather than measured performance.
65 sources across 2145 references, grouped by site from 496 recalled names. The top 5 carry 45% of them.
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
Anthropic sits at the front of this category with unusual consistency: a median rank of 1 (Q1 1, Q3 1) across 160 ranked answers from all 7 models, and every one of the seven models individually returns a median rank of 1 as well. There is effectively no disagreement about placement — the divergence is only in emphasis. The benchmark-anchored models (claude-opus-5, gemini-3.7-flash, kimi-k3) lean on measured performance, citing SWE-bench Verified leadership and integration into tools like Cursor and GitHub Copilot, drawing on the SWE-bench leaderboard, Aider LLM leaderboards, and LMArena. The workflow-oriented models (gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra) frame the same ranking through repository-scale reasoning, multi-file edits, and agentic tooling such as Claude Code, sourcing Anthropic documentation and the Claude Code overview.
“Claude models (Sonnet/Opus 4.x series) are widely regarded as the strongest at real-world software engineering tasks, leading benchmarks like SWE-bench Verified and powering tools such as Claude Code, Cursor and GitHub Copilot.”
Underlying the top placement is a shared theme of large-codebase competence — long context, coordinated multi-file changes, careful instruction-following, and clear explanations rather than isolated snippet generation. Both camps land on the same conclusion from different angles, and one model even flags its own conflict of interest while retaining the ranking.
“Claude models are especially strong at understanding large codebases, making coordinated multi-file changes, and explaining code clearly.”
Consistently emphasizes that Claude (Sonnet/Opus 4.x) leads real-world coding benchmarks like SWE-bench Verified and excels at agentic, multi-file refactoring in large codebases. Repeatedly cites Claude Code and integration into tools like Cursor and GitHub Copilot as reinforcing its top position.
Repeatedly frames Anthropic's Claude models (especially Claude 3.5 Sonnet) as the industry gold standard for code generation, refactoring, and architectural reasoning. Consistently highlights large context windows, strong instruction-following, and minimal hallucinations.
Consistently recommends Anthropic for its strength in understanding large codebases, debugging, refactoring, and following detailed instructions to produce maintainable code. Emphasizes careful reasoning, clear explanations, and agentic/repository-level workflows.
Consistently positions Anthropic as the strongest overall coding recommendation, citing repository-scale reasoning, multi-file edits, and clear explanations. Emphasizes long context, agentic tooling like Claude Code, and reliable instruction-following.
Repeatedly recommends Anthropic for demanding, professional coding work, emphasizing repository-scale reasoning, code review, debugging, and following detailed engineering constraints. Highlights long-context capability and agentic workflows over lowest cost.
Consistently names Anthropic as its default choice for complex, general-purpose coding, focused on understanding existing codebases, coordinated multi-file changes, and explaining design decisions. Emphasizes Claude Code and agentic, repository-oriented workflows over snippet generation.
Consistently cites Claude models topping coding benchmarks like SWE-bench Verified and strong developer sentiment, with emphasis on multi-file reasoning and agentic tools like Claude Code. Notes precise instruction-following and integration into tools such as Cursor.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
anthropic.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
OpenAI lands at a consistent second across the field, with a median rank of 2 (Q1 2, Q3 2) over 159 ranked answers, and every one of the 7 models reports the same median of 2. The agreement is unusually tight: each model frames OpenAI as a close runner-up to Anthropic rather than a distant one. The recurring theme is a split between recognized strengths — algorithmic problem solving, debugging, reasoning-heavy tasks, and the broadest tooling ecosystem — and the reason it stops just short of first, namely that Claude is perceived as more reliable on large-repo, multi-file agentic work. claude-opus-5, kimi-k3, and gemini-3.7-flash all lean on ecosystem breadth, citing sources like the Aider LLM leaderboards, LMArena, GitHub Copilot, and SWE-bench, while the gpt-5.6 variants and gpt-6-astra emphasize Codex tooling and API maturity, backed by SWE-bench and OpenAI Codex documentation.
“It is a very close second and often better for algorithmic reasoning and competitive-programming style problems.”
The second theme is variability: several models qualify their ranking by noting that the best OpenAI choice depends on the specific model, task, or workflow, which is the main reason it settles narrowly behind Anthropic rather than tying or leading. This shows up plainly in the gpt-5.6 reasonings, where the ranking is repeatedly described as task-dependent, and in gpt-6-astra's advice to test both providers on one's own repository rather than assume a universal winner.
“They are particularly strong at algorithmic problem-solving and explaining code, though some developers find them slightly less consistent than Claude on large, multi-step engineering tasks.”
GPT-5/o-series and Codex models are praised for algorithmic problem solving, debugging, and competitive-programming tasks with the broadest tooling ecosystem, but consistently placed a very close second to Anthropic due to Claude's edge on large-repo agentic coding.
Emphasizes GPT-4o and o1 reasoning models as industry-leading for algorithmic problem solving and debugging, with unmatched ecosystem integration into tools like GitHub Copilot and Cursor, while noting it is slightly edged out on nuanced multi-file/full-stack tasks.
Describes OpenAI as a highly capable, versatile general-purpose option for code generation, debugging, and tool use, ranked slightly below Anthropic because coding quality varies by model and workflow.
Highlights strong code generation, debugging, and agentic/Codex tooling backed by a mature API ecosystem, ranking it just behind Anthropic since quality and model selection can vary by task.
Frames OpenAI as an excellent, safe general-purpose default with strong reasoning, mature APIs, and broad tooling, ranked just below Anthropic because the best choice depends on the specific environment and workflow.
Positions OpenAI as a close second, strong for code generation, debugging, tests, and Codex-supported workflows with a major ecosystem advantage, while slightly favoring Anthropic for repository-heavy editing and recommending testing both on one's own repo.
Notes GPT/o-series models excel at algorithmic and reasoning-heavy tasks with the most mature ecosystem (GitHub Copilot, API), but rank just behind Anthropic as developers find Claude more reliable on complex, multi-file software engineering.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
openai.com
platform.openai.com
openai.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named Google DeepMind
Google occupies a stable third position in the models' coding recommendations, with a median rank of 3 (Q1 3, Q3 3) across 159 ranked answers from 7 models. Six of the seven models—claude-opus-5, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra, and kimi-k3—all converge on a median of 3, while gemini-3.7-flash sits slightly lower at a median rank of 4 (Q1 3, Q3 4). The models refer to the brand as both Google and Google DeepMind. The recurring theme behind this placement is a distinctive strength in large context windows paired with a perceived gap in day-to-day coding consistency relative to the top two brands.
The dominant point of agreement is that Gemini's very large context window—cited by several models as reaching one to two million tokens—makes it well suited to reasoning across entire repositories and documentation-heavy projects, reinforced by multimodal capabilities and Google Cloud, Android, and developer-ecosystem integration. This is backed by sources including the Google DeepMind Gemini model page (cited 19 times by claude-opus-5), LMArena (24 times by kimi-k3), SWE-bench, the Aider LLM leaderboards, and Gemini API and Code Assist documentation. The consistent caveat, and the reason Google lands third rather than higher, is that its agentic and iterative coding output is seen as less consistent than Anthropic's or OpenAI's, with gemini-3.7-flash also flagging occasional shortfalls in code precision on complex edge cases.
“Gemini 2.5/3 Pro offers a huge context window that is great for reasoning across entire repositories, plus strong performance and generous free tiers via AI Studio and Gemini CLI.”
“They rank third because, while highly capable and improving rapidly, developer consensus generally places their coding output quality just behind Anthropic and OpenAI on day-to-day tasks.”
Consistently highlights Gemini 2.5/3 Pro's very large context window for repository-wide reasoning, strong benchmarks, multimodal/frontend strength, and generous free-tier access via AI Studio and Gemini CLI, while ranking it a close third due to less consistent agentic code editing than Claude or GPT.
Frames Google as a strong option for long-context, multimodal work and for developers in the Google Cloud/Android ecosystem, but consistently ranks it behind Anthropic and OpenAI due to less consistent coding quality across model versions and integrations.
Stresses Gemini's very large context windows for analyzing large repositories and documentation plus Google Cloud integration, while noting coding consistency varies across model tiers and trails the top two choices.
Highlights large context windows, multimodal inputs, and Google Cloud/ecosystem integration as strengths, positioning coding performance as competitive but variable by model version and workflow, generally below the top two for demanding agent tasks.
Consistently recommends Gemini as a strong context-heavy and multimodal option ranked third for general coding, but potentially a first choice for workflows involving large documentation, many files, or Google's developer ecosystem.
Emphasizes Gemini's exceptionally large context window for reasoning across entire repositories and competitive benchmark scores and pricing, while ranking it third because developer mindshare and day-to-day coding consistency slightly trail Anthropic and OpenAI.
Repeatedly emphasizes massive multi-million-token context windows for ingesting entire codebases and documentation, plus multimodal and ecosystem integration, while noting occasional shortfalls in code precision or idiomatic output on complex edge cases relative to top competitors.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
cloud.google.com
ai.google.dev
ai.google.dev
ai.google.dev
deepmind.google
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
DeepSeek settles at a median rank of 4 across all seven models (Q1 4, Q3 4) over 141 ranked answers, marking it as a consistent mid-table recommendation rather than a top pick. Agreement is notably tight: six of the seven models place it at a median of 4, with only gemini-3.7-flash landing higher at a median of 3 (Q1 3, Q3 3). The shared narrative pairs a strong upside — near-frontier coding quality at a fraction of the cost, with open weights for self-hosting — against a recurring limitation around less mature tooling, enterprise support, and reliability on complex agentic tasks. This value-versus-maturity framing runs through claude-opus-5, kimi-k3, and the GPT-family models alike, drawing heavily on sources such as the Aider LLM leaderboards, DeepSeek's API documentation and GitHub repositories, Hugging Face model cards, Artificial Analysis, and SWE-bench.
The cost-and-openness theme is where models converge most tightly, typically citing the V3 and R1 lineage as the anchor for their reasoning.
“DeepSeek's V3 and R1 models deliver coding performance that rivals much more expensive proprietary models, making them exceptional value and a leading open-weight option.”
The offsetting theme — why it rarely climbs above rank 4 — centers on deployment effort, governance, and ecosystem maturity, drawn from the API documentation and GitHub-heavy source pool that the GPT models lean on.
“I rank it fourth because its surrounding tools, enterprise support, and overall developer ecosystem are less mature than those of the leading providers.”
Gemini's higher placement reflects a heavier emphasis on benchmark parity and self-hosting appeal, supported by references to the SWE-bench Leaderboard, Hugging Face, and the LMSYS Chatbot Arena.
Repeatedly presents DeepSeek as an open-weights powerhouse rivaling proprietary models on coding benchmarks at a fraction of the cost, ideal for self-hosting, with a less mature developer tooling ecosystem as its main limitation.
Consistently frames DeepSeek's V3/R1-lineage models as delivering near-frontier coding performance at a fraction of the cost with open weights for self-hosting, while trailing the top proprietary labs on complex agentic tasks and tooling polish.
Emphasizes DeepSeek's capable coding performance and cost efficiency with open-weight availability, ranking it below top providers due to less predictable reliability, tooling consistency, and deployment/ecosystem maturity.
Consistently describes DeepSeek as offering capable coding and reasoning at competitive prices with open-weight/self-hosting options, ranking it lower because support, governance, privacy, and ecosystem maturity require more evaluation.
Frames DeepSeek as a value-oriented option strong on cost efficiency, open weights, and self-hosting flexibility, placing it below leading providers because deployment, support, safety controls, and production tooling need more evaluation.
Positions DeepSeek as attractive for cost efficiency and access to open model weights, ranking it below the leading integrated assistants because deployment choices, self-hosting, and workflow tooling require more setup effort.
Consistently highlights DeepSeek's V3 and R1 models delivering near-frontier coding at very low cost with open weights for self-hosting, ranking it below the top labs due to less mature tooling, reliability, and consistency on complex tasks.
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
huggingface.co
github.com
github.com
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
also named Mistral
Mistral AI settles at a median rank of 5 (Q1 5, Q3 5) across 57 ranked answers from all 6 models, and the tight quartiles signal strong agreement: every model places it fifth, with only gemini-3.7-flash showing marginally more spread (Q1 4, Q3 5). The consensus is less about weakness than positioning. Models frame Mistral as a specialist and deployment-oriented pick rather than a top capability choice, repeatedly citing its coding-focused models Codestral and Devstral. claude-opus-5 and gemini-3.7-flash lean on the technical strengths — fast, cheap, open-weight models tuned for autocomplete and fill-in-the-middle — while noting they trail frontier models on hard multi-file work. These themes rest on sources such as the Hugging Face Devstral model card, Mistral AI docs, the Codestral announcement, and Artificial Analysis.
“Codestral and Devstral are purpose-built coding models with permissive/open weights and excellent latency for fill-in-the-middle autocomplete, plus EU data-residency appeal.”
The second recurring theme, most prominent in the gpt-5.6 family, is that Mistral appeals when openness, deployment flexibility, and European hosting matter more than maximizing raw coding performance. gpt-5.6-luna, gpt-5.6-sol, and gpt-5.6-terra consistently rank it fifth on the grounds of a less established ecosystem and weaker demonstrated results on demanding repository-scale tasks, drawing on Mistral AI documentation alongside benchmark references like Aider LLM Leaderboards and SWE-bench. gpt-6-astra echoes the completion-workflow angle, treating it as a targeted rather than general-purpose fit.
“Mistral is a good choice when openness, deployment flexibility, and control over infrastructure matter more than achieving the strongest possible coding results.”
Consistently highlights Codestral and Devstral as fast, cheap, open-weight models good for autocomplete, fill-in-the-middle, and EU/on-prem deployment, while noting they lag frontier models on hard multi-file tasks, making Mistral a privacy/cost choice rather than a top capability pick.
Emphasizes Codestral as an efficient, lightweight model optimized for low-latency code completion, fill-in-the-middle, and broad language support, ideal for IDE autocompletion and self-hosting, but ranked fifth because it trails frontier models on complex multi-step reasoning.
Positions Mistral as appealing for open-weight options, deployment flexibility, and European hosting, but ranks it fifth because its coding ecosystem, tooling, and performance on complex multi-file tasks are less consistently strong than leading providers.
Frames Mistral as a good option for deployment flexibility, European hosting, and open-weight/Codestral coding models, but ranks it below others due to a less established ecosystem and weaker demonstrated performance on demanding repository-scale tasks.
Recommends Mistral when European hosting, deployment flexibility, or open-weight options matter, while generally preferring higher-ranked providers first for the hardest end-to-end software-engineering tasks.
Values Mistral for specialized code completion and fill-in-the-middle workflows with deployment flexibility, ranking it fifth for broad complex coding assistance while noting it can fit targeted completion-oriented workloads (and advises checking licensing).
Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.
mistral.ai
docs.mistral.ai
Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.
where
where
Each one is a single question put to every active AI model, many times over.
See every category with its question and leaders.
Tell us what to ask and how wide to sample it — we run it and send back the report.