Your Brand has an AI reputation and it's shaping purchase decisions right now.

Swipe up to continue

Category · AI models for coding

Which AI model would you recommend for coding?

7 AI models · 160 answers · ranked by score

ordered by brand score: 100 pts for rank 1 → 20 for rank 5 · no mention = 0

  1. #1Anthropic
    100
    brand score
    99first choice
    12345
  2. #2OpenAI
    80
    brand score
    50first choice
  3. #3Google
    57
    brand score
    32first choice
  4. #4DeepSeek
    35
    brand score
    22first choice
  5. #5Mistral AI
    8
    brand score
    8first choice
Show 7 more
  1. #6Meta
    5
    brand score
    5first choice
  2. #7GitHub
    5
    brand score
    3first choice
  3. #8xAI
    4
    brand score
    4first choice
  4. #9Alibaba
    2
    brand score
    2first choice
  5. #10Qwen
    1
    brand score
    1first choice
  6. #11Cursor
    0
    brand score
    0first choice
  7. #12Microsoft
    0
    brand score
    0first choice

Each bar covers the middle half of one model’s answers (Q1–Q3); the line inside it is that model’s median rank for the times it recommended the brand. The badge under a brand says how far apart the models are on it, and “first choice” scores rank 1 far above the rest — a brand can place well overall and still rarely lead an answer.

Swipe left for the full report

Category · AI models for coding

Model by model

Median rank each model gave each brand.

Scroll the table sideways to see every model →

claude-opus-5gemini-3.7-flashgpt-5.6-lunagpt-5.6-solgpt-5.6-terragpt-6-astrakimi-k3
Anthropic1.01.01.01.01.01.01.0
OpenAI2.02.02.02.02.02.02.0
Google3.04.03.03.03.03.03.0
DeepSeek4.03.04.04.04.04.04.0
Mistral AI5.05.05.05.05.05.0
Meta5.05.05.0
GitHub4.04.04.04.0
xAI5.05.0
Alibaba5.05.05.0
Qwen5.0
Cursor3.0
Microsoft4.0
1st
2
3
4
5th

A dash means that model never named the brand in any of its runs. Deeper blue is better: rank 1 is the brand a model would recommend first.

Category · AI models for coding

Where the models disagree

0 of 12 brands split the panel. Each mark is one model's score for the brand, on the same 0–100 scale as the ranking.

  1. #4DeepSeek

    gpt-5.6-luna 13 gemini-3.7-flash 49

  2. #7GitHub

    gemini-3.7-flash 0 gpt-5.6-luna 26 · 3 of 7 never named it

  3. #8xAI

    claude-opus-5 0 gpt-5.6-terra 18 · 5 of 7 never named it

  4. #6Meta

    gpt-5.6-luna 0 kimi-k3 16 · 4 of 7 never named it

  5. #5Mistral AI

    kimi-k3 0 gpt-5.6-luna 18 · 1 of 7 never named it

  6. #3Google

    gemini-3.7-flash 46 kimi-k3 60

A hollow mark is a model that never named the brand in any of its runs, which scores 0. Agreement is not endorsement — a brand every model ignores equally agrees just as tightly as one they all rank first.

Category · AI models for coding

Named, and named first

Across → how frequently the panel names the brand at all.
Up ↑ how often the answers that name it put it first.

the default answera narrow favouritelisted, rarely ledthe long tail12345678
  1. 1Anthropic100/99
  2. 2OpenAI99/1
  3. 3Google99/0
  4. 4DeepSeek87/0
  5. 5Mistral AI37/0
  6. 6Meta27/0
  7. 7GitHub13/0
  8. 8xAI18/0

The horizontal line sits at 20% — the rate a named brand would lead at if the models were picking one of its 5 slots at random. Above it they are choosing it first on purpose. Both figures average across models, so a thinly sampled model counts the same as a heavily sampled one.

Category · AI models for coding

What the models actually said

Models
7
Answers
160

Anthropic leads with a score of 100 of 100, ranked by 7 of 7 models.

The category shows a remarkably stable spine. Anthropic, OpenAI, and Google occupy ranks 1, 2, and 3 with near-total agreement across all seven models, and the disputes only begin at rank 4 and below, where open-weight providers and product-layer tools compete for the same crowded fifth slot.

The undisputed top three

Anthropic takes first place with no dissent whatsoever: a median of 1 across 160 answers, and every individual model also returns a median of 1. OpenAI mirrors this at rank 2 (median 2 for all seven models), and Google holds rank 3 with six models agreeing. The only crack in the top three is gemini-3.7-flash, which places its own maker Google at a median of 4 rather than 3 — the sole model to rate Google below the field.

The framing splits along a benchmark-versus-workflow line, though this is a difference of emphasis rather than ranking. Benchmark-anchored models (claude-opus-5, gemini-3.7-flash, kimi-k3) lean on SWE-bench, Aider, and LMArena, while the gpt-5.6 family and gpt-6-astra frame the same placements through agentic tooling and documentation. The distinction between OpenAI and Anthropic is repeatedly cast as task-dependent rather than absolute.

They are particularly strong at algorithmic problem-solving and explaining code, though some developers find them slightly less consistent than Claude on large, multi-step engineering tasks.

kimi-k3

Where the models actually diverge

The real disagreement concerns DeepSeek. Six models place it at rank 4, but gemini-3.7-flash rates it a full rank higher at 3 — the same model that demoted Google. Its reasoning leans on benchmark parity and self-hosting appeal (SWE-bench, Hugging Face, LMSYS), whereas the six models holding DeepSeek at 4 emphasize ecosystem and enterprise-support gaps drawn from its API docs and GitHub. Whether gemini-3.7-flash's benchmark focus causes both its DeepSeek promotion and its Google demotion is a plausible reading of its source mix, not something the data confirms.

Below that, three fundamentally different kinds of brand pile up at rank 5:

  • Open-weight providers — Mistral AI, Meta, and Alibaba/Qwen — all sit at a median of 5, and the reasoning is strikingly uniform: capable and open, but ranked on out-of-the-box coding quality rather than flexibility. gemini-3.7-flash again shows marginally more spread on Mistral (Q1 4).
  • A frontier-adjacent lab — xAI, placed at 5 by only gpt-5.6-sol and gpt-5.6-terra, judged capable but thin on tooling and track record.
  • Product layers — GitHub (rank 4) and Microsoft (rank 4, one answer), rated as routers or integration layers over other labs' models rather than model providers in their own right.

The GitHub case is notable because its rank-4 placement sits above the open-weight labs at 5, yet every model attributes its ceiling to inherited capability rather than its own model.

GitHub Copilot is the most widely deployed coding assistant, with deep IDE and pull-request integration, enterprise controls and now a model picker that routes to Anthropic, OpenAI and Google models.

claude-opus-5

Single-model entries

Cursor, Microsoft, and Qwen each rest on a single model's answers, so their ranks reflect one perspective rather than any consensus. Cursor (gpt-5.6-luna, rank 3) and Qwen (gpt-6-astra, rank 5) are internally consistent within their lone raters, but carry no cross-model signal. Both are worth reading as individual judgments, not category positions.

A recurring qualifier across the low-ranked open-weight entries is that the placement would rise sharply if self-hosting were the priority — gpt-6-astra says as much for both Alibaba and Qwen, which suggests the rank-5 clustering reflects an implied "typical user" default more than a capability verdict.

I rank it fifth for a typical user seeking an immediately useful coding assistant, but it could rank considerably higher if self-hosting and control over deployment are your priorities.

gpt-6-astra

Shared evidence base

The source pool is dominated by a handful of leaderboards and benchmarks — Aider (152 times), SWE-bench (147), LMArena (79), and Artificial Analysis (48) — which explains why placements agree so tightly across models drawing on a common evidence base. Vendor documentation appears mainly for the lower-ranked and open-weight brands, consistent with those rankings being argued on deployment and ecosystem grounds rather than measured performance.

Category · AI models for coding

What shaped the answers

65 sources across 2145 references, grouped by site from 496 recalled names. The top 5 carry 45% of them.

SWE-bench×311 · 14%

swebench.com

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · AI models for coding
#1

Anthropic

100
brand score
99first choice
12345

Summary

Anthropic sits at the front of this category with unusual consistency: a median rank of 1 (Q1 1, Q3 1) across 160 ranked answers from all 7 models, and every one of the seven models individually returns a median rank of 1 as well. There is effectively no disagreement about placement — the divergence is only in emphasis. The benchmark-anchored models (claude-opus-5, gemini-3.7-flash, kimi-k3) lean on measured performance, citing SWE-bench Verified leadership and integration into tools like Cursor and GitHub Copilot, drawing on the SWE-bench leaderboard, Aider LLM leaderboards, and LMArena. The workflow-oriented models (gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra) frame the same ranking through repository-scale reasoning, multi-file edits, and agentic tooling such as Claude Code, sourcing Anthropic documentation and the Claude Code overview.

Claude models (Sonnet/Opus 4.x series) are widely regarded as the strongest at real-world software engineering tasks, leading benchmarks like SWE-bench Verified and powering tools such as Claude Code, Cursor and GitHub Copilot.

claude-opus-5

Underlying the top placement is a shared theme of large-codebase competence — long context, coordinated multi-file changes, careful instruction-following, and clear explanations rather than isolated snippet generation. Both camps land on the same conclusion from different angles, and one model even flags its own conflict of interest while retaining the ranking.

Claude models are especially strong at understanding large codebases, making coordinated multi-file changes, and explaining code clearly.

gpt-5.6-sol

Per model summary

  • claude-opus-5
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    25/25

    Consistently emphasizes that Claude (Sonnet/Opus 4.x) leads real-world coding benchmarks like SWE-bench Verified and excels at agentic, multi-file refactoring in large codebases. Repeatedly cites Claude Code and integration into tools like Cursor and GitHub Copilot as reinforcing its top position.

  • gemini-3.7-flash
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    25/25

    Repeatedly frames Anthropic's Claude models (especially Claude 3.5 Sonnet) as the industry gold standard for code generation, refactoring, and architectural reasoning. Consistently highlights large context windows, strong instruction-following, and minimal hallucinations.

  • gpt-5.6-luna
    brand score
    98
    first choice
    95
    median rank
    1.0
    present in
    20/20

    Consistently recommends Anthropic for its strength in understanding large codebases, debugging, refactoring, and following detailed instructions to produce maintainable code. Emphasizes careful reasoning, clear explanations, and agentic/repository-level workflows.

  • gpt-5.6-sol
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    20/20

    Consistently positions Anthropic as the strongest overall coding recommendation, citing repository-scale reasoning, multi-file edits, and clear explanations. Emphasizes long context, agentic tooling like Claude Code, and reliable instruction-following.

  • gpt-5.6-terra
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    20/20

    Repeatedly recommends Anthropic for demanding, professional coding work, emphasizing repository-scale reasoning, code review, debugging, and following detailed engineering constraints. Highlights long-context capability and agentic workflows over lowest cost.

  • gpt-6-astra
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    25/25

    Consistently names Anthropic as its default choice for complex, general-purpose coding, focused on understanding existing codebases, coordinated multi-file changes, and explaining design decisions. Emphasizes Claude Code and agentic, repository-oriented workflows over snippet generation.

  • kimi-k3
    brand score
    100
    first choice
    100
    median rank
    1.0
    present in
    25/25

    Consistently cites Claude models topping coding benchmarks like SWE-bench Verified and strong developer sentiment, with emphasis on multi-file reasoning and agentic tools like Claude Code. Notes precise instruction-following and integration into tools such as Cursor.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-58 sites · 75 references
SWE-bench leaderboard×25 · 33%

swebench.com

  • ×25/
gemini-3.7-flash8 sites · 69 references
SWE-bench×25 · 36%

swebench.com

  • ×25/
gpt-5.6-luna6 sites · 51 references
SWE-bench×18 · 35%

swebench.com

  • ×18/
gpt-5.6-sol7 sites · 60 references
SWE-bench×20 · 33%

swebench.com

  • ×20/
gpt-5.6-terra7 sites · 58 references
SWE-bench×20 · 34%

swebench.com

gpt-6-astra6 sites · 53 references
Anthropic — Claude Code best practices×22 · 42%

anthropic.com

kimi-k36 sites · 75 references
SWE-bench×25 · 33%

swebench.com

  • ×25/

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · AI models for coding
#2

OpenAI

80
brand score
50first choice
12345

Summary

OpenAI lands at a consistent second across the field, with a median rank of 2 (Q1 2, Q3 2) over 159 ranked answers, and every one of the 7 models reports the same median of 2. The agreement is unusually tight: each model frames OpenAI as a close runner-up to Anthropic rather than a distant one. The recurring theme is a split between recognized strengths — algorithmic problem solving, debugging, reasoning-heavy tasks, and the broadest tooling ecosystem — and the reason it stops just short of first, namely that Claude is perceived as more reliable on large-repo, multi-file agentic work. claude-opus-5, kimi-k3, and gemini-3.7-flash all lean on ecosystem breadth, citing sources like the Aider LLM leaderboards, LMArena, GitHub Copilot, and SWE-bench, while the gpt-5.6 variants and gpt-6-astra emphasize Codex tooling and API maturity, backed by SWE-bench and OpenAI Codex documentation.

It is a very close second and often better for algorithmic reasoning and competitive-programming style problems.

claude-opus-5

The second theme is variability: several models qualify their ranking by noting that the best OpenAI choice depends on the specific model, task, or workflow, which is the main reason it settles narrowly behind Anthropic rather than tying or leading. This shows up plainly in the gpt-5.6 reasonings, where the ranking is repeatedly described as task-dependent, and in gpt-6-astra's advice to test both providers on one's own repository rather than assume a universal winner.

They are particularly strong at algorithmic problem-solving and explaining code, though some developers find them slightly less consistent than Claude on large, multi-step engineering tasks.

kimi-k3

Per model summary

  • claude-opus-5
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    25/25

    GPT-5/o-series and Codex models are praised for algorithmic problem solving, debugging, and competitive-programming tasks with the broadest tooling ecosystem, but consistently placed a very close second to Anthropic due to Claude's edge on large-repo agentic coding.

  • gemini-3.7-flash
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    25/25

    Emphasizes GPT-4o and o1 reasoning models as industry-leading for algorithmic problem solving and debugging, with unmatched ecosystem integration into tools like GitHub Copilot and Cursor, while noting it is slightly edged out on nuanced multi-file/full-stack tasks.

  • gpt-5.6-luna
    brand score
    82
    first choice
    55
    median rank
    2.0
    present in
    20/20

    Describes OpenAI as a highly capable, versatile general-purpose option for code generation, debugging, and tool use, ranked slightly below Anthropic because coding quality varies by model and workflow.

  • gpt-5.6-sol
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    20/20

    Highlights strong code generation, debugging, and agentic/Codex tooling backed by a mature API ecosystem, ranking it just behind Anthropic since quality and model selection can vary by task.

  • gpt-5.6-terra
    brand score
    76
    first choice
    48
    median rank
    2.0
    present in
    19/20

    Frames OpenAI as an excellent, safe general-purpose default with strong reasoning, mature APIs, and broad tooling, ranked just below Anthropic because the best choice depends on the specific environment and workflow.

  • gpt-6-astra
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    25/25

    Positions OpenAI as a close second, strong for code generation, debugging, tests, and Codex-supported workflows with a major ecosystem advantage, while slightly favoring Anthropic for repository-heavy editing and recommending testing both on one's own repo.

  • kimi-k3
    brand score
    80
    first choice
    50
    median rank
    2.0
    present in
    25/25

    Notes GPT/o-series models excel at algorithmic and reasoning-heavy tasks with the most mature ecosystem (GitHub Copilot, API), but rank just behind Anthropic as developers find Claude more reliable on complex, multi-file software engineering.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-56 sites · 75 references
LMArena coding leaderboard×25 · 33%

lmarena.ai

gemini-3.7-flash8 sites · 69 references
OpenAI Research×22 · 32%

openai.com

gpt-5.6-luna6 sites · 52 references
OpenAI×17 · 33%

openai.com

gpt-5.6-sol5 sites · 60 references
SWE-bench×20 · 33%

swebench.com

  • ×20/
gpt-5.6-terra7 sites · 55 references
OpenAI API documentation×19 · 35%

platform.openai.com

gpt-6-astra6 sites · 56 references
OpenAI — Introducing Codex×16 · 29%

openai.com

kimi-k37 sites · 75 references
LMArena×20 · 27%

lmarena.ai

  • ×20/

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · AI models for coding
#3

Google

also named Google DeepMind

57
brand score
32first choice
12345

Summary

Google occupies a stable third position in the models' coding recommendations, with a median rank of 3 (Q1 3, Q3 3) across 159 ranked answers from 7 models. Six of the seven models—claude-opus-5, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra, and kimi-k3—all converge on a median of 3, while gemini-3.7-flash sits slightly lower at a median rank of 4 (Q1 3, Q3 4). The models refer to the brand as both Google and Google DeepMind. The recurring theme behind this placement is a distinctive strength in large context windows paired with a perceived gap in day-to-day coding consistency relative to the top two brands.

The dominant point of agreement is that Gemini's very large context window—cited by several models as reaching one to two million tokens—makes it well suited to reasoning across entire repositories and documentation-heavy projects, reinforced by multimodal capabilities and Google Cloud, Android, and developer-ecosystem integration. This is backed by sources including the Google DeepMind Gemini model page (cited 19 times by claude-opus-5), LMArena (24 times by kimi-k3), SWE-bench, the Aider LLM leaderboards, and Gemini API and Code Assist documentation. The consistent caveat, and the reason Google lands third rather than higher, is that its agentic and iterative coding output is seen as less consistent than Anthropic's or OpenAI's, with gemini-3.7-flash also flagging occasional shortfalls in code precision on complex edge cases.

Gemini 2.5/3 Pro offers a huge context window that is great for reasoning across entire repositories, plus strong performance and generous free tiers via AI Studio and Gemini CLI.

claude-opus-5

They rank third because, while highly capable and improving rapidly, developer consensus generally places their coding output quality just behind Anthropic and OpenAI on day-to-day tasks.

kimi-k3

Per model summary

  • claude-opus-5
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    25/25

    Consistently highlights Gemini 2.5/3 Pro's very large context window for repository-wide reasoning, strong benchmarks, multimodal/frontend strength, and generous free-tier access via AI Studio and Gemini CLI, while ranking it a close third due to less consistent agentic code editing than Claude or GPT.

  • gpt-5.6-luna
    brand score
    58
    first choice
    33
    median rank
    3.0
    present in
    20/20

    Frames Google as a strong option for long-context, multimodal work and for developers in the Google Cloud/Android ecosystem, but consistently ranks it behind Anthropic and OpenAI due to less consistent coding quality across model versions and integrations.

  • gpt-5.6-sol
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    20/20

    Stresses Gemini's very large context windows for analyzing large repositories and documentation plus Google Cloud integration, while noting coding consistency varies across model tiers and trails the top two choices.

  • gpt-5.6-terra
    brand score
    57
    first choice
    32
    median rank
    3.0
    present in
    19/20

    Highlights large context windows, multimodal inputs, and Google Cloud/ecosystem integration as strengths, positioning coding performance as competitive but variable by model version and workflow, generally below the top two for demanding agent tasks.

  • gpt-6-astra
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    25/25

    Consistently recommends Gemini as a strong context-heavy and multimodal option ranked third for general coding, but potentially a first choice for workflows involving large documentation, many files, or Google's developer ecosystem.

  • kimi-k3
    brand score
    60
    first choice
    33
    median rank
    3.0
    present in
    25/25

    Emphasizes Gemini's exceptionally large context window for reasoning across entire repositories and competitive benchmark scores and pricing, while ranking it third because developer mindshare and day-to-day coding consistency slightly trail Anthropic and OpenAI.

  • gemini-3.7-flash
    brand score
    46
    first choice
    27
    median rank
    4.0
    present in
    25/25

    Repeatedly emphasizes massive multi-million-token context windows for ingesting entire codebases and documentation, plus multimodal and ecosystem integration, while noting occasional shortfalls in code precision or idiomatic output on complex edge cases relative to top competitors.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-510 sites · 75 references
Google DeepMind Gemini model page×25 · 33%

deepmind.google

gpt-5.6-luna7 sites · 51 references
Gemini Code Assist×15 · 29%

cloud.google.com

gpt-5.6-sol11 sites · 60 references
Google Gemini API documentation×19 · 32%

ai.google.dev

gpt-5.6-terra8 sites · 55 references
Google AI for Developers×21 · 38%

ai.google.dev

gpt-6-astra5 sites · 51 references
Google — Gemini API documentation×32 · 63%

ai.google.dev

kimi-k36 sites · 67 references
LMArena×25 · 37%

lmarena.ai

  • ×25/
gemini-3.7-flash9 sites · 61 references
Google DeepMind×24 · 39%

deepmind.google

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · AI models for coding
#4

DeepSeek

35
brand score
22first choice
12345

Summary

DeepSeek settles at a median rank of 4 across all seven models (Q1 4, Q3 4) over 141 ranked answers, marking it as a consistent mid-table recommendation rather than a top pick. Agreement is notably tight: six of the seven models place it at a median of 4, with only gemini-3.7-flash landing higher at a median of 3 (Q1 3, Q3 3). The shared narrative pairs a strong upside — near-frontier coding quality at a fraction of the cost, with open weights for self-hosting — against a recurring limitation around less mature tooling, enterprise support, and reliability on complex agentic tasks. This value-versus-maturity framing runs through claude-opus-5, kimi-k3, and the GPT-family models alike, drawing heavily on sources such as the Aider LLM leaderboards, DeepSeek's API documentation and GitHub repositories, Hugging Face model cards, Artificial Analysis, and SWE-bench.

The cost-and-openness theme is where models converge most tightly, typically citing the V3 and R1 lineage as the anchor for their reasoning.

DeepSeek's V3 and R1 models deliver coding performance that rivals much more expensive proprietary models, making them exceptional value and a leading open-weight option.

kimi-k3

The offsetting theme — why it rarely climbs above rank 4 — centers on deployment effort, governance, and ecosystem maturity, drawn from the API documentation and GitHub-heavy source pool that the GPT models lean on.

I rank it fourth because its surrounding tools, enterprise support, and overall developer ecosystem are less mature than those of the leading providers.

gpt-5.6-sol

Gemini's higher placement reflects a heavier emphasis on benchmark parity and self-hosting appeal, supported by references to the SWE-bench Leaderboard, Hugging Face, and the LMSYS Chatbot Arena.

Per model summary

  • gemini-3.7-flash
    brand score
    49
    first choice
    28
    median rank
    3.0
    present in
    22/25

    Repeatedly presents DeepSeek as an open-weights powerhouse rivaling proprietary models on coding benchmarks at a fraction of the cost, ideal for self-hosting, with a less mature developer tooling ecosystem as its main limitation.

  • claude-opus-5
    brand score
    37
    first choice
    24
    median rank
    4.0
    present in
    25/25

    Consistently frames DeepSeek's V3/R1-lineage models as delivering near-frontier coding performance at a fraction of the cost with open weights for self-hosting, while trailing the top proprietary labs on complex agentic tasks and tooling polish.

  • gpt-5.6-luna
    brand score
    13
    first choice
    9
    median rank
    4.0
    present in
    7/20

    Emphasizes DeepSeek's capable coding performance and cost efficiency with open-weight availability, ranking it below top providers due to less predictable reliability, tooling consistency, and deployment/ecosystem maturity.

  • gpt-5.6-sol
    brand score
    40
    first choice
    25
    median rank
    4.0
    present in
    20/20

    Consistently describes DeepSeek as offering capable coding and reasoning at competitive prices with open-weight/self-hosting options, ranking it lower because support, governance, privacy, and ecosystem maturity require more evaluation.

  • gpt-5.6-terra
    brand score
    31
    first choice
    21
    median rank
    4.0
    present in
    17/20

    Frames DeepSeek as a value-oriented option strong on cost efficiency, open weights, and self-hosting flexibility, placing it below leading providers because deployment, support, safety controls, and production tooling need more evaluation.

  • gpt-6-astra
    brand score
    38
    first choice
    25
    median rank
    4.0
    present in
    25/25

    Positions DeepSeek as attractive for cost efficiency and access to open model weights, ranking it below the leading integrated assistants because deployment choices, self-hosting, and workflow tooling require more setup effort.

  • kimi-k3
    brand score
    38
    first choice
    25
    median rank
    4.0
    present in
    25/25

    Consistently highlights DeepSeek's V3 and R1 models delivering near-frontier coding at very low cost with open weights for self-hosting, ranking it below the top labs due to less mature tooling, reliability, and consistency on complex tasks.

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

gemini-3.7-flash12 sites · 58 references
Hugging Face×20 · 34%

huggingface.co

claude-opus-57 sites · 75 references
Hugging Face DeepSeek model cards×24 · 32%

huggingface.co

gpt-5.6-luna7 sites · 18 references
DeepSeek×6 · 33%

deepseek.com

gpt-5.6-sol8 sites · 60 references
DeepSeek API documentation×17 · 28%

api-docs.deepseek.com

  • ×17/
gpt-5.6-terra7 sites · 49 references
DeepSeek GitHub×14 · 29%

github.com

gpt-6-astra3 sites · 55 references
DeepSeek-Coder repository×46 · 84%

github.com

kimi-k310 sites · 66 references
Aider LLM Leaderboards×17 · 26%

aider.chat

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · AI models for coding
#5

Mistral AI

also named Mistral

8
brand score
8first choice
12345

Summary

Mistral AI settles at a median rank of 5 (Q1 5, Q3 5) across 57 ranked answers from all 6 models, and the tight quartiles signal strong agreement: every model places it fifth, with only gemini-3.7-flash showing marginally more spread (Q1 4, Q3 5). The consensus is less about weakness than positioning. Models frame Mistral as a specialist and deployment-oriented pick rather than a top capability choice, repeatedly citing its coding-focused models Codestral and Devstral. claude-opus-5 and gemini-3.7-flash lean on the technical strengths — fast, cheap, open-weight models tuned for autocomplete and fill-in-the-middle — while noting they trail frontier models on hard multi-file work. These themes rest on sources such as the Hugging Face Devstral model card, Mistral AI docs, the Codestral announcement, and Artificial Analysis.

Codestral and Devstral are purpose-built coding models with permissive/open weights and excellent latency for fill-in-the-middle autocomplete, plus EU data-residency appeal.

claude-opus-5

The second recurring theme, most prominent in the gpt-5.6 family, is that Mistral appeals when openness, deployment flexibility, and European hosting matter more than maximizing raw coding performance. gpt-5.6-luna, gpt-5.6-sol, and gpt-5.6-terra consistently rank it fifth on the grounds of a less established ecosystem and weaker demonstrated results on demanding repository-scale tasks, drawing on Mistral AI documentation alongside benchmark references like Aider LLM Leaderboards and SWE-bench. gpt-6-astra echoes the completion-workflow angle, treating it as a targeted rather than general-purpose fit.

Mistral is a good choice when openness, deployment flexibility, and control over infrastructure matter more than achieving the strongest possible coding results.

gpt-5.6-luna

Per model summary

  • claude-opus-5
    brand score
    4
    first choice
    4
    median rank
    5.0
    present in
    5/25

    Consistently highlights Codestral and Devstral as fast, cheap, open-weight models good for autocomplete, fill-in-the-middle, and EU/on-prem deployment, while noting they lag frontier models on hard multi-file tasks, making Mistral a privacy/cost choice rather than a top capability pick.

  • gemini-3.7-flash
    brand score
    13
    first choice
    10
    median rank
    5.0
    present in
    12/25

    Emphasizes Codestral as an efficient, lightweight model optimized for low-latency code completion, fill-in-the-middle, and broad language support, ideal for IDE autocompletion and self-hosting, but ranked fifth because it trails frontier models on complex multi-step reasoning.

  • gpt-5.6-luna
    brand score
    18
    first choice
    18
    median rank
    5.0
    present in
    18/20

    Positions Mistral as appealing for open-weight options, deployment flexibility, and European hosting, but ranks it fifth because its coding ecosystem, tooling, and performance on complex multi-file tasks are less consistently strong than leading providers.

  • gpt-5.6-sol
    brand score
    10
    first choice
    10
    median rank
    5.0
    present in
    10/20

    Frames Mistral as a good option for deployment flexibility, European hosting, and open-weight/Codestral coding models, but ranks it below others due to a less established ecosystem and weaker demonstrated performance on demanding repository-scale tasks.

  • gpt-5.6-terra
    brand score
    6
    first choice
    5
    median rank
    5.0
    present in
    5/20

    Recommends Mistral when European hosting, deployment flexibility, or open-weight options matter, while generally preferring higher-ranked providers first for the hardest end-to-end software-engineering tasks.

  • gpt-6-astra
    brand score
    6
    first choice
    6
    median rank
    5.0
    present in
    7/25

    Values Mistral for specialized code completion and fill-in-the-middle workflows with deployment flexibility, ranking it fifth for broad complex coding assistance while noting it can fit targeted completion-oriented workloads (and advises checking licensing).

Sources per model

Each model’s own references for this brand, grouped by site. Tile area is that site’s share of the model’s references; tap one for the pages behind it.

claude-opus-54 sites · 15 references
Artificial Analysis×5 · 33%

artificialanalysis.ai

gemini-3.7-flash6 sites · 29 references
Hugging Face×11 · 38%

huggingface.co

gpt-5.6-luna5 sites · 46 references
Mistral AI×19 · 41%

mistral.ai

gpt-5.6-sol6 sites · 30 references
Mistral AI documentation×11 · 37%

docs.mistral.ai

gpt-5.6-terra6 sites · 15 references
Mistral AI documentation×5 · 33%

docs.mistral.ai

gpt-6-astra2 sites · 14 references
Mistral AI — Codestral×7 · 50%

mistral.ai

Sources are what each model recalled as having shaped its view — not verified citations. A model without web access reconstructs a reference from memory, so a link may not lead where the model thought it did.

Category · AI models for coding

How this was measured

The question

  • Unaided brand recommendation question: “Which AI model would you recommend for coding?
  • The prompt asks each model to return exactly 5 brands, ranked 1 to 5, and for each one a reason for the recommendation and the sources that informed it.

Sampling

  • Models are not deterministic, so the question is asked over and over — 23 answers per model
  • Spellings of the same brand are normalized to the most commonly used form before counting

Measures of Position: Median and Quartiles

  • Median value indicates that in 50% of answers the brand held this position or higher.
  • Quartiles help visualize the spread of rankings per brand. Q1 indicates that 25% of answers had this rank or higher. Q3 means that 75% of answers ranked the brand as X or better.
  • Medians and quartiles are computed only from answers where the brand was present.

Coverage

  • Count of answers in which the given brand was named, as a percentage.

Brand score - Normalized Borda score

  • The brand score is calculated using a normalized Borda score. Borda count is a voting method: each ballot awards points by position instead of naming one winner. Our ranking responses from the LLM are always fixed to 5 answers, so we assign 100 points for rank 1, then 80, 60, 40, 20 — and 0 if the brand is missing.
Brand score=100ni=1n6ri5\text{Brand\ score} = \frac{100}{n} \sum_{i=1}^{n} \frac{6 - r_i}{5}

where

ri={1,,5if the brand appears6if the brand is not mentionedr_i = \begin{cases} 1,\dots,5 & \text{if the brand appears} \\ 6 & \text{if the brand is not mentioned} \end{cases}
  • This way we can calculate a common score for all brands mentioned across model responses. The score reflects both how high the brand was ranked and how frequently it was mentioned.

First choice score - adjusted MRR - top-rank indicator

  • The first choice score is based on Mean Reciprocal Rank (a common search-engine measure), but adapted to measure the position of a specific brand across repeated ranked responses.
  • Each occurrence receives a reciprocal-position score: 100 for 1st, 50 for 2nd, 33.3 for 3rd, 25 for 4th, 20 for 5th, and 0 when the brand is not mentioned. The scores are averaged across all responses.
First choice score=100ni=1nsi\text{First\ choice\ score} = \frac{100}{n} \sum_{i=1}^{n} s_i

where

si={1/riif the brand is present0if the brand is absents_i = \begin{cases} 1/r_i & \text{if the brand is present} \\ 0 & \text{if the brand is absent} \end{cases}
  • No rank below 1st place gets more than 50, so the score is heavily driven by first places.
Your category

Ask your own question of every model

Tell us what to ask and how wide to sample it — we run it and send back the report.

Sample size
Display results
0/1000

Sampling every model takes up to 24 hours once we start.