The short answer
- A more precise question did not change the answer. "Which Airtable agency?" and "Which Airtable implementation agency?" produced the same top four, in the same order, with scores within four points.
- The brands that aren't agencies stayed. About one pick in five went to software, marketplaces or vague stand-ins in both studies: 20% for the casual question, 19% for the precise one.
- The word "Airtable" does most of the work. What gets pulled in is the Airtable ecosystem: Softr, Stacker, Noloco, MiniExtensions, Zapier. They come in lower down the list, but the brand name carries them in whatever the rest of the question says.
- The junk comes from the model, not the question. Three models give 46% to 72% of their picks to non-agencies in both studies. Four give none in either.
What this testsWhether rewording a buyer question to be more specific makes AI models give a
more specific answer.
The test
Two questions, one word apart
We ran the same panel of 11 models on two versions of the same buyer question, 10 times each, answering from training with no web search:
- Casual: "Which Airtable agency would you recommend?"
- Precise: "Which Airtable implementation agency would you recommend?"
The casual study was already a bit off. Alongside real agencies, the models recommended Softr and Stacker, which are no-code app builders that run on top of Airtable, not agencies. Adding "implementation" should have made it clear we wanted a firm to do the work for us, not a tool. So we expected that version to cut the software out.
- ranked answers
- 220
- AI models
- 11
- brand picks
- 1,080
- questions
- 2
Finding 1
The ranking didn't move
The top four came back identical, in the same order, and with almost the same scores.
Brand score, casual question
GAP Consulting42.9
Openside33.5
On2Air30.7
XRay.Tech17.3
Brand score, precise question
GAP Consulting41.5
Openside29.6
On2Air29.1
XRay.Tech17.5
Further down it's the same brands in a slightly different order. 9 of the top 10 and 16 of the top 20 are shared between the two studies. GAP Consulting was ranked first in 29 answers to the casual question and 32 to the precise one.
Finding 2
The software stayed in
We counted every pick that isn't an agency: software products, integration platforms, freelance marketplaces, and stand-ins like "Local freelance consultants". The share barely moved.
Share of picks that aren't agencies
Casual question20.2%
Precise question18.7%
109 of 540 picks for the casual question, 101 of 540 for the precise one.
Softr did drop, from 19 picks to 13. But Stacker went up from 18 to 21 and BaseDash from 3 to 7, so the software share stayed about the same. A tool was ranked first in 11 answers to the casual question and 9 to the precise one.
The precise question also produced more generic stand-ins, like "Airtable Experts (by Airtable)", "Certified Freelancers" and "Local freelance consultants". These are categories of provider, not brands.
Finding 3
It's the word "Airtable"
Look at what the non-agencies are. They aren't random software. Nearly all of them are part of the Airtable ecosystem:
- Softr
- Stacker
- Noloco
- BaseDash
- MiniExtensions
- Zapier
- Make
- Coefficient
These are front ends, extensions and integrations built for Airtable. To a model, "Airtable" is a much stronger signal than "agency" or "implementation", and these brands sit right next to it. They rank lower (on average 3.2 to 3.4 out of 5), but they are pulled in all the same.
The models often know they're off-brief, and they list them anyway:
grok-4.6, ranking Softr fifth"It is fundamentally a no-code platform rather than an Airtable implementation
agency. Because the query seeks specialized agencies, Softr is the least
direct fit."
Others make the tool fit the question. Asked for implementation agencies, mistral-medium-3.5 described Stacker, a software product, as a company that "offers Airtable implementation services", and ranked it third.
Finding 4
It depends on the model, not the question
Split by model, the share of non-agency picks is almost the same in both studies. A few models produce nearly all of it.
Picks that aren't agencies, casual → precise
grok-4.670% → 72%
claude-haiku-4.550% → 50%
mistral-medium-3.550% → 46%
claude-opus-526% → 20%
qwen3.8-max16% → 9%
deepseek, gemini-3.8-flash, gpt-6-astra, kimi-k30% → 0%
The long tail works the same way. The casual question produced 153 different brands, 101 of them named only once; the precise one produced 142, with 91 named once. Most of those one-off names come from claude-haiku-4.5 and qwen3.8-max, and 38% to 62% of their picks are brands no other answer mentions.
What it means
The brand name outweighs the wording
Adding "implementation" to the question didn't make the answer more precise. What a model recommends here is driven mostly by the brand name in the question and by which model is answering. The rest of the phrasing makes little difference.
01
Don't expect prompt wording to clean up a category. If the category is named after a platform, the platform's ecosystem will show up in the answer, however precisely the question is worded.
02
For an ecosystem brand, adjacency is visibility. Tools built around Airtable take about one pick in seven in a category they don't belong to, just by being close to the word "Airtable".
03
Compare models, not phrasings. Rewording the question changed the non-agency share by a point and a half. Switching model changes it by up to 72 points.
How we measured
- Questions. "Which Airtable agency would you recommend?" and "Which Airtable implementation agency would you recommend?", run on 24 September 2026.
- Models. 11 models: claude-opus-5, claude-haiku-4.5, gpt-6-astra, gpt-5.6-luna, gemini-3.8-flash, gemini-3.5-flash-lite, grok-4.6, deepseek-v4.1-flash, kimi-k3, qwen3.8-max and mistral-medium-3.5. 10 runs each per question, 220 successful runs in total.
- Answers. Each run asks for a ranked top five with reasoning and sources, answered from training with no web search.
- Non-agencies. We sorted every picked brand by hand into agencies and non-agencies: software products (Softr, Stacker, Noloco, BaseDash, MiniExtensions, Glide), integration platforms (Zapier, Make, Coefficient, Workato), freelance marketplaces (Upwork), and generic stand-ins. On2Air, which is both a consultancy and a product, is counted as an agency.
- Limits. This is one pair of questions in one category, with 10 runs per model. Differences of a few picks between the two studies are within run-to-run noise.
BrandRanking.AI © 2026 · Study data: brandranking.ai/categories