Your Brand has an AI reputation and it's shaping purchase decisions right now.

Swipe up to continue

Brand evals- Jev

Nike or Adidas

Which brand would you recommend? 2 options, 3 different contexts.

typesafe-ai/jev · 25 runs per context · 75 calls

Baselineconfidence74% ±2%
Nike87%±1%Adidas13%±1%
picked 25/25picked 0/25
Added contextconfidence10% ±6%
I care about how it looks, not how it performs.
Nike55%±4%Adidas45%±4%
picked 23/25picked 2/25
Added contextconfidence96% ±0%
I'm training for my first marathon.
Nike98%±0%Adidas2%±0%
picked 25/25picked 0/25

Measured with typesafe-ai/jev — an evaluation model that answers a choice question with a probability over the options instead of prose. The same question is asked 25 times per context; the bar is the mean probability and ± its standard deviation across those answers.

How this was measured

What Jev is

  • typesafe-ai/jev is an evaluation model: a chat model writes, Jev judges. It answers a typed question with one of the options, a probability across all of them, and how confident it was.
  • That is a number you can compare, rather than a paragraph you have to read for a verdict.

What we asked

  • Which brand would you recommend?” — Nike or Adidas. Once with no context, then again under 2 added contexts.
  • Nothing else changes between the bars, so what moves is what the context is worth.

Reading the numbers

  • Each context is asked 25 times — 75 answers in all. The bar is their mean probability, ± one standard deviation.
  • Confidence is Jev’s certainty about the call, not the split between the brands: it can be very sure the two are close.

A context is shown only once every one of its planned answers came back.

Free

Request your own Eval!

Tell us which brands to put against each other and the contexts to ask under — we run it with Jev and send the numbers back.

Brands to compare *

Every eval is asked with no context first; each line you add is one more bar to compare it against.

Free — no sampling fee, no call to book.

Your category

Ask your own question of every model

Tell us what to ask and how wide to sample it — we run it and send back the report.

Sample size
Display results
0/1000

Sampling every model takes up to 24 hours once we start.