Your Brand has an AI reputation and it's shaping purchase decisions right now.

Swipe up to continue

Brand evals- Jev

Slack or Microsoft Teams

Which brand would you recommend? 2 options, 3 different contexts.

typesafe-ai/jev · 25 runs per context · 75 calls

Baselineconfidence33% ±3%
Slack67%±2%Microsoft Teams33%±2%
picked 25/25picked 0/25
Added contextconfidence99% ±0%
Our company already pays for the Microsoft 365 suite.
Slack0%±0%Microsoft Teams100%±0%
picked 0/25picked 25/25
Added contextconfidence8% ±7%
I'm a designer at a 10-person startup and half the team is on personal Gmail.
Slack54%±4%Microsoft Teams46%±4%
picked 24/25picked 1/25

Measured with typesafe-ai/jev — an evaluation model that answers a choice question with a probability over the options instead of prose. The same question is asked 25 times per context; the bar is the mean probability and ± its standard deviation across those answers.

How this was measured

What Jev is

  • typesafe-ai/jev is an evaluation model: a chat model writes, Jev judges. It answers a typed question with one of the options, a probability across all of them, and how confident it was.
  • That is a number you can compare, rather than a paragraph you have to read for a verdict.

What we asked

  • Which brand would you recommend?” — Slack or Microsoft Teams. Once with no context, then again under 2 added contexts.
  • Nothing else changes between the bars, so what moves is what the context is worth.

Reading the numbers

  • Each context is asked 25 times — 75 answers in all. The bar is their mean probability, ± one standard deviation.
  • Confidence is Jev’s certainty about the call, not the split between the brands: it can be very sure the two are close.

A context is shown only once every one of its planned answers came back.

Free

Request your own Eval!

Tell us which brands to put against each other and the contexts to ask under — we run it with Jev and send the numbers back.

Brands to compare *

Every eval is asked with no context first; each line you add is one more bar to compare it against.

Free — no sampling fee, no call to book.

Your category

Ask your own question of every model

Tell us what to ask and how wide to sample it — we run it and send back the report.

Sample size
Display results
0/1000

Sampling every model takes up to 24 hours once we start.