The Confidence Game
Play How it works The bank Live

You answer with a probability.

Not true or false — a number. How sure are you? Then Jev, a model trained to have honest probabilities, answers the same claim. At the end you get your reliability diagram next to the machine's.

You can't win by bluffing. The score is a strictly proper scoring rule, so your best possible strategy is to report exactly what you believe. Saying 90% to look bold costs you. Saying 50% to play safe costs you too. How the scoring works →

—
claims in the bank
—
Jev's Brier score
—
cost to build the bank
0
API keys needed to play
Pick a run

50% likely to be true
certainly falseno ideacertainly true
You said
Jev said
Coin flip
50%
0.250

YouJevCoin flip
By category
Share it

X and LinkedIn preview the headline as text; attach the PNG for the picture.

What this is

A calibration test with an opponent.

Calibration is the match between how sure you say you are and how often you turn out to be right. If you say “90%” a hundred times, about ninety of them should happen. Almost nobody is calibrated. Most people are overconfident, and the gap is invisible from the inside — you need to be scored to see it.

Training tools for this have existed for years: Quantified Intuitions, Open Philanthropy's calibration app, Clearer Thinking's quiz, Metaculus practice questions. The research says brief practice measurably helps. Every one of those tools is human-only, and it had to be: a chat model's stated “90%” is just generated text, and the probability of it typing those characters tells you nothing about whether it's right.

How a round works The scoring What we measured What it's useful for Fairness Honest limits
The gap this fills Calibration training exists. Nobody has ever had a calibrated opponent to measure against. Jev's probabilities are trained against outcomes, so — unusually — the number it gives you means something.

How a round works

You get a claim with mechanical ground truth — generated from World Bank population figures and a country-borders dataset, never hand-written. You answer with a probability between 1% and 99%. Your answer is locked on the server before anything is revealed, then you see the truth, Jev's probability for the same sentence, and what each of you scored.

After 20 (or 100) of those you get a scorecard: a reliability diagram with your curve against Jev's, a Brier score, the Murphy decomposition, and a verdict.

The scoring

Every question is binary. You give p = P(claim is true). The outcome o is 0 or 1.

Brier score = (1/N) · Σ (pᵢ − oᵢ)² lower is better

That's mean squared error on probabilities, and it is strictly proper: your expected score is minimised by reporting exactly what you believe. There is no clever strategy. Three anchors worth remembering:

  • Answering 0.5 to everything scores exactly 0.250, whatever the outcomes. That's the coin-flip column, and it's always on screen.
  • Perfect and certain scores 0.
  • Confidently wrong every time scores 1.

Scoring worse than 0.250 means you'd have done better by admitting you had no idea. That happens to most people, and it's the point of the game.

Why one number isn't enough

A Brier score mixes up two completely different virtues. The Murphy decomposition separates them:

Brier = Reliability − Resolution + Uncertainty ↓ calibration ↑ knowledge the quiz's own difficulty

Reliability asks: when you said 80%, did it happen 80% of the time? Zero is perfect — and it's free. You can drive it to zero without knowing a single fact, by answering 50% every time. Resolution asks whether your answers carried any information at all. High resolution means you actually separated the true claims from the false ones.

ReliabilityResolutionWhat you are
lowlowCalibrated but cluelessAnswered 50% to everything. Unfalsifiable and useless.
lowhighWell calibratedRare. Confident when right, hedged when not.
highhighKnows things, oversells themThe expert failure mode.
highlowConfidently wrongThe worst square.

A single score can't tell those apart. That's the argument the whole game is making, and it's why the scorecard shows three numbers instead of one.

What we measured about Jev

The design this was built from claimed Jev is bad at numbers, so humans would win the numeric categories. The data says otherwise. Jev got 100% of the easy numeric comparisons right. It only degrades as the two values get close — and as it does, its confidence drops in step:

Numeric claimsJev's accuracyJev's mean confidence
one value ≥ 8× the other100%0.97
2.5–8×98%0.86
1.3–2.5×80%0.74
closer than 1.3×61%0.68

It doesn't get close comparisons wrong so much as it gets uncertain about them. So there's no category where you beat this thing by knowing more. On its worst subset it still edges the coin flip — Brier 0.239 against 0.250.

The only gap you can exploit On those closest comparisons Jev's reliability is 0.022. A player who simply types 50% on all of them has reliability 0.000 — perfectly calibrated — and beats Jev outright in 37% of rounds. You don't win by knowing more. You win by admitting what you don't know. That's Jev's blind spot mode.

What it's useful for

Forecasting practice

If you write on Metaculus, Good Judgment or a prediction market, this is a three-minute drill with an opponent whose numbers are worth comparing to.

Teaching probability

Proper scoring rules, reliability diagrams and the calibration/resolution split are hard to motivate on a whiteboard and obvious after one bad scorecard.

A worked reference

The scoring module is pure, tested against hand-computed fixtures, and documented with derivations. Useful if you're implementing Brier or Murphy yourself and want something to check against.

Evaluating model calibration

The bank, the fixtures and the eval are committed. Point the same 937 claims at another model and the comparison is apples to apples.

Team decision hygiene

Run it before a planning meeting. Knowing who in the room is systematically overconfident changes how you read the room's estimates.

Seeing what a calibrated model looks like

Bring your own key and ask it anything. Watching a model answer 0.62 instead of bluffing 0.95 is the fastest way to understand what “calibrated” buys you.

Fairness, and how each rule is enforced

  • Jev sees exactly the claim you see. The request is the claim text and one question, Is the claim in the state true? — nothing appended, nothing explained.
  • One claim per request. Packing claims into a shared request doubles Jev's Brier score and flips individual answers. Measured, not assumed.
  • Jev answered before you played. Answers were fetched once and cached, keyed by a hash of the claim text. It never sees your answer and can't change its mind. You can check that yourself.
  • Your forecast locks server-side. The truth and Jev's number aren't in your browser until you've committed.
  • Every round is exactly 50/50 true/false, and evenly split across categories and difficulties — so "always say true" scores 0.25, same as the coin flip.
  • Everything is published. The bank, the fixtures, the scoring code. Browse the bank →

Honest limits

Twenty questions is not a calibration measurement At 20 rounds the 95% interval on your accuracy is about ±20 points — wide enough to contain almost any conclusion. The scorecard shows that interval every time. Play the short one for fun; don't post the verdict as a fact about yourself. The 100-question run is the one worth quoting.
The bank is Western-skewed and English-only

Claims come from World Bank population figures and a country-borders dataset, so "general knowledge" here means a particular kind of general knowledge. Someone who has never needed to know whether Zambia or Senegal is larger isn't badly calibrated — they're being asked the wrong questions.

Difficulty labels aren't validated on humans

They're construction parameters — a ratio between two values, a hop count in the border graph. They predict Jev's accuracy monotonically (100% → 99% → 89% → 76%). Whether they predict yours is untested.

Jev isn't perfectly deterministic

Identical repeated requests differ by a median of 0.01, and up to 0.12 in the worst case we measured — six times the documented ±0.02. No claim flipped sides across eight repeats, so accuracy is stable and Brier wobbles in the third decimal. It's also why answers are cached rather than fetched live: two players on the same question set would otherwise see different opponents.

Sixteen countries are excluded from border questions

Their land adjacency depends on a disputed or overseas territory — France borders Brazil via French Guiana, Spain borders the UK via Gibraltar. A claim whose truth depends on which reading you take isn't a calibration question, it's a trick. The full list and the reason for each is in the published bank.

There's no human data here yet

Everything about player archetypes in the repo is simulated from an invented skill model. It validates that the scorecard separates player types; it is not data about people.

The published bank

Every claim, every answer.

All 937 claims, their ground truth, and the probability Jev gave each one — the same data the game scores you against. Sort by Jev's worst to see where a calibrated model gets confidently wrong.

Green pill: Jev on the right side of 50%. Red: the wrong side. The number is its probability that the claim is true. Download the raw files: bank.json · jev-fixtures.json

Bring your own key

Ask the model yourself.

The game needs no key — it runs on cached answers. This page is for the two things caching can't do: asking Jev something of your own, and checking that our cached answers are real.

Where your key goes Jev refuses requests from browser origins, so a key has to pass through this server to reach it. It is never stored — not on disk, not in a session, not in a log. It's used for one upstream call and discarded. In your browser it lives in memory only, unless you tick “remember in this tab”. If you'd rather not paste a key into someone else's page, that's the correct instinct: clone the repo and run it locally.
No key set.

Free key with $5 of credit at console.typesafe.ai/keys. A single question costs about $0.000012 — the credit is roughly 400,000 questions.

Ask your own claim

Write a statement that's either true or false. Commit your own probability first — then see what a calibrated model says. That ordering is the whole exercise.

50% — your probability
certainly falseno ideacertainly true

Jev is asked exactly Is the claim in the state true? about your text — the same question the whole bank was built with. There's no ground truth here, so nothing is scored: this is a comparison, not a round.

Verify the cached answers

The game's fairness case rests on Jev having answered before you played. You don't have to take that on trust. This re-asks the live model a random sample of cached claims and shows the drift.

Expect small non-zero drift: the model isn't perfectly deterministic. Our own measurement over 192 repeated calls found a median spread of 0.01 and a worst case of 0.12, with no claim ever changing which side of 50% it was on. A flip here would be notable; a delta of 0.01 is the model being itself.

The Confidence Game · scored on Brier, not accuracy Source & data Docs Jev by TypeSafe Quantified Intuitions