You answer with a probability.
Not true or false — a number. How sure are you? Then Jev, a model trained to have honest probabilities, answers the same claim. At the end you get your reliability diagram next to the machine's.
You can't win by bluffing. The score is a strictly proper scoring rule, so your best possible strategy is to report exactly what you believe. Saying 90% to look bold costs you. Saying 50% to play safe costs you too. How the scoring works →
| You | Jev | Coin flip |
|---|
X and LinkedIn preview the headline as text; attach the PNG for the picture.
A calibration test with an opponent.
Calibration is the match between how sure you say you are and how often you turn out to be right. If you say “90%” a hundred times, about ninety of them should happen. Almost nobody is calibrated. Most people are overconfident, and the gap is invisible from the inside — you need to be scored to see it.
Training tools for this have existed for years: Quantified Intuitions, Open Philanthropy's calibration app, Clearer Thinking's quiz, Metaculus practice questions. The research says brief practice measurably helps. Every one of those tools is human-only, and it had to be: a chat model's stated “90%” is just generated text, and the probability of it typing those characters tells you nothing about whether it's right.
How a round works
You get a claim with mechanical ground truth — generated from World Bank population figures and a country-borders dataset, never hand-written. You answer with a probability between 1% and 99%. Your answer is locked on the server before anything is revealed, then you see the truth, Jev's probability for the same sentence, and what each of you scored.
After 20 (or 100) of those you get a scorecard: a reliability diagram with your curve against Jev's, a Brier score, the Murphy decomposition, and a verdict.
The scoring
Every question is binary. You give p = P(claim is true). The
outcome o is 0 or 1.
That's mean squared error on probabilities, and it is strictly proper: your expected score is minimised by reporting exactly what you believe. There is no clever strategy. Three anchors worth remembering:
- Answering 0.5 to everything scores exactly 0.250, whatever the outcomes. That's the coin-flip column, and it's always on screen.
- Perfect and certain scores 0.
- Confidently wrong every time scores 1.
Scoring worse than 0.250 means you'd have done better by admitting you had no idea. That happens to most people, and it's the point of the game.
Why one number isn't enough
A Brier score mixes up two completely different virtues. The Murphy decomposition separates them:
Reliability asks: when you said 80%, did it happen 80% of the time? Zero is perfect — and it's free. You can drive it to zero without knowing a single fact, by answering 50% every time. Resolution asks whether your answers carried any information at all. High resolution means you actually separated the true claims from the false ones.
| Reliability | Resolution | What you are |
|---|---|---|
| low | low | Calibrated but cluelessAnswered 50% to everything. Unfalsifiable and useless. |
| low | high | Well calibratedRare. Confident when right, hedged when not. |
| high | high | Knows things, oversells themThe expert failure mode. |
| high | low | Confidently wrongThe worst square. |
A single score can't tell those apart. That's the argument the whole game is making, and it's why the scorecard shows three numbers instead of one.
What we measured about Jev
The design this was built from claimed Jev is bad at numbers, so humans would win the numeric categories. The data says otherwise. Jev got 100% of the easy numeric comparisons right. It only degrades as the two values get close — and as it does, its confidence drops in step:
| Numeric claims | Jev's accuracy | Jev's mean confidence |
|---|---|---|
| one value ≥ 8× the other | 100% | 0.97 |
| 2.5–8× | 98% | 0.86 |
| 1.3–2.5× | 80% | 0.74 |
| closer than 1.3× | 61% | 0.68 |
It doesn't get close comparisons wrong so much as it gets uncertain about them. So there's no category where you beat this thing by knowing more. On its worst subset it still edges the coin flip — Brier 0.239 against 0.250.
What it's useful for
Forecasting practice
If you write on Metaculus, Good Judgment or a prediction market, this is a three-minute drill with an opponent whose numbers are worth comparing to.
Teaching probability
Proper scoring rules, reliability diagrams and the calibration/resolution split are hard to motivate on a whiteboard and obvious after one bad scorecard.
A worked reference
The scoring module is pure, tested against hand-computed fixtures, and documented with derivations. Useful if you're implementing Brier or Murphy yourself and want something to check against.
Evaluating model calibration
The bank, the fixtures and the eval are committed. Point the same 937 claims at another model and the comparison is apples to apples.
Team decision hygiene
Run it before a planning meeting. Knowing who in the room is systematically overconfident changes how you read the room's estimates.
Seeing what a calibrated model looks like
Bring your own key and ask it anything. Watching a model answer 0.62 instead of bluffing 0.95 is the fastest way to understand what “calibrated” buys you.
Fairness, and how each rule is enforced
- Jev sees exactly the claim you see. The request is the
claim text and one question,
Is the claim in the state true?— nothing appended, nothing explained. - One claim per request. Packing claims into a shared request doubles Jev's Brier score and flips individual answers. Measured, not assumed.
- Jev answered before you played. Answers were fetched once and cached, keyed by a hash of the claim text. It never sees your answer and can't change its mind. You can check that yourself.
- Your forecast locks server-side. The truth and Jev's number aren't in your browser until you've committed.
- Every round is exactly 50/50 true/false, and evenly split across categories and difficulties — so "always say true" scores 0.25, same as the coin flip.
- Everything is published. The bank, the fixtures, the scoring code. Browse the bank →
Honest limits
The bank is Western-skewed and English-only
Claims come from World Bank population figures and a country-borders dataset, so "general knowledge" here means a particular kind of general knowledge. Someone who has never needed to know whether Zambia or Senegal is larger isn't badly calibrated — they're being asked the wrong questions.
Difficulty labels aren't validated on humans
They're construction parameters — a ratio between two values, a hop count in the border graph. They predict Jev's accuracy monotonically (100% → 99% → 89% → 76%). Whether they predict yours is untested.
Jev isn't perfectly deterministic
Identical repeated requests differ by a median of 0.01, and up to 0.12 in the worst case we measured — six times the documented ±0.02. No claim flipped sides across eight repeats, so accuracy is stable and Brier wobbles in the third decimal. It's also why answers are cached rather than fetched live: two players on the same question set would otherwise see different opponents.
Sixteen countries are excluded from border questions
Their land adjacency depends on a disputed or overseas territory — France borders Brazil via French Guiana, Spain borders the UK via Gibraltar. A claim whose truth depends on which reading you take isn't a calibration question, it's a trick. The full list and the reason for each is in the published bank.
There's no human data here yet
Everything about player archetypes in the repo is simulated from an invented skill model. It validates that the scorecard separates player types; it is not data about people.
Every claim, every answer.
All 937 claims, their ground truth, and the probability Jev gave each one — the same data the game scores you against. Sort by Jev's worst to see where a calibrated model gets confidently wrong.
Green pill: Jev on the right side of 50%. Red: the wrong side. The number is its probability that the claim is true. Download the raw files: bank.json · jev-fixtures.json
Ask the model yourself.
The game needs no key — it runs on cached answers. This page is for the two things caching can't do: asking Jev something of your own, and checking that our cached answers are real.
Free key with $5 of credit at console.typesafe.ai/keys. A single question costs about $0.000012 — the credit is roughly 400,000 questions.
Ask your own claim
Write a statement that's either true or false. Commit your own probability first — then see what a calibrated model says. That ordering is the whole exercise.
Jev is asked exactly Is the claim in the state true?
about your text — the same question the whole bank was built with.
There's no ground truth here, so nothing is scored: this is a comparison,
not a round.
Verify the cached answers
The game's fairness case rests on Jev having answered before you played. You don't have to take that on trust. This re-asks the live model a random sample of cached claims and shows the drift.
Expect small non-zero drift: the model isn't perfectly deterministic. Our own measurement over 192 repeated calls found a median spread of 0.01 and a worst case of 0.12, with no claim ever changing which side of 50% it was on. A flip here would be notable; a delta of 0.01 is the model being itself.