Arctic or Andes?

How confident dare you be?

🧭 The science behind the game

How does this strange scoring system work?

The quiz is not primarily testing whether you can guess Peru or Greenland. It is testing whether you know how sure you should be. The goal is to make your probability estimates match reality.

1. The basic idea

Start here

For every picture you have 100% to distribute between Peru and Greenland. You are not required to pick one and pretend to be certain. You can say 50/50, 60/40, 80/20, or anything in between.

When you submit, the true country is revealed immediately. Your score depends on the probability you assigned to the country that turned out to be correct.

Too uncertain?
If you really think the answer is 80% Peru, putting only 55% on Peru costs you some expected points. You are leaving information value on the table.
Too confident?
If you think Peru is only 60% likely but put 99% on it, you get a spectacular reward when you happen to be right β€” but a spectacular penalty when you are wrong.
The central lesson: confidence is not the same thing as being correct. The scoring system rewards calibrated confidence.

2. The graph

The whole scoring system at a glance

The graph below asks a hypothetical question: suppose that, in situations like this one, your chosen answer would actually be correct q of the time. What happens to your average score if you report a probability p?

q = your actual chance of being right p = the probability you report p = q = perfect calibration
Contour plot of average payout as a function of actual chance of being right and reported probability
Average expected payout for a binary question. The dotted diagonal is perfect calibration; the thick curve is where expected payout is zero.
Annotated average payout graph explaining overconfidence, underconfidence and calibration
The same graph with the main ideas called out.
On the diagonal: if you report 70%, situations like this really are correct about 70% of the time. That is honest probability reporting.
Above or below the diagonal: you are systematically reporting probabilities that do not match your actual success rate. The further away you are, the more you pay in expected score.
Near 100%: a little extra confidence makes a relatively small difference to the reward if you are right, but can make the loss enormous if you are wrong.
At 50/50: if you genuinely know nothing, you get an expected score of zero. You neither claim information you do not have nor lose points for failing to invent it.

3. What do the points actually feel like?

Two possible countries

For a two-option question, if you put probability p on an option, a correct answer gives you 100 Γ— ln(2p) points. If it is wrong, the other probability is 1 βˆ’ p, so you get 100 Γ— ln(2(1 βˆ’ p)) points.

Your probabilityIf correctIf wrong
50%00
60%+18βˆ’22
70%+34βˆ’51
80%+47βˆ’92
90%+59βˆ’161
99%+68βˆ’391

The values are rounded to whole points. The logarithm is deliberately unforgiving near 0% and 100%.

4. Why the logarithm?

For the mathematically curious

Here is the neat part. Suppose your genuine belief that an answer is correct is q, but you report p. Your expected score is

E(p) = 100 [ q ln(2p) + (1 βˆ’ q) ln(2(1 βˆ’ p)) ]

We can ask: which value of p gives the highest expected score?

dE/dp = 100 [ q/p βˆ’ (1 βˆ’ q)/(1 βˆ’ p) ]

Set this equal to zero:

q/p = (1 βˆ’ q)/(1 βˆ’ p)

q(1 βˆ’ p) = p(1 βˆ’ q)

p = q

And the second derivative is

dΒ²E/dpΒ² = βˆ’100 [ q/pΒ² + (1 βˆ’ q)/(1 βˆ’ p)Β² ] < 0

so this is not merely a stationary point: it is the unique maximum. In other words, if you want to maximize your expected score, you should report your actual probability.

🧠 The extra-nerdy information-theory version

The loss from reporting p instead of your true belief q can be written as

E(q) βˆ’ E(p) = 100 [ q ln(q/p) + (1 βˆ’ q) ln((1 βˆ’ q)/(1 βˆ’ p)) ]

The expression in brackets is the Kullback–Leibler divergence between two Bernoulli distributions:

100 Γ— DKL(Bern(q) βˆ₯ Bern(p))

KL divergence is always non-negative and is zero only when p = q. So every deviation from your genuine probability has a measurable expected cost.

This is why the logarithmic score is called a proper scoring rule. It makes honest probability reporting optimal. It is not the only possible proper scoring rule β€” the Brier score is another β€” but the logarithmic score has a particularly direct connection to information and makes extreme overconfidence very expensive.

5. So how do I actually play well?

Practical advice
  1. Separate β€œwhich is more likely?” from β€œhow likely?”
    You might be fairly sure that Peru is more likely than Greenland while still thinking Peru is only 65% likely.
  2. Use your first impression, then ask how often that impression would be wrong.
    That second question is often where useful uncertainty appears.
  3. Pay attention to the downside.
    If an 80% guess is wrong, the βˆ’92 point result is much larger than the +47 you get when it is right. That is not a bug: it is what prevents you from simply choosing the most likely-looking answer at 99% every time.
  4. 50/50 is not failure.
    If the picture genuinely contains no useful information, 50/50 is the mathematically sensible answer.
  5. Learn your own calibration.
    The statistics page shows whether your 60%, 70%, 80% and 90% predictions actually come true at roughly those rates.
A person who gets 16 out of 20 pictures right can still lose to somebody who gets fewer right but is much better calibrated. The game rewards the quality of the probabilities, not just the number of correct guesses.

6. A few practical details

For completeness

How are pictures selected?

Every attempt receives 10 random Peru pictures and 10 random Greenland pictures, shuffled together.

Can I change an answer?

You can change the probabilities while working on the current picture. Once you press Answer, that question is locked and scored.

What happens if I quit?

Your progress is saved after every answer. An abandoned 4/20 attempt remains in the data as a partial quiz.

What do the public statistics show?

The site can show completed high scores, aggregate quiz statistics, calibration statistics, and per-picture difficulty. IP addresses are kept out of the public displays.