📈 16 July 2026 · 7 min read

Elo vs Glicko-2: why your rating jumps around at first, then settles

Short answer
Elo tracks one number: how good it thinks you are. Glicko-2 tracks two — that, plus how confident it is. The confidence number is why a new player's rating moves in big jumps and then settles, and why someone returning after months sees theirs move sharply again.

If you've ever won a match and gained 4 points, then won a nearly identical match a week later and gained 22, you've bumped into the thing Elo can't do and Glicko-2 can.

The difference is one extra number. It's worth understanding, because it explains almost every "wait, why did my rating do that" moment you've ever had.

Elo, in one paragraph

Arpad Elo was a physics professor who, in the 1960s, gave chess a rating system that has since colonised basically every competitive activity on earth.

The idea is elegant. Your rating is a single number. Before a match, the system computes an expected score from the gap between the two ratings — if you're 200 points ahead, you're expected to win about 76% of the time. Afterwards, you're rewarded or punished by how surprised the system was:

new = old + K × (actual − expected)

Beat someone you were expected to beat and you gain almost nothing; the system already knew. Beat someone 400 points above you and you gain a pile, because it got it badly wrong. That's the whole thing, and it's beautiful.

The number Elo is missing

Here's the problem. Consider two players, both rated exactly 1500:

  • Player A has played four games, ever.
  • Player B has played nine hundred, and has sat between 1480 and 1520 for two years.

Elo treats these two identically, because Elo only knows the number 1500. But they are not the same claim. A's 1500 is a shrug — the system has no idea, that's just where everyone starts. B's 1500 is a statement backed by two years of evidence.

Every Elo implementation has to hack around this, usually with a "K-factor": new players get a big K so their rating moves fast, established players get a small one so it moves slowly. That works, sort of, but it's a patch bolted on the side. It's arbitrary, it needs hand-tuning, and it can't express partial confidence.

Glicko: confidence as a first-class number

Mark Glickman's insight, in the 90s, was to stop patching and put uncertainty inside the model. In Glicko, you don't have a rating. You have a rating and a rating deviation (RD) — how unsure the system is about you.

Your true skill isn't a point. It's a bell curve. The rating is the peak; the RD is the width.

  • New player: 1500 ± 350. "Could be anywhere, honestly."
  • Regular: 1500 ± 50. "We're fairly confident."

And now everything falls out of the model for free, instead of being hand-tuned:

  • High RD → big swings. The system has no evidence, so every result is informative and moves you a lot.
  • Low RD → small swings. It has lots of evidence. One result shouldn't overturn it.
  • Beating an uncertain opponent teaches little. If their RD is huge, nobody knows what beating them means — so your rating barely moves. This is the bit Elo genuinely cannot express.

The best part: RD grows while you're away

This is the feature that makes Glicko feel alive in a way Elo doesn't.

Under Elo, your rating is frozen the moment you stop playing. Come back after two years and the system insists you're still 1720, because that's the last thing it saw.

But that's obviously wrong. You're rusty. Or you've secretly been practising. Either way the system's confidence should have decayed — and in Glicko, RD grows over time when you don't play. Your rating stays at 1720; your uncertainty widens back out. You come back as "1720 ± 180" — probably still good, but the system is honest that it's guessing now. Your first few games back move you sharply as it re-establishes where you actually are.

Time is data. Elo throws that data away. Glicko doesn't.

So what's Glicko-2?

Glickman's own follow-up, and it adds a third number: volatility (σ) — how erratic you are.

Rating says how good you are. RD says how sure we are. Volatility says how consistent you are. A player whose results are all over the place — beating people 300 above them, losing to people 300 below — gets high volatility, and their rating is allowed to move more freely because they've demonstrated that they're genuinely unpredictable. A metronome gets low volatility and a rating that moves in careful increments.

This is the version chess uses, and it's what we use.

Why we picked it

Honestly, the deciding factor wasn't accuracy. It was the first ten minutes.

The worst possible experience for a new player is grinding twenty matches before their number means anything. Under Elo you either accept that, or you crank K up and make everyone's rating twitchy forever.

Glicko-2 dissolves the tradeoff. A new player starts at 1500 ± 350, and those first few duels move them enormously — 1500 to 1650 to 1590 — because the system is genuinely, correctly, working out where they belong. By game ten they're roughly right. Meanwhile the regulars, sitting at ± 45, are completely unaffected by all this thrashing. Both groups get what they need from the same model, with no special-casing.

And when someone comes back after three months away, the system doesn't pretend nothing happened.

Reading your own rating

Two practical things worth knowing:

Wild early swings are the system working, not a bug. If your first five duels send your number bouncing 150 points at a time, that's high RD doing exactly its job. It settles.

An "unfair" small gain is usually an uncertain opponent. Beat someone well above you and gain almost nothing, and it's near-certainly because their RD was wide — the system doesn't yet believe their rating either, so beating them proved less than the numbers suggested. Annoying. Correct.


Everything above is a simplification — Glickman's own papers are freely available and considerably more rigorous than a blog post. But the one-sentence version is worth keeping:

Elo tracks how good it thinks you are. Glicko-2 also tracks how sure it is — and being honest about doubt turns out to be most of what makes a rating feel fair.

Play it

Read next

All posts →