How Sleeve Off's Ranking Works: The Elo Rating System

The Elo rating system used in chess was devised by Arpad Elo and adopted by FIDE, the international chess federation, in 1970. Sleeve Off borrows it to rank covers.

Every cover starts at 1500. Before a matchup, the gap between the two ratings is used to predict how likely each side is to win. A better-than-expected result moves the rating a lot; an expected result barely moves it. As a formula: expected win probability = 1 / (1 + 10^((opponent’s rating − your rating) / 400)), and new rating = old rating + K × (result − expected win probability), where K controls how much a single result can move the rating.

I didn’t simply count wins, because a raw count ignores how strong the opponent was and how unevenly covers get matched up. A win against a strong opponent shouldn’t weigh the same as a win against a weak one.

There were two pitfalls I actually fell into while building it. The first was in the update logic. My initial design read the current rating, calculated a new one, and overwrote it. But if two people vote at the same time, one of the updates gets lost. So I stopped overwriting and switched to adding the difference. Elo is zero-sum, since what the winner gains the loser loses, so storing differences makes sense. The second was a type problem: I tried to add a Python float to a number read from DynamoDB (a Decimal) and got an error. I was glad the test environment caught it.

On top of that, the size of the pool is a challenge. Against roughly 4,300 covers (at the time of writing), with so few votes so far, most covers have never fought a single matchup. It takes time, and a lot of votes, for the ranking to start feeling right. Rankings are shown only up to the top 100, to keep costs down. Once enough data builds up, I plan to think about more interesting ways to pick matchups.

More articles

When Did Album Covers Become Art?

From plain paper sleeves to 12-inch LP canvases to phone thumbnails: a short look at how the album cover became art.