Old Light Play now
All posts

Essay

Why my strategy game has no Elo rating

The Old Light leaderboard ranking empires by systems held and structure score

Old Light, the persistent browser strategy game I run, has a leaderboard but no Elo rating. Elo, or one of its MMR descendants, is the default answer whenever a game needs a ranking system, so leaving it out was a real decision, and working out why turned into a decent tour of what a persistent game actually is.

How the Elo rating system works#

Arpad Elo was a physics professor and chess master who designed the Elo rating system for the US Chess Federation; FIDE, the international chess federation, adopted it in 1970. Every player carries a number, and the difference between two numbers predicts the expected result of their next game. A 400-point gap means the stronger player is expected to take roughly ten times as many points as the weaker one. After each game the rating moves by K * (actual - expected): beat the opponent you were favoured against and you gain almost nothing, upset someone far above you and you gain a lot. The K-factor sets how quickly ratings move, and when both players carry the same K, whatever one side gains the other loses.

The system has descendants. Glicko attaches an uncertainty to each rating, so the system knows how much to trust its own number. Microsoft's TrueSkill extends the idea to teams, built so Xbox Live could matchmake group games. The matchmaking rating, or MMR, behind most modern ranked ladders is some cousin of these. All of them refine the same underlying question: how strong is this player, right now, at winning the next match?

That question is worth answering because of what it feeds. A rating's day job is matchmaking. The number exists so a system can pair strangers into a fair game, over and over, and stay fair as everyone improves. Elo is infrastructure for a specific kind of game: discrete, symmetric matches between players who need to be introduced to each other. Old Light fails every one of those preconditions.

No matches for Elo to score#

Elo needs an event with a clear start and end, and a result it can score. Conflict in Old Light has none of that shape. A probe slips across a border at 3am. A frontier grinds two hexes sideways over a month of pressure that never once becomes a battle. Where does the match begin? If my fleet catches your mining system empty while you sleep, did I win a game of anything, or did you lose one you never knew you were playing?

The symmetry is missing too. A chess result carries information about skill precisely because both players start from the same position. In a persistent 4X, nobody fights fair on purpose. Picking the moment a rival's defences are thin, or the neighbour who has overextended, is the skill. An Elo update would score that result as if both sides had agreed to a fair fight.

And there is no matchmaker to feed. Everyone plays in one shared galaxy, and you fight your neighbours because they are your neighbours. The game never pairs you with anyone; geography does. In a lobby game the rating has a job the moment you queue. Here, nobody queues, so a rating would have nothing to do.

No season resets, and empires die#

Elo tracks a flow: your current form, sampled match after match, with old results diluted as new ones arrive. An empire is a stock. It accumulates. The systems you claimed in week one are still yours in month six, and the leaderboard position they support is a measure of what you hold, not of your recent form. There are no rounds and no season wipes; the galaxy just continues.

When an empire does end, it ends completely. Lose your last star and you are eliminated; the game offers a respawn with a fresh empire in a new corner of the galaxy. What number would carry across? Porting a rating over would declare that the player is the unit the galaxy measures. But the galaxy never sees players. It sees empires, and the empire that died and the one that replaces it share nothing except the person at the keyboard.

A public skill rating would leak intel#

In a matchmade game, the rating exists to make information symmetric. Fair pairing requires the system to know how strong everyone is, and publishing the rating simply shares that knowledge with the players.

In a persistent war game, information asymmetry is the gameplay. Ownership and borders are public across the whole galaxy, but what an empire has built stays hidden until you spend a probe to find out. Intelligence is something you buy with risk and travel time. Now pin a precise, honest skill rating to every name on the map. That modest-looking empire in the next sector, the one with unremarkable territory? Rated 2100. Suddenly you know to leave it alone, and you learned it for free, without launching a single recon mission. A good rating system would leak exactly what the fog exists to hide, and a more accurate rating would only leak more.

What Old Light uses instead of an Elo rating#

So the leaderboard number had to be designed backwards from what it must not reveal. Score in Old Light comes from structures. Every system rolls its structure levels up into one public number, your empire score is the sum across every system you hold, and that sum is the leaderboard. It travels with the ownership data every player already sees, so it is the one strength hint you get about any empire without scouting it.

It is deliberately coarse. A high-scoring system tells you serious investment happened there and nothing about what kind: it could be mostly extractors or mostly reactor levels, and two systems with identical scores can be wildly different opponents. Score says nothing about fleets either; a low-scoring system can have a massive one parked in orbit. Building a few systems tall and spreading many systems wide both raise it. All the number really carries is size.

A GDP figure, not a rating#

The closest real-world analogue is not a chess rating but a national GDP figure. A GDP number is public and coarse. It moves when a country builds, not when it schemes, and rivals find it useful because of what it leaves out: knowing an economy's size tells you nothing about its armies. Nations publish the figure and still fund intelligence services, because the headline number and the classified detail are different products. Old Light's score works the same way, with probes as the intelligence service.

Score is also a signal the galaxy itself responds to. Computer empires pace themselves against the strongest human empire near them: race ahead and the grey neighbours grow with you. So the number is a broadcast as much as a ladder position, and other actors, human and otherwise, act on it.

For a session game with a queue, Elo remains the right answer. Old Light never gets to ask that question, so its leaderboard counts what each empire has built and leaves everything else to your probes.