RookSac
← All articles

Why Reviewing 100 Games Beats Reviewing One

A single game review tells you what happened in that game. A hundred games tell you what happens to you. The difference matters because most of the mistakes that cost rating points are habits, and a habit only shows up when you can see it more than once.

By the RookSac teamSeptember 29, 20266 min read

One game is an anecdote

Chess results and chess accuracy are noisy. Two games against the same opponent, played an hour apart, can differ by twenty points of accuracy without your understanding of chess changing at all: you slept badly, you got a position you knew, you flagged in a won ending. If you draw conclusions from one game, you will mostly be drawing conclusions from luck.

RookSac's three bundled sample reports (100 real games each, described on the sample report page) make this concrete. The spread in single-game accuracy is large even for elite players: the standard deviation of game accuracy is MagnusCarlsen 5.4, hikaru 5.3, GothamChess 6.7, and the best and worst single games in each sample are more than 25 points apart. Now suppose you review only some of those games and take the average accuracy. This table shows the range that 90% of random picks fall into:

Games reviewedMagnusCarlsenhikaruGothamChess
1 game79.8–97.279.2–97.072.7–95.3
5 games86.2–93.985.5–92.980.1–89.6
10 games87.5–93.086.5–91.981.7–88.1
25 games88.8–91.987.8–90.783.0–86.8
50 games89.5–91.388.4–90.183.8–86.0
All 100 games90.489.384.9

Average accuracy of randomly chosen subsets of each sample (4,000 draws each); 90% of draws fall inside the range shown. Because the subsets come from the same 100 games, the ranges for the larger ones are somewhat narrower than they would be with fresh games.

Look at the first row. If you review a single game, its accuracy can land almost anywhere in a range that is 17 to 23 points wide, so it says very little about how you normally play. By 25 games the range has shrunk to three or four points, and by 50 it is about two. The lesson is not that you should analyze all your games forever. It is that the size of a difference only means something once you have enough games behind each number.

How many games do you need?

For percentages such as win rate, error rate or "how often I convert a winning position", there is a simple rule for how much noise to expect. The standard error of a rate measured over n games is at most √(0.25 / n), which is about 0.5 divided by the square root of n. Doubling that gives an approximate 95% range:

GamesTypical noise (1 standard error)Approximate 95% range around a 50% rate
10±16 points18% to 82%
20±11 points28% to 72%
50±7 points36% to 64%
100±5 points40% to 60%
400±3 points45% to 55%

So if you score 60% with the Black pieces in 20 games and 50% with White, you have learned almost nothing: both numbers are inside each other's noise. With 100 games each, a ten-point gap starts to be worth looking at. A practical rule of thumb: aim for at least 30 games in any slice you want to compare, and treat differences smaller than about ten points as noise until you have 100.

Slices worth looking at

Once you have a pool of games, split it. The split is where the insight comes from. These are the ones that tend to be most revealing:

What to look at first

Resist the urge to read every number. Take the totals in this order:

  1. Error rate. Inaccuracies, mistakes and blunders per 100 moves. It is the most direct measure of how much you are giving away and it does not depend on opponent strength as much as results do.
  2. The leakiest phase. The phase with the lowest accuracy or the highest error rate.
  3. Conversion. Winning positions that did not become wins.
  4. One opening. The opening you play most, or the one that scores worst.

That gives you a short, ranked list of things to work on, instead of a wall of statistics. Pick one and work on it for a month.

Traps in bulk analysis

A four-week plan

  1. Week 1: collect. Gather your last 50 to 100 games in one time control. Analyze them all.
  2. Week 2: find one leak. Look at phase, error rate, conversion and your main openings. Choose the single biggest problem.
  3. Week 3: study it. Review three to five games where the problem showed up (see how to review a game), then do targeted exercises.
  4. Week 4: play and re-measure. Play 30 to 50 new games, analyze them, and compare the same number. If it moved by more than the noise, it worked.

You can do all of this in a spreadsheet by hand, one game at a time. It is slow, which is why most players never do it. That is the gap RookSac is built to close: it pulls in your whole history, analyzes every game, and shows these splits on one screen. The method still matters more than the tool, though: pick a pool, split it, find one leak, fix it, re-measure.