Why Reviewing 100 Games Beats Reviewing One
A single game review tells you what happened in that game. A hundred games tell you what happens to you. The difference matters because most of the mistakes that cost rating points are habits, and a habit only shows up when you can see it more than once.
One game is an anecdote
Chess results and chess accuracy are noisy. Two games against the same opponent, played an hour apart, can differ by twenty points of accuracy without your understanding of chess changing at all: you slept badly, you got a position you knew, you flagged in a won ending. If you draw conclusions from one game, you will mostly be drawing conclusions from luck.
RookSac's three bundled sample reports (100 real games each, described on the sample report page) make this concrete. The spread in single-game accuracy is large even for elite players: the standard deviation of game accuracy is MagnusCarlsen 5.4, hikaru 5.3, GothamChess 6.7, and the best and worst single games in each sample are more than 25 points apart. Now suppose you review only some of those games and take the average accuracy. This table shows the range that 90% of random picks fall into:
| Games reviewed | MagnusCarlsen | hikaru | GothamChess |
|---|---|---|---|
| 1 game | 79.8–97.2 | 79.2–97.0 | 72.7–95.3 |
| 5 games | 86.2–93.9 | 85.5–92.9 | 80.1–89.6 |
| 10 games | 87.5–93.0 | 86.5–91.9 | 81.7–88.1 |
| 25 games | 88.8–91.9 | 87.8–90.7 | 83.0–86.8 |
| 50 games | 89.5–91.3 | 88.4–90.1 | 83.8–86.0 |
| All 100 games | 90.4 | 89.3 | 84.9 |
Average accuracy of randomly chosen subsets of each sample (4,000 draws each); 90% of draws fall inside the range shown. Because the subsets come from the same 100 games, the ranges for the larger ones are somewhat narrower than they would be with fresh games.
Look at the first row. If you review a single game, its accuracy can land almost anywhere in a range that is 17 to 23 points wide, so it says very little about how you normally play. By 25 games the range has shrunk to three or four points, and by 50 it is about two. The lesson is not that you should analyze all your games forever. It is that the size of a difference only means something once you have enough games behind each number.
How many games do you need?
For percentages such as win rate, error rate or "how often I convert a winning position", there is a simple rule for how much noise to expect. The standard error of a rate measured over n games is at most √(0.25 / n), which is about 0.5 divided by the square root of n. Doubling that gives an approximate 95% range:
| Games | Typical noise (1 standard error) | Approximate 95% range around a 50% rate |
|---|---|---|
| 10 | ±16 points | 18% to 82% |
| 20 | ±11 points | 28% to 72% |
| 50 | ±7 points | 36% to 64% |
| 100 | ±5 points | 40% to 60% |
| 400 | ±3 points | 45% to 55% |
So if you score 60% with the Black pieces in 20 games and 50% with White, you have learned almost nothing: both numbers are inside each other's noise. With 100 games each, a ten-point gap starts to be worth looking at. A practical rule of thumb: aim for at least 30 games in any slice you want to compare, and treat differences smaller than about ten points as noise until you have 100.
Slices worth looking at
Once you have a pool of games, split it. The split is where the insight comes from. These are the ones that tend to be most revealing:
- By phase. Is your accuracy lowest in the opening, the middlegame or the endgame? That tells you which part of chess to study first. In all three RookSac samples, even at 2,900–3,400 rating, the middlegame is the least accurate phase.
- By colour. A large gap between your White and Black results usually points to your repertoire for one colour, not to your general strength.
- By opening. Which opening do you play most, and how does it score? An opening you play in a fifth of your games and score badly with is worth far more study than a rare line you lose in.
- By time control. Blitz and rapid are different skills. Mixing them hides both.
- By clock. How often do your errors happen with under a minute left? In the samples, moves played with 30 seconds or less on a 3-minute clock made up about 15% of moves but about a quarter of the mistakes and blunders. If your own number is higher, the fix is time management, not tactics. See time management in blitz and rapid.
- By opponent rating. Do you play worse against stronger opponents, or against weaker ones? Both are common, for different reasons.
- Converted or not. In how many games did you reach a clearly winning position, and how many of those did you win? Every one that got away is a specific moment you can go back to.
What to look at first
Resist the urge to read every number. Take the totals in this order:
- Error rate. Inaccuracies, mistakes and blunders per 100 moves. It is the most direct measure of how much you are giving away and it does not depend on opponent strength as much as results do.
- The leakiest phase. The phase with the lowest accuracy or the highest error rate.
- Conversion. Winning positions that did not become wins.
- One opening. The opening you play most, or the one that scores worst.
That gives you a short, ranked list of things to work on, instead of a wall of statistics. Pick one and work on it for a month.
Traps in bulk analysis
- Mixing conditions. A pool that contains 1-minute bullet and 30-minute classical will average out to something that describes neither. Filter first.
- Reading noise as a trend. Your accuracy will bounce a few points from week to week. Judge change over dozens of games, not five.
- Cherry-picking. If you only analyze games you lost, every statistic will be worse than reality. That is fine for finding problems, but do not compare it with an all-games number.
- Chasing tiny differences. "My accuracy is 0.8 lower with the Sicilian" is almost never real. Look for gaps large enough to survive the noise table above.
- Forgetting the opponent. Compare your accuracy with your opponents' in the same games. If both of you play at 78, the game was probably messy for both sides, and there may be less to learn from it.
A four-week plan
- Week 1: collect. Gather your last 50 to 100 games in one time control. Analyze them all.
- Week 2: find one leak. Look at phase, error rate, conversion and your main openings. Choose the single biggest problem.
- Week 3: study it. Review three to five games where the problem showed up (see how to review a game), then do targeted exercises.
- Week 4: play and re-measure. Play 30 to 50 new games, analyze them, and compare the same number. If it moved by more than the noise, it worked.
You can do all of this in a spreadsheet by hand, one game at a time. It is slow, which is why most players never do it. That is the gap RookSac is built to close: it pulls in your whole history, analyzes every game, and shows these splits on one screen. The method still matters more than the tool, though: pick a pool, split it, find one leak, fix it, re-measure.