Claude Code Found My Pirc and King's Indian Mistakes

Piotr Grudzień
Published 25 min read
A chess diagram seen from Black's side with one green arrow marking the recommended move, next to three cards reading 2,000 games downloaded, 791 as Black in the Pirc and King's Indian, and 9 positions, one PDF, under the title Claude Code Found My Pirc and King's Indian Mistakes

Intro

Setting long-running tasks for AI Agents to work on computers is the way to work in 2026. I have 15-20 Claude Code / Codex agents running for me every day. I wanted to see if they can do useful work for me in the realm of chess training.

At my chess level, you have to have some kind of an opening repertoire. Even if you spend ~0h per week actually training chess. You do find yourself often falling into trouble in similar kinds of positions. I wanted Claude to go ahead and do two things for me:

  1. Find 5-10 positions I often misplay
  2. Explain to me how to do better in those positions in a way digestible for a 2000-2200 ELO player (not grandmaster- or engine-speak)

I found both Claude’s results and process quite impressive so I’m sharing them here. Key elements that caught my attention:

  • the scoring system Claude devised for detecting and ranking teachable positions
  • how it calibrated findings to my playing strength
  • the actual 9 positions it ended up highlighting

What I did

I used Claude Code and Stockfish 19 on a two-core Linux VM to go through 2,000 of my chess.com blitz games, find the Pirc and King’s Indian positions I keep going wrong in, and turn the result into an eleven-page PDF brochure written for a 2000 to 2200 player. The session ran for four hours and thirteen minutes on its own while I was at home with the laptop closed, and I checked on it from my phone.

The nine positions and what I learned from them are in What I learned about my chess. The brochure itself is a PDF you can download.

Why the Pirc and the King’s Indian keep getting me into trouble

As Black I have played the Pirc against 1.e4 and the King’s Indian against 1.d4 for a very long time. Both give good counterattacking chances and both are on thin ice as was once explained to me by GM Mateusz Bartel. White plays simple, natural moves, I make one small inaccuracy, and by move 15 or 20 I am close to lost. It happens far less than it used to, but it still happens, and I had a feeling it happened in the same handful of positions.

Isn’t a chess engine or chess.com enough?

At my level, learning openings from engine recommendations makes no sense. You will learn what the best move is but won’t understand why and how to proceed. It’s almost as bad as writing blog posts using AI.

chess.com’s Game Review is fun but it’s aimed at beginner players and its suggestions are very basic. I wonder if the team at chess.com is working on making Game Review better calibrated to individual player’s strength?

By far the best way to train chess is with a coach or with chess books. But that takes a lot of time. So let’s use AI instead.

The same position seen from Black's side in the brochure, with a red arrow for the move I keep playing and a blue arrow for the move I should play Position 6 of the brochure, rendered from its data: the Austrian Attack after 6.Be3 Ng4 7.Bg1. I played 7…e5 in all three games that reached it. The engine wants 7…c5.

What I wanted from 2,000 games

I wanted five to ten positions: the specific moments where my habits and the engine disagree, with the move I should play and the reason in words a 2200 player uses. No ocean of engine lines, and no advice to develop my pieces and castle early. A recurring-position audit is a run that groups every game by the exact position where my results start to slip, so the positions are ranked by how often I reach them and how much they cost, not by how badly one game went.

Doing that by hand means loading hundreds of games into an analysis board, noting the move where the evaluation drops, and keeping a tally. It is weeks of evenings, which is why nobody does it.

Why run this as an unattended Claude Code session on a VM?

These days I only ever work with ten to twenty AI sessions running at the same time, so doing this project the same way was the natural choice. Each session gets its own virtual machine, which changes three things. It can run for hours without my laptop being open. It can install whatever it needs, in this case a PDF typesetter and two fonts. And it can be reached from a phone, because the terminal is a web page.

The machine for this run had two CPU cores, 4 GB of memory, Stockfish 19 and the python-chess library already installed, and a Claude Code session that starts with permission prompts switched off. The VM does not idle or sleep, which matters when the engine is going to run for three and a half hours. It also does not notify you when the work is done; if you want a ping, you ask for one, and I did.

What was the first prompt?

The first message asks for a plan, not for work. It states who I am, what I want, what the machine has, how long it may run, and what the deliverable is. It also contains the one rule that let me publish this post: my username and my opponents’ names stay out of everything it prints.

I'm a 2200+ blitz player on chess.com. As Black I play the Pirc and the King's Indian. They give good counterattacking chances but they're risky: White plays simple natural moves, I make one small inaccuracy, and by move 15-20 I'm close to lost. It happens less than it used to, but it still happens, and I want to know exactly where.

I want you to go through all my games and find the 5-10 positions where this keeps happening, then teach me those positions.

What you have here: Stockfish 19, python-chess in .venv, 2 CPU cores and 4 GB of RAM. My chess.com username is in /home/dev/chess/.chesscom_username. Read it from there and never print it. I'm going to publish screenshots of this session, so keep my username, my opponents' names and anything about the hosting of this machine out of logs, filenames, printed output and the report. Call my opponents "White".

I think I have about 1,900 blitz games. Keep the ones where I'm Black and the opening is the Pirc, the King's Indian or the Modern. Classify by the actual move sequence, not only by the ECO code chess.com attaches.

What I want at the end is a PDF brochure in results/. Page one is a one-pager: the positions at a glance. Then one page per position with a diagram from my side of the board, the line that gets there, what I usually play there and how often (from my own games), what I should play instead, and why, in the words a 2000-2200 FIDE player would use. No beginner advice. No pages of engine lines. If a move is only good for reasons Stockfish finds at depth 30, I can't learn it; prefer plan-level explanations and say when the engine's reason is not a human reason. Make it look good: this is something I want to print and keep.

You can run for several hours and I won't be watching, so budget the engine time for two cores, save intermediate results so nothing is lost if something crashes, and keep a short progress file in reports/ that I can read from my phone.

I did a planning session earlier and saved its notes in pipeline/DESIGN.md. Use them as a starting point, but it's your plan: check what holds up and change what doesn't.

Before you run anything, write me a plan: how you'll download and classify the games, how you'll find the mistakes with the engine within the time budget, how you'll decide which positions make the cut, how you'll make sure the advice is pitched at my level and not a beginner's, and how you'll check your own work. Don't start until I say go.

The browser window of the workspace: a header reading Chess workspace, tabs for Terminal, Files, Connections and Sharing, and the dark terminal below with the whole kickoff prompt just sent The desktop browser at 19:03: the workspace page with its tabs, and the terminal with the first message just sent. Click to read it at full size; the text is the block above.

“Say when the engine’s reason is not a human reason” is what separates a brochure for me from a brochure for a computer. The planning note the prompt mentions, pipeline/DESIGN.md, is a file from an earlier session: about 5,000 words of measurements and arithmetic, with the engine speed on this machine, the node budgets, the scoring ideas and the privacy rules. It opens like this:

Design note: finding my recurring Pirc / King’s Indian trouble spots. Written in a planning session on 2026-10-03. It records what was measured on this machine and the approach that came out of that planning. It is a starting point, not a spec. Check the numbers, change what does not hold up, and present your own plan before running anything.

Its decision table, which the session kept almost unchanged apart from the budgets:

DecisionChoiceWhy
Unit of engine workunique position, not gameopening positions repeat across a repertoire: dedupe for free, honest ETA
Parallelism2 worker processes, 1 thread eachsame speed as one 2-thread engine, deterministic
Screening budget150k nodes, 1 line (about 0.6 s, depth 12 to 14)every position in the window
Confirmation1M nodes, 3 lines (about 4.7 s, depth 17)flagged moves, habitual positions, drift samples
Publication20 s, 4 lines, 2 threads (depth 22 to 24)at most about 60 printed positions
Severity unitwin-probability loss, not centipawnscentipawns over-weight positions that are already decided
RendererTypst plus python-chess boardsJSON in, PDF and PNG pages out, no system libraries

The full note is published next to the brochure. Handing the next session a file is cheaper than typing the same context twice, and it is also what let the session’s plan come back in five minutes.

What did Claude Code plan?

The plan came back five and a half minutes later, in seven numbered sections and about 1,200 words. It covered the download and the classification, an engine budget in nodes rather than seconds, the selection rules, how it would keep the advice at my level, how it would check its own work, what it had changed from my notes, and a timeline. I was supposed to read it on the laptop before leaving, and I did.

The same browser window at 19:10: the terminal shows the end of Claude Code's plan, a box-drawn table of steps with wall-clock estimates, the deliverables, the assumptions to change before go, and an empty prompt box The end of the plan, captured in the desktop browser at 19:10: the timeline table, the deliverables and the assumptions it wanted confirmed.

The plan’s estimates and what actually happened:

StepPlanActual
Toolchain install and smoke render5 minunder 1 min
Download, classify, review file10 min4 min
Screening pass90 min91 min
Confirmation pass1.5 to 2 h110 min, stopped at its cap
Selection and candidate report5 min1 min
Review and writing1 habout 1 h, overlapping the confirmation pass
Publication re-check, verification, render40 min37 min
Totalabout five hours4 h 13 min from setup to the emailed brochure

I sent two small adjustments and said go. The reply confirmed both and started installing the typesetter in the background while it wrote the first scripts.

How Claude Code downloaded and classified 2,000 chess.com games

The chess.com public API serves a player’s games as monthly archives, 29 of them in my case, with the full PGN of every game. The download took 41 seconds: 2,000 games, 1,904 of them 3+2 blitz, 955 of those with me as Black.

Terminal output of the download and classification scripts: 29 months fetched, 2,000 games, then 791 games in the family split into 532 Pirc and 259 King's Indian Download and classification output, rendered from the session transcript.

Classification was by move sequence, not by the opening name chess.com attaches. A game counts as mine if Black plays both …d6 and …g6 in the first ten moves and does not push …d5 or an early …c5; it is a King’s Indian if White has both c4 and d4 by move 10, and a Pirc if White has e4 without c4. That gave 791 games: 532 Pirc and 259 King’s Indian. The classifier was cross-checked against chess.com’s labels afterwards and kept 99 percent of the games chess.com calls Pirc and 92 percent of the ones it calls King’s Indian. My Modern Defense bucket stayed empty, because I play …Nf6 early in every game.

Two facts came out of this stage before any engine ran. I score 54 percent in the Pirc and 49 percent in the King’s Indian. And a bigger gap: after move 12 my expected score is 40 percent or less in 43 percent of my King’s Indian games, against 28 percent of my Pirc games. The thin ice was mostly on the 1.d4 side. I would have guessed the opposite.

How Stockfish screened 791 games on two cores

The engine work was three passes with fixed node budgets, so that a throttled core would stretch the clock without changing a result. Before launching, the session ran a smoke test, measured about 170,000 nodes per second per core, less than my notes had assumed, and cut the budgets to fit: 100,000 nodes per position for screening and 600,000 for confirmation.

The concept diagram of the pipeline: games, classification, screening, clusters, confirmation, selection, verified write-up, Typst to PDF Concept diagram of the pipeline. Every box is a script the session wrote during the run; the cache in the middle is what makes it resumable.

The unit of work was a position, not a game. Opening positions repeat heavily within one player’s repertoire, so every position was stored once under its FEN in a small SQLite cache, keyed by the budget it was analysed at. Everything downstream (the per-game tables, the error list, the clusters, the brochure) is computed from that cache plus the raw games. Any stage can be re-run at any time, and a crash costs one position.

The screening pass covered moves 4 to 20 of every game in two phases: 9,155 unique positions out of 14,203 raw ones in the first, 11,563 in the second, with games already decided by move 12 skipped in the second. 20,718 positions in 91 minutes, two single-threaded engine workers, median depth 15, half a second each.

Terminal output while the screening pass runs: the supervisor launched, phase A finished in 42 minutes, then the progress file showing phase B at 10 percent with an ETA, positions per minute and depth The screening pass under way, rendered from the session transcript. The progress file was rewritten every 30 seconds and was what I read on the phone.

Bar chart: flagged moves per 100 games by Black's move number from 4 to 20, rising from move 7 to a peak at move 13 and staying high to move 20 Where the flagged moves are, by my move number. Chart generated from the screening and confirmation data.

The screen flagged 2,377 of my moves as losing eight or more points of expected score. The confirmation pass then re-searched the positions around every flagged move, plus my 300 most frequently reached positions, at six times the budget with three lines each: 4,536 positions at a median depth of 17, 110 minutes, stopped by its own time cap with 93 percent of the queue done. 84 percent of the flagged moves survived the deeper look.

While the engine ran, the session did something I had not asked for: it spawned a second agent to review its own analysis code. That reviewer found a precedence bug in one condition, a place where results from the two budgets were mixed in one calculation, and an off-by-one in the move window. All three were fixed before the confirmation pass started.

Checking on the run from my phone

I kicked the run off at the desk at 19:13 and closed the laptop. For the rest of the evening the session was a web page on my phone: the same terminal, with a row of keys the phone keyboard lacks.

At 19:38, still from the phone, I asked for the one thing that would tell me when it was done:

when the brochure is done, email it to me and put the one-pager as a png next to it in results. then keep going

Two phone screens side by side, each showing the mobile page with the workspace switcher, the Terminal tab, the dark terminal and a row of Esc, Tab, arrow and Keyboard keys: on the left my status question at 19:21 and the reply about the download and the classifier, on the right the email request at 19:58 and the reply folding it into the plan Two check-ins from the phone, 19:21 and 19:58, captured on the phone. The same terminal as on the desktop, with a row of keys the phone keyboard lacks.

The most useful exchange of the evening was at 21:07, while the confirmation pass was running. I typed nine words and had an answer in eight seconds:

give me a quick status and an ETA
Status: screening is complete (791 games, 2,377 flagged moves, cluster table above); the confirmation pass is 8 percent through 4,883 positions at depth 16-20, on both cores.
ETA: confirmation finishes about 20:50 UTC, then my review and writing (about an hour), then the publication-depth re-check and render, so the brochure and your email around 22:30 UTC.

Two phone screens side by side: on the left the status question being typed into the terminal at 21:07 under the end of the cluster-table message, on the right the two-line status reply with the ETA at 21:08 The question at 21:07 and the answer at 21:08, captured on the phone. The ETA was 47 minutes optimistic.

Timeline of the evening from 19:00 to 24:00: the engine stages as bars, my messages as dots below them The evening, stage by stage, drawn from the transcript timestamps. Diagram.

What I did not do from the phone is read anything long. The cluster table, the draft pages and the brochure itself waited for the laptop. A phone is good for three things here: reading a status, sending a short instruction, and seeing that nothing has stopped.

The tricks that turned 2,000 games into nine positions

This is the part of the run that decides whether the brochure is for me or for a computer. Seven rules, all of them cheap, all of them in the first prompt or the planning note in some form, and all of them implemented by the session as code.

How the score adds up

Each of the seven rules below is a number, and the diagram is the whole formula. Every flagged move gets a weight: its severity in points, capped at 40, times the recency, opponent and clock factors. The weights are summed per position, and the sum is multiplied by the visibility factor, a results factor, the teachability bonus and an exposure bonus for positions I reach often. The selection ranks those sums. The rules, one by one:

Diagram of the selection score: four cards for the per-move weight (severity, recency, opponent and clock), four cards for the per-position multipliers (visibility, results, teachability and exposure), and a footer with the selection rules How a position’s selection score adds up, as implemented in the pipeline. Diagram.

Score mistakes in expected score, not centipawns

A centipawn is the wrong unit for a mistake. Going from 0.0 to +1.5 is a different game; going from +1.5 to +3.0 is the same lost game, lost harder. Stockfish ships its own win, draw and loss model, and the session used that to turn every evaluation into an expected score for Black and every one of my moves into a loss in points of that score. A move entered the error pool at eight or more points lost; the stricter cut of ten points is what the brochure’s tile counts, 2.34 such moves per game.

Chart of White's expected score against the evaluation in pawns: Stockfish's own model is steep around zero and flat beyond plus two, the Lichess centipawn curve is drawn dashed for comparison Expected score against evaluation over 4,615 analysed positions. The two marked gaps are the point: the first pawn and a half costs 45 points, the next pawn and a half costs 5. The dashed line is the curve Lichess uses for its accuracy score. Chart generated from the eval cache.

Keep the mistakes Stockfish sees at low depth

If I could keep only one of the seven rules, it would be this one. An engine’s verdict comes with a depth attached, and the depth tells you whether a human can learn the lesson.

Level-calibrated engine analysis means filtering Stockfish’s findings by how shallow the refutation is: a mistake the engine already sees at depth 8 is a pattern a 2000 to 2200 player can learn, one it only sees at depth 25 is engine nuance.

Getting this information costs nothing. A UCI engine prints one line per completed depth while it searches, so every search left a short history behind: what the engine thought at depth 1, at depth 2, and so on up to its final depth. The session kept those histories in the cache and read two numbers out of every flagged move:

  • The visibility depth: the shallowest depth from which the position after my move already shows at least half of the final damage, and keeps showing it. A knight hanging in one move is visible at depth 1; a pawn structure that only turns bad after a six-move regrouping is visible at depth 12 or later.
  • The stability depth: the depth from which the better move stays the engine’s first choice. A move that flips between candidates until depth 16 is not a move I will find over the board.

The larger of the two is the move’s visibility, and it enters the selection score as a multiplier: 1.0 up to depth 10, 0.7 up to depth 14, 0.4 up to depth 18, and 0.15 beyond. The same number sets a category: a mistake with visibility beyond 18 and a loss under 15 points is “engine-only” and drops out of the running; a stable, quiet better move is a “plan”; a forcing refutation is a “tactic”.

Histogram of the depth at which Stockfish first sees the damage of a flagged move, with bands labelled human, hard, deep and engine-only 48 percent of my flagged moves are visible by depth 10. Only 33 of 1,937 needed depth 19 or more, and those did not make the brochure. Chart generated from the confirmation data.

The chart is what the rule looks like over the 1,937 confirmed moves: 48 percent are visible by depth 10, the band a club player can see at the board, and only 33 needed depth 19 or more. None of the nine positions comes from that tail. Three of the nine are visible at depth 1 (positions 4, 8 and 9: the engine sees the problem the moment the move is played), and the deepest printed one is position 5, at depth 11, where the damage only shows after 6.e5 Nfd7 7.Qf3. Every page of the brochure states both depths in its footer, so I know whether a rule is a pattern to memorise or a line to work through.

The rule does not favour cheap mistakes over costly ones. Severity is a separate factor and the two multiply. A shallow mistake with a small loss scores low. A deep mistake with a big loss scores lower than it would in any other engine ranking. What rises to the top is expensive and learnable at the same time.

Compare what I usually play with what the engine plays

For every position I had reached three or more times, the session kept my move distribution with results, and the engine’s top three moves with evaluations. A position where my usual move is six or more points below the engine’s choice is a repertoire leak even if no single game crossed the blunder threshold. Eighteen of those came out of 791 games, and three of the nine pages are leaks rather than blunders.

Paired horizontal bars for the nine positions: my expected score after my usual move in rust and after the better move in teal My expected score after my usual move and after the better one, nine positions, at the publication depth. Chart generated from the data behind the brochure.

Separate single blunders from slow drift

The “thin ice” feeling has two causes with different cures: one bad move, or a whole setup that drifts from equal to worse without any single error. The session measured drift as the change in expected score between move 4 and move 12 in games where no single move lost ten points, grouped by White’s system. It found 17 drift games in two groups, not enough for a page of their own. My problem turned out to be moves, not setups, which was good news.

Discount flagged clocks and pre-moves

Every chess.com PGN carries the clock after each move. A move played with under fifteen seconds left counted a quarter; a pre-move in a position I had not reached before, or a probable mouse-slip, counted half; a pre-move in a position I reach every week counted in full, because that is exactly the autopilot habit I wanted to find. 348 of the 2,107 moves in the error pool came from games I lost on time, and a loss on time never counted as losing because of the opening.

Weight losses to lower-rated opponents

A mistake punished by a 2400 counts 1.4 times, one against an 1800 counts 0.6, because the stronger the opponent, the more standard the setup that caused it. The brochure’s “punished” count on each page comes from the same data: for position 5, opponents found the refutation in five of five games.

Prefer plan moves over long tactics

The last rule separates two kinds of better move. A plan move is stable from depth 8 and quiet: a pawn break, a reroute, an exchange, a castle. A tactic is a forcing sequence that wins material within a few moves. Both can make a page, but a plan move whose alternative loses by force is the most teachable kind of mistake, and it got a 30 percent bonus in the selection score. Of 2,107 flagged moves, 155 were plan errors, 712 tactics and 1,238 a mix.

The selection then ran a diversity rule across 154 candidates: no more than two positions per White system, both openings represented, and a stop when the next candidate scored less than 35 percent of the first. One exclusion is worth quoting from the cluster table the session printed at 20:58:

Terminal output of the session's reading of the cluster table before confirmation: 4...a6 against Be3 is an engine preference, the King's Indian Classical with 6...Nbd7 is the biggest theme, and two single-position tactical blunders repeat The session’s early reading of the cluster table, rendered from the transcript. 4…a6 against Be3 appears in 50 games and the engine dislikes it, but I score 54 percent from there, so it got no page.

The top rows of that table, at screening depth:

#SystemMy moveReachedEngineGames with the errorLoss rateVisible from depth
1Pirc, 150 Attack4…a650Bg75040%12
2King’s Indian, Sämisch7…Nbd713e61155%9
3Pirc, Nf3 and h39…Nxe42Bg72100%1
5Pirc, Be3 without Qd25…b57Bg7743%11
6King’s Indian, Classical9…c68exd4838%5
9King’s Indian, Fianchetto7…Qc713Bf51338%1
11Pirc, Austrian7…e53c5333%5

How Claude Code checked its own work

Nothing printed in the brochure comes from the screening or confirmation passes. Every position on a page, every drill, and the end of every printed line was re-searched for 20 seconds with four lines on both cores, 70 positions at depths between 18 and 26, typically 22. The session then ran a verifier over the draft: the recommended move must be the engine’s first choice at that depth or within 15 centipawns of it; every line must replay legally; every diagram must equal the position reached in a real game; every claim such as “equal” or “White is winning” must match the evaluation; and no evaluation may be typed by hand, only filled in from the fact table. If any check fails, the renderer refuses to build the PDF.

Terminal output of the verifier: a result line with five failures and the list of failed checks, then a patched run with one failure, the renderer refusing to render, then all checks passing and the privacy scan clean The verifier at work at 23:14 and 23:22, rendered from the transcript.

It failed five times on the first draft. The most interesting failure was position 5: at depth 22 the engine’s first choice changed from the move the draft recommended to another, and the page had to say so. The brochure now prints both, with their expected scores, and names the engine’s choice as the engine’s choice.

The verifier checks moves, evaluations and legality. It cannot judge prose. Before publishing I went through the brochure once more with a second session, which found one wrong sentence on position 9: the rule line said a tactic “wins the queen for a rook” when it wins a piece, and the rook only after a recapture. The sentence was fixed and the page re-verified. The public copy of the PDF below also has the links to my games removed, since each one shows both players’ names.

What the brochure looks like

The brochure is an A4 PDF: a one-pager with the nine positions at a glance, one page per position, and a closing page with a table of my results by White’s system and the method. It is typeset with Typst from a JSON fact table, and the diagrams come from python-chess. The session picked Typst because it reads JSON, embeds vector boards and needs no browser or system libraries. The draft went through eight renders, from 22 pages down to 11, before the layout held.

The top of the brochure's first page: the title Where the Pirc and the King's Indian go wrong, the subtitle 791 games as Black, and four tiles with 791 games analysed, 53 percent score as Black, 33 percent worse by move 12, 2.34 flagged moves per game The top of page 1. Page 1 of the PDF.

The first row of the one-pager's grid: three mini boards from Black's side with coloured arrows, each with an opening badge, the White system and a one-line rule The first three of the nine positions on page 1: red for what I usually play, blue for what to play, yellow for White’s plan.

The whole first page of the brochure: title, tiles and a three by three grid of boards with one-line rules Page 1 in full. Click to enlarge.

Two position pages side by side: each has a large board, the line to reach it, a bar of what I played, sections for what I usually play and what to play instead, the reason, a pattern box, a lesson bar and a drill board Pages 7 and 2 of the PDF, the Austrian Attack after 7.Bg1 and the Classical King’s Indian with …Nbd7.

The top half of a position page at reading size: the board, the move order, the usual move with its explanation and line, and the better move with its line The top of page 10 at reading size. Every number in the footer comes from the 20-second re-search.

At 23:27 the session emailed me the PDF and the one-pager, the one notification of the whole run, because I had asked for it at 19:38.

A phone lock screen showing Saturday 3 October, 23:29, and a mail notification announcing the brochure email with the PDF and the one-pager attached 23:29: the email arriving on the phone, captured on the phone.

The browser window on the Files tab of the workspace: a results folder listing the one-pager image, the brochure PDF and the PGN with Download buttons, and below it a Recent emails entry marked Accepted for delivery, the address blurred The same files in the workspace’s Files tab, captured in the desktop browser at 23:40, with the email listed as accepted for delivery. The PDF is downloaded from here.

The session's closing message listing the deliverables and the nine positions in one line each The closing message, rendered from the transcript: the nine positions as the session summarised them.

The complete PDF is at pirc-kid-key-positions.pdf.

What I learned about my chess

Five King’s Indian positions and four Pirc positions made the cut. The ones below are the pages of the brochure I keep coming back to; all nine are in the table at the end of this section. The numbers are my expected score as Black after my usual move and after the better one, from the 20-second search.

Position 9: with the queen on e7, e4 is poisoned

Board from Black's side after 9.Be2 in the 4.Be3 a6 5.a4 Pirc with ...Nc6 and ...e5, a red arrow for ...Nxe4, a blue arrow for ...Bg7 and a yellow arrow for Nxc6, next to the line and the numbers 1.e4 d6 2.d4 Nf6 3.Nc3 g6 4.Be3 a6 5.a4 Nc6 6.h3 e5 7.Nf3 exd4 8.Nxd4 Qe7 9.Be2. Diagram from the data behind the brochure.

I played 9…Nxe4 twice here and lost both games. The capture looks safe because the queen on e7 seems to cover the knight, but 10.Nxe4 Qxe4 11.Bf3 attacks the queen while c6 hangs: 11…Qh4 12.Nxc6 wins a piece, and 12…bxc6 runs into Bxc6+ and Bxa8. The engine sees it at depth 1. 9…Bg7 and castling is simply fine. Expected score: 0 percent after …Nxe4, 49 percent after …Bg7.

I must have seen this before, but it is a trap I should be recognising immediately: queen on the e-file, knight on d4 against my knight on c6, bishop able to reach f3. The pattern is the one from the Scotch, and I had never filed it under the Pirc.

Position 6: after 7.Bg1, hit d4, not e5

The Austrian Attack with 4.f4 Bg7 5.Nf3 O-O 6.Be3 Ng4 7.Bg1 (the first image of this post). 7…e5 is the thematic Pirc blow and I played it in all three games that reached the position, but 8.dxe5 dxe5 9.h3 Nf6 10.Qxd8 Rxd8 11.Nxe5 just wins the pawn. 7…c5 hits d4 instead, and after 8.h3 cxd4 9.Bxd4 Nf6 Black is equal. Expected score: 1 percent against 49.

I may have seen it, but I would not remember it, and during a game I would have to think. The one-line rule the brochure prints is “7.Bg1 c5”, which is all I need.

Position 1: take on d4 before …c6

Board from Black's side after 9.Bf1 in the Classical King's Indian with ...Nbd7, a red arrow for ...c6, a blue arrow for ...exd4 and a yellow arrow for d5, next to the line and the numbers 1.c4 Nf6 2.Nc3 g6 3.e4 d6 4.d4 Bg7 5.Nf3 O-O 6.Be2 Nbd7 7.O-O e5 8.Re1 Re8 9.Bf1. Diagram from the data behind the brochure.

Eight games reached this position and I played 9…c6 in all eight, scoring 44 percent. 9…c6 looks like natural preparation, but it allows 10.d5, and after 10…c5 11.h3 White has the position he wants. 9…exd4 10.Nxd4 Ne5 first, and …c6 only afterwards, keeps the centre open and the knight on e5 active. Expected score: 26 percent against 46.

I genuinely thought it was better to wait with …exd4 in these positions. This is the page that changed a belief rather than reminding me of something.

Position 4: 6…c5, every time

Board from Black's side after 6.Nf3 in the Four Pawns Attack, a red arrow for ...Nbd7, a blue arrow for ...c5 and a yellow arrow for the e4 to e5 push, next to the line and the numbers 1.d4 Nf6 2.c4 g6 3.Nc3 d6 4.e4 Bg7 5.f4 O-O 6.Nf3. Diagram from the data behind the brochure.

The Four Pawns is my worst system by result. In this position I have played three different moves; 6…Nbd7 in half the games, and after 7.e5 Ne8 8.c5 White rolls. 6…c5 7.d5 e6 is the only move I should ever play here. Expected score: 22 percent after the knight move, 48 after the break. I knew this at some point and definitely forgot it.

Position 8: after 10.Nd5, 10…Nd7 and let e7 go

Board from Black's side after 10.Nd5 in the Sämisch Gambit, a red arrow for ...Kf8, a blue arrow for ...Nd7 and a yellow arrow for Nxe7, next to the line and the numbers 1.d4 Nf6 2.c4 g6 3.Nc3 d6 4.e4 Bg7 5.f3 O-O 6.Be3 c5 7.dxc5 dxc5 8.Qxd8 Rxd8 9.Bxc5 Nc6 10.Nd5. Diagram from the data behind the brochure.

I reach the Sämisch Gambit regularly because I always answer 6.Be3 with 6…c5, and in this position I have played four different moves in six games. The book move is 10…Nd7: 11.Nxe7+ Nxe7 12.Bxe7 Bxb2 13.Rd1 Re8 and the bishop pair and the b2 pawn pay for the e7 pawn. 10…Kf8, my most frequent choice, keeps the pawn and gives White the better game. Expected score: 41 percent against 50. The engine sees the difference at depth 1, which is a polite way of saying I should have learned the line.

Position 5: against 5.h3, develop before …b5

Board from Black's side after 5.h3 in the 4.Be3 a6 Pirc, a red arrow for ...b5, a blue arrow for ...Bg7 and a yellow arrow for e5, next to the line and the numbers 1.e4 d6 2.d4 Nf6 3.Nc3 g6 4.Be3 a6 5.h3. Diagram from the data behind the brochure.

I play 4…a6 against Be3 in every game, and after 5.Qd2 the follow-up …b5 is correct. After 5.h3 I played the same …b5 in seven of seven games, and in five of them White found 6.e5 Nfd7 7.Qf3, with the queen still on d1 a move earlier, and I was worse by move 12 in most of them. 5…Bg7, and …b5 only after castling, is the difference between an autopilot move and the same move played at the right time. The engine’s first choice is 5…Nbd7; both are better than what I do.

All nine positions

#Opening and systemAfterMy usual moveBetterPoints lostPage
1King’s Indian, Classical with …Nbd79.Bf19…c6 (8 of 8)9…exd4192
2King’s Indian, Sämisch Benoni7.d57…Nbd7 (11 of 13)7…e653
3King’s Indian, Fianchetto with …c67.Nc37…Qc7 (13 of 13)7…Bf5 or 7…d5114
4King’s Indian, Four Pawns6.Nf36…Nbd7 (3 of 6)6…c5275
5Pirc, 4.Be3 a6 5.h35.h35…b5 (7 of 7)5…Bg7 or 5…Nbd7276
6Pirc, Austrian with 6.Be3 Ng47.Bg17…e5 (3 of 3)7…c5487
7Pirc, Austrian, the e5 wedge against …Nc613.O-O13…Na5 (pattern over 7 games)13…f6458
8King’s Indian, Sämisch Gambit10.Nd510…Kf8 (2 of 6)10…Nd799
9Pirc, 4.Be3 a6 5.a4 with …Nc6 and …e59.Be29…Nxe4 (2 of 2)9…Bg74910

Points lost are points of expected score at the publication depth. Position 7 is a pattern rather than a single position: seven games reached the same structure, White pawns on d4 and e5 against my knight on c6, by different move orders, and in every one of them I moved a knight instead of playing …f6.

What I learned about working this way

The run was 181 tool calls, 20 scripts, one reviewer agent, eight background monitors and three detached processes, and I typed seven messages. The things that made it work are not chess-specific.

Plan first, then go. The first message asks for a plan and forbids starting. Reading the plan took me three minutes and it is where I would have caught a wrong assumption, had there been one. The plan’s time estimates were within ten minutes of reality for every stage except its own writing.

Write the level into the prompt. “No beginner advice” is a wish. “If a move is only good for reasons Stockfish finds at depth 30, I can’t learn it” is a specification, and it turned into a number the session recorded for every flagged move.

Give it a machine and a deliverable path. A VM that stays on, a results folder, and a sentence saying what file I want at the end. Everything in between is the session’s problem, and it solved the parts I had not thought about, such as caching every position so a crash would cost one search rather than an hour.

Ask it to check itself at a higher budget than it worked at. Screening is cheap and noisy. The verifier that refuses to render the PDF until the deep search agrees with every printed claim is the most valuable part of the pipeline. It caught a real change of mind at depth 22.

The review round is where my expertise goes in. I had nothing to change when I read the brochure, which is a credit to the first prompt, not to me. The one wrong sentence was found by a second pass over the prose, which the verifier cannot judge.

Hours of machine time, minutes of attention. Four hours and thirteen minutes of VM time, seven messages, and the only interruptions were the ones I chose.

What I would change next time: a shorter confirmation cap, because the pass stopped at 93 percent and the last 7 percent would have taken twenty more minutes; the “worse by move 12” statistic computed before the engine runs, because it is the most useful number in the brochure and needs no engine at all; and the publication pass before the writing, so that the text is written against final numbers from the start.

How to reproduce this with your own games

You need a Linux machine that stays on, Claude Code, Stockfish, python-chess, and a PDF typesetter such as Typst. The kickoff prompt above is the whole specification; change the openings, the rating and the number of games. The rules that matter, in the order the session applied them:

  1. Download the monthly archives from the chess.com public API and keep the raw JSON; it holds the username and the opponents’ names, so it never leaves the machine.
  2. Classify games by the moves actually played, and cross-check against the site’s opening labels afterwards.
  3. Analyse positions, not games, with fixed node budgets, and cache every result under its FEN and budget.
  4. Score every move in expected score using the engine’s own win, draw and loss model.
  5. Record the depth at which the engine first sees the damage, and prefer shallow.
  6. Group errors by exact position, by opening line and by structure, and keep your own move distribution per position.
  7. Re-search everything you intend to print at a much higher budget, and let a verifier block the render on any disagreement.
  8. Tell the session what must never be printed, and check the output for it before you publish.

About this series

At Quickchat AI we work with ten to twenty AI sessions running at the same time, each on its own machine, each with one task and one deliverable. This series shows how, one specific task per post, with the prompts and outputs as they were. The next posts will cover work that is not chess.

Frequently asked questions

Can Claude Code analyze my chess games with Stockfish?

Yes. Claude Code ran on a Linux VM with Stockfish 19 and python-chess installed, downloaded my 2,000 chess.com blitz games through the public API, and wrote and ran the analysis scripts itself: 20,718 unique positions at a screening budget, 4,536 at a confirmation budget, and 70 at a 20-second publication budget. The engine does the evaluating; Claude Code does the bookkeeping, the grouping and the write-up.

How do I find recurring mistakes across all my games instead of reviewing one game at a time?

Group the games by the exact position they reach, not by game. Every position is stored once under its FEN, every one of my moves gets a loss in points of expected score, and the losses are grouped by position, by opening line and by structure. In my 791 games as Black, 68 exact positions carried a mistake in two or more games, and 18 habitual moves were measurably below the engine’s choice even where no single game looked like a blunder. A per-game review never shows that, because each game is scored on its own.

Is the Pirc Defense too risky for Black at club level?

In my 532 Pirc games at around 2200 blitz I scored 54 percent, and the trouble sits in specific positions rather than in the opening as a whole: the Austrian Attack with 6.Be3 Ng4 7.Bg1, the 4.Be3 a6 lines where 5.h3 or 5.a4 punish an early …b5 or a careless …Nxe4, and the e5 wedge against …Nc6. The King’s Indian was the bigger problem in my games: 49 percent, and worse by move 12 in 43 percent of games against 28 percent in the Pirc.

Which King’s Indian and Pirc lines give Black the most trouble?

In my games: the Classical King’s Indian with …Nbd7, where 9…c6 lets 10.d5 in and …exd4 first is better; the Sämisch Benoni, where 7…Nbd7 should be 7…e6; the Fianchetto with …c6, where 7…Qc7 gives up e4 and …Bf5 or …d5 is right; the Four Pawns Attack, where 6…c5 is the only move I should ever play; the Sämisch Gambit, where 10.Nd5 should be met by 10…Nd7; and in the Pirc the Austrian Attack after 7.Bg1 (…c5, not …e5), the e5 wedge against …Nc6 (…f6 at once), the 4.Be3 a6 5.h3 line (develop before …b5) and the 4.Be3 a6 5.a4 line with …Nc6 and …e5, where 9…Nxe4 loses a piece.

Can Claude Code run overnight on a VM, and how do you check on it from your phone?

Yes. My run took 4 hours and 13 minutes from setup to the emailed brochure, unattended, on a two-core VM that stays on when the laptop is closed. I opened the same terminal in the phone’s browser and typed one-line questions; a status with an ETA came back in eight seconds. The session also kept a short progress file I could read from the phone, and it emailed me the PDF when it was done because I had asked it to. It does not notify you on its own.

Why not just use chess.com Game Review or Lichess Insights?

Game Review scores one game at a time and Lichess Insights gives averages per opening family. Neither tells you which exact position you keep reaching and losing from, how often you play the wrong move there, or what a 2000 to 2200 player should play instead. The pipeline in this post answers those three questions for nine positions and puts the answers on one page each.

What does level-calibrated engine analysis mean?

Level-calibrated engine analysis means filtering Stockfish’s findings by how shallow the refutation is: a mistake the engine already sees at depth 8 is a pattern a 2000 to 2200 player can learn, one it only sees at depth 25 is engine nuance. The run recorded, for every flagged move, the depth at which Stockfish first saw the damage, kept the shallow ones and dropped the deep ones. 48 percent of my flagged moves were visible by depth 10.