# Design note: finding my recurring Pirc / King's Indian trouble spots Written in a planning session on 2026-10-03. It records what was measured on this machine and the approach that came out of that planning. **It is a starting point, not a spec.** Check the numbers, change what does not hold up, and present your own plan before running anything. ## 0. Privacy rules for this run (non-negotiable) - The chess.com username is in `/home/dev/chess/.chesscom_username`. Read it from there (strip whitespace; lowercase it for API paths). Never print it: not in logs, progress files, filenames, commit messages, PGN headers you write, the report or the brochure. Refer to the player as "you" or "Black", and to every opponent as "White". - Never log a full API URL (it contains the username). Log the month (`2026/09`) instead. - Opponent names sit in every PGN header. Strip them from anything you write out. Example-game references use the date, both ratings, the result, the move count and the game link only. - Say nothing about how or where this machine is hosted in any output. This session will be screenshotted and published. ## 1. What is on this machine (measured) - **Stockfish 19** at `/usr/local/bin/stockfish` (x86-64-bmi2 code path chosen automatically, NNUE embedded). About 480 knps with Threads=2, Hash=256: depth 20 in 2 s from the start position. With Threads=1, Hash=128 on opening positions: 150k nodes ≈ 0.5–0.7 s (depth 12–14), 1M nodes ≈ 4 s (depth 17–18). Two single-threaded engines give the same throughput as one two-threaded engine, and node limits make results reproducible. - **python-chess 1.11.2** in `/home/dev/chess/.venv` (no requests/httpx/matplotlib/PIL/numpy: use `urllib` and hand-written SVG). `chess.svg.board()` renders boards; coordinates are drawn as paths, so no font is needed. Arrow colours: use the NAMED colours (`"red"`, `"green"`, `"blue"`, `"yellow"`) and override them via `colors={"arrow red": "#rrggbbaa"}`; a literal hex colour passed to `Arrow(color=...)` is emitted as 8-digit hex and breaks some renderers. - `engine.analysis(board, chess.engine.Limit(nodes=N), multipv=k, game=)` streams one `info` per completed depth (a free per-depth history). Passing a new `game` object triggers `ucinewgame` (fresh hash, deterministic node-limited results). `root_moves=[move]` evaluates one specific move. **python-chess applies no timeout to node-limited searches**: a hung engine hangs its worker forever, so run a heartbeat watchdog. Never put MultiPV in `engine.configure()`; pass `multipv=` to `analyse`/`analysis`. - `GameNode.clock()` parses chess.com `[%clk 0:02:53.7]` (tenths). `board.epd()` drops the move counters, which makes it a good cache key. - **`chess-analyze`** (CLI) analyses one position per call with about 1.3 s of engine start-up each time: fine for spot checks, wrong for batch work. Its Python source (path in the wrapper: `cat "$(command -v chess-analyze)"`) exports `serialize_score(score, perspective)`; reuse it so scores keep the same JSON schema (`{"type":"centipawns","value":35}` or `{"type":"mate","value":N,"winner":...}`, mates never converted to centipawns). - 2 CPU cores, 3.8 GB RAM, no swap, about 74 GB free disk, passwordless sudo, internet. - Not installed: node, chromium, LaTeX, weasyprint, pandoc, fontconfig, imagemagick, rsvg. Installed fonts: DejaVu only. - **chess.com public API** (no auth; send a User-Agent header): `https://api.chess.com/pub/player//games/archives` lists 29 monthly archive URLs (2024/06 → 2026/10). Each month holds about 50–70 games, all 180+2 blitz in the sample. Game JSON keys: `white`/`black` (`rating`, `result`, `username`), `eco` (an opening URL), `end_time`, `pgn`, `time_class`, `time_control`, `rules`, `rated`, `uuid`, `url`, sometimes `accuracies`. PGNs carry `[ECO]`, `[ECOUrl]` and `%clk`, no `%eval`. - Opening names: `https://raw.githubusercontent.com/lichess-org/chess-openings/master/{a,b,c,d,e}.tsv` (CC0; columns `eco`, `name`, `pgn`). Replay each `pgn` to an EPD and look up the deepest hit along a game's path for transposition-aware names. No opening-explorer API is reachable from here (401), so any "what strong players do" framing comes from your own chess knowledge and must be labelled as such in the method notes. - Earlier one-game review scripts live in `/home/dev/chess/scratchpad/` (`game_pass2.py`: one MultiPV-3 search per position, with the eval after a move taken as the next position's eval; `build_review.py`: a hand-written SVG eval chart with an explicit palette, an annotated PGN, and an `EXPECT_BEST` assertion that aborts if the narrative's "best moves" disagree with the engine). Reuse the ideas. - Sample month (2026/09): 24 games as Black, 21 of them in the Pirc / King's Indian / Modern family; every reply to 1.e4 was 1...d6. Expect about 1,900 blitz games in total, about 950 as Black, about 800 in the family. ## 2. Design in one paragraph Engine stages never "analyze games". They fill an **EPD-keyed SQLite eval cache** (`(epd, budget_id)` → JSON score + per-depth history). Everything downstream (per-game move tables, errors, clusters, drift, habits, stats, the brochure) is a pure, idempotent function of that cache plus the raw game JSON, so any stage can be re-run at any time, partial results exist while the engine is still running, and a crash loses at most one position. Two single-threaded Stockfish workers run node-limited searches (deterministic); a deeper "publish" pass re-checks everything that ends up in print. The brochure renders from one fact table (`brochure.json`); narrative text may reference evaluations only through placeholders that a verifier resolves and asserts, so no hand-typed number can reach the PDF. | Decision | Choice | Why | |---|---|---| | Unit of engine work | unique EPD (not game) | opening positions repeat heavily across a repertoire: dedupe for free, honest ETA | | Parallelism | 2 worker processes × Threads=1, Hash=128 | same nps as one 2-thread engine, deterministic | | Screening budget `screen.v1.n150k.mpv1` | 150k nodes, MultiPV 1 (~0.6 s, depth 12–14) | every window position | | Confirmation `confirm.v1.n1000k.mpv3` | 1M nodes, MultiPV 3 (~4.7 s, depth ~17) | flagged moves, habitual positions, drift samples | | Publication `publish.v1.t20.mpv4.th2` | 20 s, MultiPV 4, Threads 2, Hash 512 (depth 22–24) | at most ~60 printed positions | | Severity unit | win-probability loss: `wp(cp) = 50 + 50 * (2 / (1 + exp(-0.00368208 * cp)) - 1)` (cp clamped to ±1000, mates → ±1000) | centipawns over-weight already-decided positions | | Renderer | Typst (static binary) + python-chess SVG boards | JSON in, PDF and PNG pages out, no system libraries | ## 3. Layout ``` /home/dev/chess/pipeline/ setup.sh, common.py, progress.py, features.py, openings.py, s1_download.py s2_classify.py s3_positions.py s4_engine.py s5_assemble.py s6_aggregate.py s7_select.py s9_verify.py s10_render.py run_all.sh, brochure/{brochure.typ,theme.typ}, content/positions.json (your narrative, stage 8) /home/dev/chess/pipeline/data/ raw/*.json, games.jsonl, classified.jsonl, queue.jsonl, evals.sqlite (WAL), games_analyzed/.json, errors.jsonl, clusters.json, drift.json, habits.json, stats.json, selected.json, brochure.json, brochure/ (build dir), logs/, state/ /home/dev/chess/.tools/typst/typst static binary /home/dev/chess/reports/ progress.md (short, phone-readable), run-log.md (a dated paragraph per stage), flagged-games-review.md, pirc-kid-candidates.md (the first look), pirc-kid-report.md, verify-report.md, figures/*.svg, brochure-pages/page-NN.png /home/dev/chess/results/ pirc-kid-key-positions.pdf, pirc-kid-key-positions.pgn, one-pager.png ``` Conventions: position k = the position after k half-moves (k = 0 is the start); Black is to move at odd k; Black's move m is played from position 2m−1. The window is **positions 8–40** (baseline at 8, Black's moves 5–20 and the replies). All JSON writes are atomic (`tmp` + `os.replace`). Every script takes `--help` and `--force`, and logs `ts | stage | level | msg` to `data/logs/.log` and stdout. Raw data stays under `pipeline/data/` (it contains names); `results/` and `reports/` hold only cleaned deliverables. ## 4. Stages 1. **`s1_download.py`**: archive list, then each month to `data/raw/YYYY-MM.json` (skip existing files except the current and previous month; 1 s between requests; 3 retries with 2/8/30 s backoff on 429/5xx). Emit `data/games.jsonl`: `{uuid, url, end_time_iso, month, time_class, time_control, rated, rules, user_color, user_rating, opp_rating, user_result (win|draw|loss), result_code, lost_on_time, eco, eco_url, accuracies|null, pgn}`. Draw codes: agreed, repetition, stalemate, insufficient, 50move, timevsinsufficient. About 2 minutes. 2. **`s2_classify.py`**: games as Black, `rules == chess`, `time_class == blitz`. Rules (tested at 88 % recall on a sample month): Black's first move ∈ {…d6, …g6, …Nf6}; both …g6 and …d6 pushed by Black's 12th move; exclude a …d5 *push* (not a capture onto d5) within 12 moves and …c5 within Black's first 5 moves (a later …c5 is a thematic break and stays). Family: `kid` if White has played both c4 and d4 by move 10 (covers 1.e4 d6 2.d4 Nf6 3.f3 g6 4.c4); `pirc` if e4 without c4 and …Nf6 by Black's 6th move; `modern` if no …Nf6 by move 6; else `other`. Flags: `c4-no-d4`, `d4-no-c4-no-e4` (London/Torre/150-style), `eco-mismatch` (chess.com's ECOUrl says Pirc/Modern/Robatsch/Kings-Indian but the classifier says other → keep as `eco-rescued`, and the reverse). Add `white_system`, `black_setup` tokens, `line_key_10`/`line_key_12` (EPD after plies 10/12), `opening_name` (lichess EPD map). Write `reports/flagged-games-review.md` (flagged games with their first 8 moves, flags, chess.com's name, link; counts by family and system; a recall cross-check of ECOUrl keywords versus the classifier). 3. **`s3_positions.py`**: the work queue of unique EPDs. Phase A: positions 8–24 of every in-family game, the 300 most recent games first. Phase B (after phase A screening): positions 25–40 for games not already decided at position 24 (|cp| < 500 and not over). Expect ~8.5k unique EPDs in A and ~9.5k in B. Idempotent: re-runs only append new EPDs; `data/queue_summary.json` keeps unique counts, the dedupe ratio and a per-ply histogram. 4. **`s4_engine.py --budget screen|confirm|publish --workers 2 [--epds file]`**: each worker owns one `SimpleEngine` (Threads/Hash from the budget), takes queue indices `i % workers == id`, skips cached rows, builds `chess.Board(epd + " 0 1")`, streams `engine.analysis(...)` and records per completed depth of multipv 1 `(depth, score_json, best_uci)` (ignore lowerbound/upperbound infos), then the final infos per multipv. Cache row: `{"budget","engine","depth","seldepth","nodes","time","lines":[{"rank","score_white","score_stm","first_san","pv_san","pv_uci"[:16],"depth"}],"history":[[d,score_json,uci],...],"terminal":null|"1-0"|"0-1"|"1/2-1/2"}`; `INSERT OR IGNORE` after each position. Robustness: on `EngineError`/`EngineTerminatedError` restart the engine and retry once, else log to `data/state/engine_failures.jsonl` and continue; each worker writes a heartbeat file (`state/worker-N.json`) per position and the parent kills and respawns a worker silent for more than 120 s; graceful stop on SIGTERM. Every 30 s the parent updates progress; every 5 minutes it calls the partial assembly (stage 5) for newly covered games so "flagged so far" and provisional leaders stay fresh. Memory: 2 × (128 MB hash + ~60 MB engine) + 2 Python ≈ 0.5 GB; the publish pass is one 2-thread engine with 512 MB. 5. **`s5_assemble.py`**: per-game move tables from the cache. For each Black move: `eval_before` (PV1 at k), `best_san`, `top3` (if confirmed), `eval_after` (PV1 at k+1, the eval of the played move), `loss_cp`, `loss_wp = wp_white(after) − wp_white(best_before)` clamped at 0, `class` (<3 ok, ≥8 flag, ≥10 inaccuracy, ≥20 mistake, ≥30 blunder), `d_seen`, `d_stab` (section 5b), `think_s` (previous clock − clock + the 2 s increment; the first move from 180), `premove` (think_s ≤ 0.3), `low_clock` (clock < 15 s), `familiar` (EPD reached ≥ 3 times in your games), `possible_mouseslip` (loss ≥ 25, same from-square as the best move, adjacent destinations), `punished` (White's reply kept ≥ 70 % of the win-probability gain, using position k+2), `category` (section 5d). Pick the deepest common budget for each (before, after) pair and never mix budgets inside one loss computation. Game-level: a `drift` block (5f), coverage, result, `lost_on_time`. Output `data/games_analyzed/.json` and `data/games_index.json`; expose `assemble_game(uuid)` and `assemble_partial()`. 6. **`s6_aggregate.py`**: `errors.jsonl` (Black moves with `loss_wp ≥ 8` at positions 9–30 or `≥ 12` at 31–40, with all features plus motif, theme, line keys, system, black core, weights, confidence screen|confirm), `habits.json`, `clusters.json`, `drift.json`, `stats.json` (by family, White system, Black move number, opponent band, category; date range; `d_seen` distribution; screen-versus-confirm agreement). `--emit-confirm-list` writes `data/confirm_epds.jsonl`: before and after EPDs of every flagged move, the 300 most frequent Black-to-move EPDs, and positions 16 and 24 of drift games (about 2.2–2.8k EPDs). `--partial` (first look) and full mode both write `reports/pirc-kid-candidates.md` (top 25 clusters and top 5 drift setups: FEN with side to move, your move distribution with scores, engine top-3 with evals and depth, `d_seen`/`d_stab`, category, motif, flags, three example games as date · your rating vs theirs · result · link, a suggested title) and `reports/figures/*.svg`, and append a paragraph to `run-log.md`. 7. **`s7_select.py`** picks 5–10 positions (section 6) into `data/selected.json`. 8. **Your review** writes `pipeline/content/positions.json` (section 6.3). 9. **`s9_verify.py`** runs the publish pass and the assertions (section 7) and emits `data/brochure.json` plus `reports/verify-report.md`. 10. **`s10_render.py`** writes boards, compiles the brochure, exports PNG pages and the study PGN (section 8). 11. **`run_all.sh`** is the unsupervised supervisor: `set -euo pipefail`; setup (including a smoke render) first so toolchain problems surface in minute 3; then stages 1–3, screening A, partial aggregate (first look), queue B, screening B, assemble + confirm list, confirm pass, aggregate + select + figures + a draft PDF with placeholder text. Launch it with `nohup setsid bash pipeline/run_all.sh > pipeline/data/logs/run_all.log 2>&1 &` so it survives anything that happens to the interactive session; `data/state/run.lock` holds the PID; a failure trap writes `FAILED at , see ` into `reports/progress.md`. Re-running from the top after a crash is safe because every stage skips or recomputes cheaply. ## 5. The level-calibration tricks (the heart of it) - **(a) Severity in win-probability loss, not centipawns.** Thresholds: flag 8, error 10, serious 20, blunder 30 (the 8 at screening absorbs depth-13 noise). In cluster scoring, severity is `min(loss_wp, 40)` so one catastrophe does not outrank a recurring 15-point leak. Mates count as ±1000 cp for win probability but are stored as mates. - **(b) Shallow-detectable filter.** From the free per-depth histories: `d_seen` = the smallest depth from which the *after*-position score (White's view) stays within 30 cp of that search's final score (the refutation is "visible"); `d_stab` = the smallest depth from which the *before*-position best move stays the final best. `shallow_factor` = 1.0 if `max(d_seen, d_stab) ≤ 10`, 0.7 if ≤ 14, 0.4 if ≤ 18, else 0.15. Calibration: depth 8–10 is a 4–5-move forcing sequence or an immediate structural concession, which a 2200 can see; depth 18+ is engine nuance. Screening (depth ~13) already settles "≤ 10"; the confirm pass refines the rest. This also makes a good chart ("how deep the engine must look to see my mistakes"). - **(c) Three-level clustering with a motif key** (`features.py`; compute setups on "squares ever occupied" so transpositions merge). `white_system`: Pirc/Modern → `austrian` (f4), `150` (Be3 + Qd2), `classical` (Nf3 + Be2), `bc4` (Bc4 ± Qe2), `byrne` (Bg5 by move 6), `h3-be3`, `fianchetto` (g3), `other-e4`; King's Indian → `samisch` (f3 + Be3 or f3 + c4 by move 6), `four-pawns` (f4 + e4), `classical` (Nf3 + Be2), `fianchetto` (g3 + Bg2), `averbakh` (Be2 + Bg5), `makogonov` (h3 + Nf3), `other-kid`; `london`/`torre`/`english` for flagged games. `black_setup` tokens: `c6 a6 b5 e5 c5 Nbd7 Nc6 Bg4 b6 Qa5 Nfd7 h5 late-Bg7 early-O-O uncastled-at-error`; `black_core` = the sorted structural subset of {c6, a6, b5, e5, c5, Nc6, Nbd7, b6}. `motif` = (best Black move token, White's refutation token), e.g. `(Pc5, Pe5)`, plus a coarse `theme` from the after-PV: `e5-push`, `f5-push`, `h-pawn-attack`, `central-capture`, `sacrifice`, `knight-jump` (Nd5/Nb5/Ng5), `queen-raid`, `piece-trap` (material won within 6 plies), `king-in-centre`, `missed-break` (best is …c5/…e5/…d5/…b5/…f5 and the played move was not a pawn move). Levels: **L1** exact `epd_before`; **L2** `(family, line_key_12, or line_key_10 for errors at positions 9–11)`; **L3** `(family, white_system, black_core, theme)`. An L3 cluster with ≥ 4 members over ≥ 3 distinct EPDs becomes a "pattern" candidate: its representative is the most frequent EPD and its drill comes from a second EPD in the same cluster, which is the "same mistake in slightly different positions" case. - **(d) Cluster score and the plan-move test.** `score = Σ_members[min(loss_wp, 40) · w_rec · w_opp · w_clock] · shallow_factor(median) · (0.5 + 0.5 · loss_rate) · teach_bonus · exposure`, with `w_rec = max(0.4, 0.5^(age_months/12))` (recent habits matter), `w_opp = clip(1 + (opp − me)/400, 0.6, 1.4)`, `w_clock` = 0.25 if `low_clock`, 0.5 if a pre-move in an unfamiliar position or a possible mouse-slip, else 1.0 (a pre-move in a familiar position is exactly the autopilot habit we want, so full weight), `loss_rate` = losses/members with draws 0.5 and `lost_on_time` games 0.5, `exposure` = 1.2 if the EPD was reached ≥ 5 times, `teach_bonus` = 1.3 if the habitual move differs from the best and the best is a **plan move**, 1.15 for a tactic with `d_seen ≤ 8`, 0.6 for engine-only, else 1.0. Plan move = 2 of 3: (1) stability: the best move is unchanged from depth ≥ 8 and its depth-8 score is within ±40 cp of the final; (2) quietness: the first 6 plies of PV1 contain ≤ 1 capture and no check, and the move kind is a pawn break, a reroute, an exchange (recaptured within 2 plies at level material), prophylaxis/quiet, or castling; (3) the alternative loses by force: the after-PV's first 4 plies contain ≥ 2 captures or checks, or a ≥ 60 cp jump at one depth increment with `d_seen ≤ 8`. Category: `plan` if (1) and (2); `tactic` if (3) and not (2); `mixed` otherwise; `engine-only` if `shallow_factor ≤ 0.15` and `loss_wp < 15`. - **(e) "What you usually play here" versus the engine.** For every Black-to-move EPD reached ≥ 3 times: your move distribution with results, the engine's top-3 (the 300 most frequent EPDs get confirmed regardless of flags), and `habit_gap_wp = wp(best) − wp(habitual)` using the habitual move's eval-after (always cached, because you played it). `habit_gap ≥ 6` with ≥ 5 reaches is a "repertoire leak" candidate even when no single game crossed the error threshold: this catches slow bleeds that per-game blunder detection misses, and it feeds the brochure's "in your games" stats line. - **(f) Drift / thin-ice detector.** Per game: `drift_wp = wp_white(position 24) − wp_white(position 8)` (the baseline is the line's own entry eval, so an Austrian Attack's +0.5 is not penalised); a drift game has `drift_wp ≥ 12`, no Black move in 9–24 losing ≥ 10, and `wp_white(24) ≥ 62`. Group by `(white_system, black_core)` and by `line_key_12`: share of drift games, the main line (modal move at each ply), and the **deviation point** (the earliest ply where your habitual move is ≥ 4 wp below the engine's best at the confirm budget). Present these as plan-level pages (the setup at the deviation point, White's plan in yellow arrows, the recommended plan in blue, prose about the structure), never as "blunder" pages. - **(g) Clock artefacts.** `think_s`, `premove`, `low_clock`, `familiar`, `possible_mouseslip` from `%clk`; show them in the candidates file so the review can discount; `w_clock` does the weighting; losses on time never count as "lost because of the opening". - **(h) Opponent-rating weights** as in (d); stats split by opponent band (< 2100 / 2100–2250 / > 2250); clusters report "punished 6/9, punishers averaged 2240". **(i) Punished flag** gives "found by opponents in k of n" and surfaces "not punished yet" leaks. **(j)** Opening names come from the lichess TSVs; the "what strong players do" framing comes from your chess knowledge at review time and is labelled as such. ## 6. Selecting the final 5–10 positions - **6.1 Pool.** L1 clusters with ≥ 2 member games; single-member L1 clusters only if `loss_wp ≥ 25` and the EPD was reached ≥ 4 times; habit leaks (`habit_gap ≥ 6`, ≥ 5 reaches); L3 pattern clusters (≥ 4 members, ≥ 3 EPDs); drift setups with ≥ 6 games and a drift share ≥ 40 %. Exclude `engine-only`, clusters that live entirely after position 36, and members with `low_clock`. - **6.2 Algorithm** (`s7_select.py`). Rank-normalise scores within type (error / habit / pattern / drift); `final = 0.7 · rank_score + 0.3 · teachability` (from `shallow_factor` and category). Greedy maximal-marginal-relevance: pick `argmax final · (1 − 0.6 · sim_max)` where `sim` = 0.85 same L2 line key, 0.6 same (system, theme), 0.5 same theme, 0.35 same system, 0.15 same family. Quotas: family slots proportional to game share, with a minimum of 2 for any family holding ≥ 15 % of games (so the King's Indian is covered), ≤ 2 per White system, ≤ 1 drift page per system and ≤ 2 overall. Default N = 8; stop early (minimum 5) when the next adjusted score is < 35 % of the first pick; cap 10. Tie-breaks: newer, then higher loss rate, then lower `d_seen`, then more reaches. Representative = the highest-weight member whose before/after chain is confirmed; examples = the top-2 by weight with distinct opponents, preferring one loss and one "you got away with it". Output `selected.json` plus the next 10 as alternates. - **6.3 Review (you).** Read `reports/pirc-kid-candidates.md`, spot-check each candidate with `chess-analyze --fen … --seconds 5 --lines 4` and `--moves `, then write `pipeline/content/positions.json`: ```json {"cover": {"title": "...", "subtitle": "...", "rules": {"p01": "Vs 4.f4: play ...c5 before castling"}}, "positions": [{"id": "p01", "cluster_id": "...", "family": "pirc", "system": "austrian", "title": "Austrian Attack: the e5 push you keep allowing", "fen": "", "habitual": "b5", "recommended": "c5", "arrows_extra": [["e4", "e5", "yellow"]], "key_squares": ["e5", "d5"], "idea": "3-6 sentences; evaluations only as {eval:best} {eval:habitual} {eval:line_end} {wp_loss}", "line": {"moves": "c5 Bb5+ Bd7 e5 Ng4 e6 fxe6 Bxd7+ Qxd7 Ng5", "claim": "equal"}, "pattern": "1-2 sentences", "lesson": "one bold line", "drill": {"fen": "...", "question": "Black to play. What now?", "answer": "e5", "claim": "black_fine"}, "examples": ["", ""]}]} ``` The verifier fills in the stats line, the evaluations, the example references (date · your rating vs theirs · result · move count · link; never a name) and the summary table. Titles in the "key moments" style, without numbers. - **6.4 Page template per position.** Title → diagram with arrows (red = what you usually play, blue = recommended, yellow = White's threat or plan) and the FEN → "In your games" stats line → "The idea" → one short line with its end evaluation → "Pattern to recognise" box → drill mini-board with the answer in muted small type → bold "Lesson:" → example games footer. ## 7. Verification of the chess content (`s9_verify.py`; rendering refuses to run on any FAIL) 1. Collect every EPD referenced (positions, drills, the end position of every printed line and of the "typical continuation") into `data/publish_epds.jsonl` and run `s4_engine.py --budget publish` on them (≤ 60 positions, about 20 minutes, cached). 2. **EXPECT_BEST**: `recommended` must equal the publish best move, or be within 15 cp of it and in the top 2 (then the page says "…c5 (or …O-O)"); otherwise FAIL with both moves and evaluations printed. 3. **Legality**: replay every line, continuation and drill answer with `board.parse_san`; FAIL on the first illegal SAN. 4. **Diagram FEN = reached position**: replay the representative game to the ply and compare EPDs; the drill FEN must also be reachable from one of the cluster's games. 5. **No hand-typed evaluations**: a regex such as `[+−-]\d+\.\d|#-?\d|\d+ ?cp` over all free text must match only inside `{…}` placeholders; `habitual` must be the modal move in that EPD (≥ 2 games). 6. **Claims**: the `claim` enum is checked against the publish evaluations: `white_better ≥ +0.8`, `equal |x| ≤ 0.35`, `black_fine ≤ +0.35`, `black_better ≤ −0.8`, `unclear` in (−0.8, +0.8); mates keep mate notation. 7. Build `data/brochure.json`, the fact table (per position: `eval_best`, `eval_habitual`, `eval_line_end`, `wp_loss`, depth, budget id, stats, examples, figures). The template renders only from it. `reports/verify-report.md` lists every check with pass/fail and the engine lines. ## 8. Brochure design (Typst) - **Toolchain** (`setup.sh`, run before the engine pass): download `https://github.com/typst/typst/releases/download/v0.15.1/typst-x86_64-unknown-linux-musl.tar.xz` (about 17 MB; extract with `tar -xJ`) into `/home/dev/chess/.tools/typst/` (fallback: `.venv/bin/pip install typst==0.15.0`, which bundles the compiler); `sudo apt-get install -y fonts-inter fonts-ebgaramond`; assert `typst fonts --font-path /usr/share/fonts` lists Inter and EB Garamond; smoke-compile a one-page test with a python-chess SVG board, a stat tile and a bar chart, and export a PNG (`typst compile --root --font-path /usr/share/fonts in.typ page-{p}.png --ppi 200`). Static TTF/OTF fonts only (not the variable Inter). - **Data flow.** `s10_render.py` copies `brochure.json` to `data/brochure/data.json`, writes `boards/.svg` (main 85 mm, mini 42 mm, drill 45 mm) with `chess.svg.board(board, orientation=chess.BLACK, arrows=[Arrow(a, b, color="red"), Arrow(c, d, color="blue")], lastmove=..., fill={sq: "#eda10066"}, colors=THEME)`, copies `reports/figures/*.svg`, compiles `brochure.typ` to `results/pirc-kid-key-positions.pdf` and page PNGs to `reports/brochure-pages/page-{0p}.png --ppi 200` (also `results/one-pager.png`), and writes `results/pirc-kid-key-positions.pgn` (one `[SetUp "1"][FEN …]` game per position; mainline = the recommended line, variation = the habitual move with its refutation; ASCII comments). - **Palette** (print-friendly; charts and pages share it): paper `#fbfaf7`, ink `#0b0b0b`, secondary `#52514e`, muted `#898781`, hairline `#e1e0d9`, baseline `#c3c2b7`; families Pirc blue `#2a78d6`, King's Indian orange `#eb6834`, Modern aqua `#1baf7a` (badges, chart series, page tabs); severity (always with a label) inaccuracy `#fab219`, mistake `#ec835a`, blunder `#d03b3b`. Boards: python-chess defaults (`#ffce9e` / `#d18b47`, last-move `#cdd16a` / `#aaa23b`); arrows through the named colours overridden in `colors`: `"arrow red": "#d03b3bb3"` (habitual), `"arrow blue": "#2a78d6b3"` (recommended), `"arrow yellow": "#eda100b3"` (White's threat or plan); key-square fill `#eda10066`; margin `#52514e`, coordinates `#fbfaf7`. One constant in `theme.typ` and `s10_render.py` switches to a green board scheme (`#eeeed2` / `#769656`) if preferred. - **Typography.** EB Garamond for the title and page headings (28 / 16 pt), Inter 9.5 pt body with 1.35 leading, Inter SemiBold 11 pt sub-heads, stat tiles in Inter 24 pt with proportional figures, tables with tabular numerals, DejaVu Sans Mono 7.5 pt for FENs and lines; fallback chain `("Inter", "DejaVu Sans")` so odd glyphs never go missing. SAN in letters (figurines optional, off by default). - **Pages** (A4 portrait, 14 mm margins, 9–13 pages). Page 1, the one-pager: title band (title, "N games · 2024-06 → 2026-10 · 3+2 blitz · Stockfish 19"), a row of 4 stat tiles (games analysed; score as Black; share "already worse by move 12"; errors per game in moves 5–20 or "opponents punished X %"), a 4×2 grid of 42 mm mini boards (family badge, arrows, a one-line rule ≤ 60 characters, page reference), and a footer "how to read: red = what you play, blue = recommended". Page 2, "Where you go wrong": errors per 100 games by Black move number 5–20 (bars with the peak emphasised), by White system (horizontal bars with n and score), by family (score, errors per game, "worse by move 12"), the `d_seen` histogram, and a 6-line method box; charts are hand-written SVG from `s6_aggregate.py` reusing `build_review.py`'s style (thin marks, hairline grid, direct labels, ≤ 3 series, a legend only with ≥ 2 series). Pages 3–N: one page per key position (60/40 two-column layout: diagram left, text right, drill and lesson across the bottom, examples footer). Last page: repertoire summary table (White system → games → your score → usual line → verdict with label → key rule and page) plus method notes (engine, budgets, thresholds, caveats: no explorer, blitz only, engine-only errors omitted) and the file list. - **Template mechanics.** `theme.typ` defines colours, fonts and the components `stat-tile(label, value, note)`, `badge(family, system)`, `board-figure(path, caption, fen)`, `pattern-box(body)`, `drill-box(path, question, answer)`, `examples-footer(list)`, `summary-table(rows)`; `brochure.typ` does `#let D = json("data.json")` and loops over `D.positions`. Pass every string through JSON, never into markup, so usernames and PGN text cannot break the template. ## 9. Run plan and time budget (2 cores) | Step | Wall-clock | Cumulative | |---|---|---| | setup incl. smoke render | 3 min | 0:03 | | download 29 archives | 1–2 min | 0:05 | | classify + review file | 1 min | 0:06 | | queue phase A | < 1 min | 0:07 | | screening phase A (recent 300 games first) | ~45 min | 0:52 (first look `pirc-kid-candidates.md` at ~0:25) | | queue + screening phase B | ~50 min | 1:45 | | assemble, aggregate, confirm list | 2 min | 1:47 | | confirm pass (~2.4k EPDs × 4.7 s / 2 workers) | 1.5–2 h | ~3:40 | | aggregate, select, figures, draft PDF | 3 min | ~3:45 (the unsupervised part ends) | | your review and spot checks | 30–60 min | | | publish pass + verify (≤ 60 positions) | 20–25 min | | | final render, PNG pages, PGN, report | 2 min | | Rates are node-limited, so a throttled host only stretches the ETA and never changes results. `reports/progress.md` is rewritten every 30 s by the active stage (keep it under 40 lines): stage, percent, ETA, running time; games downloaded and in-family by family; positions done/total, rate, cache-hit share, engine depth range, RSS; "flagged so far" and the last flagged move (date, opponent rating, result, move, system, loss); the provisional top-3 leaders; a stage checklist; pointers to the first-look file and the current log. Append a dated note to `reports/run-log.md` every 20–30 minutes while polling, and a paragraph per stage with counts. ## 10. Risks and mitigations | Risk | Mitigation | |---|---| | Engine nondeterminism | screen/confirm: Threads=1, node limits, `ucinewgame` per EPD; the publish pass is time-based by design and asserted with tolerance; the fact table stores the numbers actually produced | | False positives from pre-moves, flagged clocks, mouse-slips | `%clk` features, `w_clock`, `possible_mouseslip`, `lost_on_time`; all visible in the candidates file; the review drops suspect members | | Classifier misses games | ECOUrl keyword cross-check with `eco-rescued`, the review file with first 8 moves, the recall table in the run log | | Memory (no swap) | 2 × 128 MB hash in batch; 512 MB only in the single publish engine; RSS in progress.md; the watchdog restarts workers | | Engine hang (no timeout on node-limited searches) | heartbeat files + a 120 s parent watchdog that kills and respawns; failures logged, position skipped | | Cache keyed wrongly | key = (EPD, budget id) where the id encodes nodes/multipv/threads/version; a changed budget is a new id and old rows are harmless | | Typst fonts or SVG rendering | apt fonts + the `typst fonts` assertion + a smoke render in minute 3; DejaVu fallback; coordinates as paths; named arrow colours only | | Over-trusting the screening depth | confirm pass on every flagged move, the top-300 habitual EPDs and drift samples; the screen-versus-confirm agreement rate is reported | | Narrative over-claims | placeholders only, claim enums, EXPECT_BEST with tolerance, legality and FEN-reachability checks, your own spot checks with `chess-analyze` | | chess.com API hiccups | raw JSON stored, idempotent download, backoff; the current month is refetched | | Early stop drops late errors in decided games | acceptable (the brochure targets the opening); say so in the method notes | | L3 "pattern" clusters mixing ideas | the review decides; a drill drawn from a second EPD must pass the same assertions | ## 11. End-of-run checklist - [ ] `progress.md` shows every stage complete; `run-log.md` has a dated paragraph per stage with counts - [ ] ~1,900 blitz games, ~950 as Black, ~800 in family; `flagged-games-review.md` reviewed, no obvious Pirc/KID games left in `other` - [ ] `queue_summary.json` dedupe ratio and per-ply histogram look sane; cache row counts per budget match the queue sizes; `engine_failures.jsonl` empty or explained - [ ] the share of flagged moves surviving confirmation is reported (expect 60–80 %) - [ ] `verify-report.md`: every assertion PASS (best-move match, legality, FEN reachability, no hand-typed evaluations, claims, habitual = modal move, mates kept as mates) - [ ] every diagram FEN in the PDF equals the fact table; every printed line replays legally; evaluations in the PDF match `brochure.json` - [ ] diversity: both the Pirc and the King's Indian present (if KID ≥ 15 % of games); ≤ 2 positions per White system; 5–10 positions in total - [ ] `results/pirc-kid-key-positions.pdf` opens with embedded fonts and no substitution warnings; page PNGs in `reports/brochure-pages/`; the study PGN re-imports cleanly with python-chess - [ ] no username and no opponent names anywhere in `results/` or `reports/`; game links and ratings only - [ ] figures in `reports/figures/` render and match `stats.json` - [ ] `pirc-kid-report.md` lists the top-20 clusters beyond the chosen ones, the drift analysis and the method notes - [ ] scripts and `content/positions.json` are in `/home/dev/chess/pipeline/`; `run_all.sh` re-runs to completion in minutes from a warm cache