Skip to content

Latest commit

 

History

History
108 lines (89 loc) · 5.46 KB

File metadata and controls

108 lines (89 loc) · 5.46 KB

TransformerChessEngine - Implementation Plan

1. Vision & Philosophy

Goal: specific high-performance hybrid chess engine. Core Concept: "Intuition + Reasoning".

  • Intuition: A Transformer model that "sees" the board as 64 inter-connected squares (spatially aware) and proposes candidate moves.
  • Reasoning: A GPU-accelerated Search (Batched MCTS or Beam Search) that explores the most promising "Chain of Thought" lines in parallel, rather than brute-forcing millions of nodes like traditional AlphaBeta pruning.

2. Architecture: "The Spatial Transformer"

Unlike text LLMs, our tokens are squares. The board is a fixed sequence of 64 tokens.

A. Input Representation (The Board Tokenizer)

  • Input Shape: [Batch_Size, 64, d_model]
  • Encoding: Each of the 64 squares is projected into a vector.
    • Piece Embedding: (Empty, W_Pawn, B_Pawn, ...) -> Vector.
    • Positional Embedding: A learnable vector for specific squares (A1, H8, etc.).
    • Global State: (Castling rights, En passant, Turn color) can be appended as a special "CLS" token or injected into all squares.

B. The Backbone (Encoder-Only Transformer)

  • Type: BERT-style Encoder (Bidirectional Attention).
  • Reasoning: Every square needs to attend to every other square to understand lines of attack/defense instantly.
  • Specifications (optimized for RTX 2070S 8GB):
    • Layers: 12
    • Attention Heads: 8
    • Embedding Size (d_model): 512
    • Feed Forward: 2048
    • Dropout: 0.1
    • Activation: GELU
    • Total Parameters: ~25-30 Million.

C. The Prediction Heads

  1. Policy Head (The "Move" Prediction):
    • Type: Flat Move Map Classification.
    • Output: A probability distribution over a fixed vocabulary of all ~1968 theoretically possible unique chess moves (UCI format).
    • Process: Softmax over the move vocabulary.
  2. Value Head (The Evaluator):
    • Output: A single scalar [-1, 1] via Tanh.
    • Purpose: Scores the board state for the search tree.

3. The "Hybrid" Inference Engine (GPU Search)

Strategy: "Aggressive Batched Tree Search" (Beam-like with Persistence).

  • Core Philosophy: "Trust the Intuition first, verify later."
  • Why: Optimizes for the model's strengths (openings/midgame) and maximum GPU throughput.
  • Persistence: We maintain the search tree in memory. When a move is made, we reroot the tree to the new state, preserving previously calculated branches.

Detailed Search Flow (Batched Depth-First / Beam)

  1. Selection (Aggressive Pruning):
    • From the current positions, pick the Top K (e.g., 5) moves based on Policy + Value.
    • Discard the rest (Trusting the model prevents exploring "bad" moves).
  2. Expansion & Batching:
    • Collect all candidates from the selection phase.
    • Deduplication: Check the Transposition Table (Hash Map) to see if we've already evaluated this position.
    • Batch Eval: Pack all new positions into a single Tensor (Size: 64-256).
    • Run model.forward() ONCE.
  3. Backpropagation/Update:
    • Store the results in the Tree Nodes.
    • Update the Transposition Table.
  4. Tree Persistence (The "Memory"):
    • The Search object is persistent across moves.
    • make_move(move): The tree simply updates self.root = self.root.children[move].
    • We instantly have a head start on the next search.

Handling "Opponent's Turn" (Minimax Logic)

  • The Search Tree handles turns natively.
  • White's Turn: Select moves ensuring high Value (+1.0).
  • Black's Turn: Select moves ensuring low Value (-1.0).
  • This "Negamax" logic is built into the Node selection.

Search Hyperparameters

  • Beam Width: Start aggressive (e.g., 5) to go deep fast.
  • Max Batch Size: 256 (Utilization of 8GB VRAM).
  • Depth: Adaptive (Time-based).

4. Data Pipeline

  • Source: Lichess Elite Database / CCRL (Computer Chess Rating Lists) / Stockfish self-play games.
  • Format:
    • Input: FEN (converted to 64-token tensor).
    • Label 1: The move actually played (Policy target).
    • Label 2: The final result of that game (Value target).

5. Development Roadmap

Phase 1: Foundations

  • Set up project structure (src/, data/).
  • Critical: Implement MoveVocabulary generator: Ensure 100% coverage of every geometrically possible move (including promotions) to prevent training errors.
  • Implement BoardTokenizer: Convert python-chess board objects to (64,) tensors.
  • Implement MoveTranslator: Convert moves (e.g., e2e4) to integer indices using the vocabulary.

Phase 2: The Model

  • Build the TransformerChessModel in PyTorch.
  • Implement "Spatial Attention" blocks. (Standard TransformerEncoder utilized)
  • Create the Training Loop (Supervised Learning on Game Database).

Phase 3: Training

  • Train on a small dataset (10k games) to verify loss convergence (checking if it learns legal moves).
  • Scale up to larger datasets.

Phase 4: The Interface (CLI)

  • Build the Engine class.
  • Integrate with python-chess for move validation and state management.
  • Create a CLI loop: User inputs move -> Engine thinks (Search) -> Engine plays.

6. Future Expansion

  • Web App: Once the engine is smart, wrap it in a FastAPI backend and build a React frontend.
  • Self-Play RL: Refine the model by letting it play against itself (AlphaZero style).