← Robostrat

Letting language models play against my AI

How the game is turned into text a language model can play, and which models turned out to be good at it.

Robostrat devlog · September 2026

Self-play and its discovery rate

The Robostrat AI learns by self-play: it plays against copies of itself and gradually shifts towards the moves that won. Nothing in principle stops self-play from finding any strategy the game allows. Given enough compute it would get there. What limits it in practice is how fast it discovers new strategies.

A strategy only enters training once the AI happens to play it, often enough for the payoff to show. A trick that pays off after a single move is found almost at once. A plan that needs five coordinated moves before it pays anything needs the AI to stumble on all five in a row, and the odds of that shrink with every step. The same goes for defence: a weakness is only trained away once some opponent in training exploits it. On one consumer GPU, some plans will not be found in any reasonable time.

For a while the outside help was me. Every strong scripted opponent I wrote came from someone noticing something the AI could not handle, usually me. That does not scale, and it depends on me having the idea. What I wanted was an automatic version of that playtester: something that plays differently from the AI, brings ideas of its own, and writes down what it tried. Language models were the obvious candidate, because they can reason. Given a position they have never seen, they can work out what the opponent threatens, weigh a plan against the alternatives, change course when it fails, and explain what they did.

Turning the game into text

The scouts play through an MCP server wrapped around the Python copy of the game engine that training runs on (more on why there are two engines in a later article). Any MCP client can drive it, and models without tool support play through a small script over plain chat instead.

The first version described the board the way a programmer would: every unit with its coordinates, every owned tile, and not much else. The scouts lost almost every game on the economy, early. Their transcripts showed why. Every turn they had to work out from raw coordinates which hexes are neighbours, what an enemy can reach and whether a tile still connects to their base, and they got it wrong without noticing. Language models reason well about strategy in words and badly about geometry on a hex grid. So the board text now does the geometry and leaves the strategy to the model. Every turn, the model gets:

Here is part of one board, from a test game in which Blue, the model's side, has fallen behind on income:

TURN 6 - you are BLUE. Credits 150 (enemy 140).
INCOME you 50/turn vs enemy 180/turn - you are BEHIND by 130.
  income by turn (yours v theirs): t1: 50v50  t2: 50v50  t3: 50v80  t4: 50v130  t5: 50v130  t6: 50v180
[…]
BOARD - 2 chars per hex: the TERRAIN, then whatever stands on it (or who owns it).
  r-2 |  # . ^ . $ ^ . . ^ . . @-. . # | q-6..+8
  r-1 | # . ^ . * ^+.+.I^ . $-^-. . ^ # | q-7..+8
  r+0 |# . ^ . * ^I. . *r.-. ^ *-. ^ . # | q-8..+8
  r+1 | # ^ . . ^+$ . ^ . .i^-*-. ^ . # | q-8..+7
  r+2 |  # . . @+. . ^ . . ^ $a. ^ . # | q-8..+6
[…]
SUPPLY CHAIN (your owned tiles, adjacency-connected; only chain A can heal/arm):
  A: 6 tiles, holds your base, units U2 U1
  ARTICULATION POINTS - if the enemy captures one of these, the units listed go UNSUPPLIED […]
    (-4,1) cuts off U1 U2 (4 tiles)
    (-3,0) cuts off U1 (3 tiles)
[…]
YOUR UNITS (2):
  U1 Infantry  (0,-1)   hp8   sup4/4  mv1/1  chainA  can: attack,move,overwatch
     moves: (-1,-1)Pln (-1,0)Pln (0,-2)Pln (1,-2)Pln (1,-1)For
[…]
ENEMY UNITS visible (3):
  E2 Recon     (0,0)    hp7   sup0/2  DISARMED (0 supply: cannot attack or counter)
[…]
ATTACKS AVAILABLE THIS TURN (what the defender's side would answer with):
  U1 -> E2 Recon(0,0) hp7: YOU DEAL 4, YOU TAKE 0; no overwatch pre-empt; target at 0 supply:
     NO counter-attack; on a kill you ADVANCE to (0,0), a Fac you do not own yet (+30/turn) […]
CAPTURES AVAILABLE THIS TURN (move onto an income tile you do not own):
  U2 -> (-4,0) Fac: +30/turn for you, stays joined to chain A, not in reach of any enemy afterwards

Trimmed; a full board is about three times as long.

Nothing here asks the model to work out geometry. The text has already found that the enemy Recon next to U1 has run out of supply and cannot hit back, that a kill would move U1 onto a factory, and that losing the single hex (-4,1) would cut both units off from their base. What is left is the actual decision: take the free hit, grab the factory at the back, or fix the weak supply line first.

Two rules keep this text honest. Everything it lists as possible comes from the engine itself, and the attack prices are not computed by a copy of the damage formula: the server plays the attack out on a copy of the game and reports what happened, so the text cannot drift from the rules. And every sentence in it that states a rule is registered against a test that proves the rule. Of everything we tried to make the local models play better, changes to this text were the only thing that clearly moved the results.

The model then answers with short commands, one action at a time, and gets a fresh board after each one, the way a human player clicks a unit and looks before clicking the next. (The first four scout games each committed a whole turn from one snapshot, and lost all four.)

move U1 to (-3,0)       attack U2 -> E1        overwatch U2 facing (1,0)
buy Recon at base       range U3 -> E1         end turn

A few tools sit around the commands:

ToolWhat it is for
simulateTry a line of moves without committing it, including the opponent's replies. Those replies are the scout's own guesses: the real AI is never consulted, because a win found by searching against the AI itself would say nothing. A test checks this.
planA standing plan, reprinted above every board.
strategy_noteA playbook per model that survives between games, which the other models can read.
reportFindings attached to the game's log.

Which models play best

We tried three cloud models through their own tools and two local models running on the same GPU the AI trains on. None of this is a proper strength test: the samples are small, and the AI and the rules changed between rounds. What the games do show is how the models compare with each other. For the speed column, a typical scout game lasted 15 to 25 turns.

ModelReasoningAgainst the AISpeed
Gemini 3.8 Flash, through AntigravityHighGood enough: wins regularlyAbout 15 seconds per turn
Claude Opus 5, through Claude CodeHighWon its one completed gameAbout 2 minutes per turn
Claude Sonnet 5, through Claude CodeHighToo weak: lost every game, including on the map and seat where Gemini wonUnder a minute per turn
Qwen3.8 27B, localOff, low and highestOne win, after more than thirty losses, and only with Gemini's playbook in its prompt1 to 2 minutes per turn with reasoning off, 5 to 10 with it on
Gemma 4 26B‑A4B, localOffNever wonAbout 2 minutes per turn

Both local models ran as 4-bit quantisations.

The clearest single result came from taking the simulator away from the strongest player. On one generated map Gemini won 3 of 3 with it. With the same map, seat and seeds and the simulator disabled, it lost 3 of 3, and lasted about as long as the local models do.

Weighing strength and speed together, Gemini 3.8 Flash at high reasoning is the best fit for the job so far. Sonnet was simply too weak to trouble the AI. Opus won too, but it takes about eight times as long per turn, and a scout is only useful if it can play a batch of games against each new version of the AI. The local models cost nothing to run, but Gemma never won, and Qwen's one win took more than thirty attempts, reasoning turned on, and another model's playbook.

A playtester, not a player

The scouts' most useful output has been things no test had seen. One noticed that an Artillery's ranged attack draws no counter-attack, and that the AI almost never used it: measured afterwards, in 4.5% of the turns where it could. That became a line of training work of its own. Others turned up bugs that only show when whole games are played and then replayed, such as a scripted opponent whose randomness made replays drift apart, game logs that silently pointed at the wrong version of the AI after every promotion, and an income bug with disconnected cities.

So why not let a language model be the AI? Because a scout needs seconds to minutes per move and costs money per game, while the trained network picks a move in milliseconds on a laptop and plays thousands of games in the time a scout plays one. And a language model does not get stronger by playing more games. The scouts' job is to find out what the AI should learn; the learning itself stays with self-play.