Letting language models play against my AI
How the game is turned into text a language model can play, and which models turned out to be good at it.
Self-play and its discovery rate
The Robostrat AI learns by self-play: it plays against copies of itself and gradually shifts towards the moves that won. Nothing in principle stops self-play from finding any strategy the game allows. Given enough compute it would get there. What limits it in practice is how fast it discovers new strategies.
A strategy only enters training once the AI happens to play it, often enough for the payoff to show. A trick that pays off after a single move is found almost at once. A plan that needs five coordinated moves before it pays anything needs the AI to stumble on all five in a row, and the odds of that shrink with every step. The same goes for defence: a weakness is only trained away once some opponent in training exploits it. On one consumer GPU, some plans will not be found in any reasonable time.
For a while the outside help was me. Every strong scripted opponent I wrote came from someone noticing something the AI could not handle, usually me. That does not scale, and it depends on me having the idea. What I wanted was an automatic version of that playtester: something that plays differently from the AI, brings ideas of its own, and writes down what it tried. Language models were the obvious candidate, because they can reason. Given a position they have never seen, they can work out what the opponent threatens, weigh a plan against the alternatives, change course when it fails, and explain what they did.
Turning the game into text
The scouts play through an MCP server wrapped around the Python copy of the game engine that training runs on (more on why there are two engines in a later article). Any MCP client can drive it, and models without tool support play through a small script over plain chat instead.
The first version described the board the way a programmer would: every unit with its coordinates, every owned tile, and not much else. The scouts lost almost every game on the economy, early. Their transcripts showed why. Every turn they had to work out from raw coordinates which hexes are neighbours, what an enemy can reach and whether a tile still connects to their base, and they got it wrong without noticing. Language models reason well about strategy in words and badly about geometry on a hex grid. So the board text now does the geometry and leaves the strategy to the model. Every turn, the model gets:
- the map drawn in characters, two per hex: the terrain, then the unit standing on it or the tile's owner;
- a tempo map of which hexes each side could step onto next turn, since stepping onto a tile is what captures it;
- both bases, with the six hexes an attack on each must come from, and who can reach them;
- the supply network as connected groups, including the single hexes whose loss would cut units off from their base;
- every income tile, with its distance and whether it can be taken this turn;
- every unit with its legal moves, read straight from the engine's own list of legal actions;
- every available attack, priced by playing it out on a copy of the game: damage dealt, damage taken, and what happens next;
- the turn count, both sides' income over the last few turns, and a plan the model wrote for itself, reprinted every turn with the current status of every hex it names.
Here is part of one board, from a test game in which Blue, the model's side, has fallen behind on income:
TURN 6 - you are BLUE. Credits 150 (enemy 140).
INCOME you 50/turn vs enemy 180/turn - you are BEHIND by 130.
income by turn (yours v theirs): t1: 50v50 t2: 50v50 t3: 50v80 t4: 50v130 t5: 50v130 t6: 50v180
[…]
BOARD - 2 chars per hex: the TERRAIN, then whatever stands on it (or who owns it).
r-2 | # . ^ . $ ^ . . ^ . . @-. . # | q-6..+8
r-1 | # . ^ . * ^+.+.I^ . $-^-. . ^ # | q-7..+8
r+0 |# . ^ . * ^I. . *r.-. ^ *-. ^ . # | q-8..+8
r+1 | # ^ . . ^+$ . ^ . .i^-*-. ^ . # | q-8..+7
r+2 | # . . @+. . ^ . . ^ $a. ^ . # | q-8..+6
[…]
SUPPLY CHAIN (your owned tiles, adjacency-connected; only chain A can heal/arm):
A: 6 tiles, holds your base, units U2 U1
ARTICULATION POINTS - if the enemy captures one of these, the units listed go UNSUPPLIED […]
(-4,1) cuts off U1 U2 (4 tiles)
(-3,0) cuts off U1 (3 tiles)
[…]
YOUR UNITS (2):
U1 Infantry (0,-1) hp8 sup4/4 mv1/1 chainA can: attack,move,overwatch
moves: (-1,-1)Pln (-1,0)Pln (0,-2)Pln (1,-2)Pln (1,-1)For
[…]
ENEMY UNITS visible (3):
E2 Recon (0,0) hp7 sup0/2 DISARMED (0 supply: cannot attack or counter)
[…]
ATTACKS AVAILABLE THIS TURN (what the defender's side would answer with):
U1 -> E2 Recon(0,0) hp7: YOU DEAL 4, YOU TAKE 0; no overwatch pre-empt; target at 0 supply:
NO counter-attack; on a kill you ADVANCE to (0,0), a Fac you do not own yet (+30/turn) […]
CAPTURES AVAILABLE THIS TURN (move onto an income tile you do not own):
U2 -> (-4,0) Fac: +30/turn for you, stays joined to chain A, not in reach of any enemy afterwards
Trimmed; a full board is about three times as long.
Nothing here asks the model to work out geometry. The text has already found that the enemy Recon next to U1 has run out of supply and cannot hit back, that a kill would move U1 onto a factory, and that losing the single hex (-4,1) would cut both units off from their base. What is left is the actual decision: take the free hit, grab the factory at the back, or fix the weak supply line first.
Two rules keep this text honest. Everything it lists as possible comes from the engine itself, and the attack prices are not computed by a copy of the damage formula: the server plays the attack out on a copy of the game and reports what happened, so the text cannot drift from the rules. And every sentence in it that states a rule is registered against a test that proves the rule. Of everything we tried to make the local models play better, changes to this text were the only thing that clearly moved the results.
The model then answers with short commands, one action at a time, and gets a fresh board after each one, the way a human player clicks a unit and looks before clicking the next. (The first four scout games each committed a whole turn from one snapshot, and lost all four.)
move U1 to (-3,0) attack U2 -> E1 overwatch U2 facing (1,0)
buy Recon at base range U3 -> E1 end turn
A few tools sit around the commands:
| Tool | What it is for |
|---|---|
simulate | Try a line of moves without committing it, including the opponent's replies. Those replies are the scout's own guesses: the real AI is never consulted, because a win found by searching against the AI itself would say nothing. A test checks this. |
plan | A standing plan, reprinted above every board. |
strategy_note | A playbook per model that survives between games, which the other models can read. |
report | Findings attached to the game's log. |
Which models play best
We tried three cloud models through their own tools and two local models running on the same GPU the AI trains on. None of this is a proper strength test: the samples are small, and the AI and the rules changed between rounds. What the games do show is how the models compare with each other. For the speed column, a typical scout game lasted 15 to 25 turns.
| Model | Reasoning | Against the AI | Speed |
|---|---|---|---|
| Gemini 3.8 Flash, through Antigravity | High | Good enough: wins regularly | About 15 seconds per turn |
| Claude Opus 5, through Claude Code | High | Won its one completed game | About 2 minutes per turn |
| Claude Sonnet 5, through Claude Code | High | Too weak: lost every game, including on the map and seat where Gemini won | Under a minute per turn |
| Qwen3.8 27B, local | Off, low and highest | One win, after more than thirty losses, and only with Gemini's playbook in its prompt | 1 to 2 minutes per turn with reasoning off, 5 to 10 with it on |
| Gemma 4 26B‑A4B, local | Off | Never won | About 2 minutes per turn |
Both local models ran as 4-bit quantisations.
The clearest single result came from taking the simulator away from the strongest player. On one generated map Gemini won 3 of 3 with it. With the same map, seat and seeds and the simulator disabled, it lost 3 of 3, and lasted about as long as the local models do.
Weighing strength and speed together, Gemini 3.8 Flash at high reasoning is the best fit for the job so far. Sonnet was simply too weak to trouble the AI. Opus won too, but it takes about eight times as long per turn, and a scout is only useful if it can play a batch of games against each new version of the AI. The local models cost nothing to run, but Gemma never won, and Qwen's one win took more than thirty attempts, reasoning turned on, and another model's playbook.
A playtester, not a player
The scouts' most useful output has been things no test had seen. One noticed that an Artillery's ranged attack draws no counter-attack, and that the AI almost never used it: measured afterwards, in 4.5% of the turns where it could. That became a line of training work of its own. Others turned up bugs that only show when whole games are played and then replayed, such as a scripted opponent whose randomness made replays drift apart, game logs that silently pointed at the wrong version of the AI after every promotion, and an income bug with disconnected cities.
So why not let a language model be the AI? Because a scout needs seconds to minutes per move and costs money per game, while the trained network picks a move in milliseconds on a laptop and plays thousands of games in the time a scout plays one. And a language model does not get stronger by playing more games. The scouts' job is to find out what the AI should learn; the learning itself stays with self-play.