Robostrat

A strategy game and a workflow

Runs in your browser.

This project started as two questions:

The game

To fully explore these two questions, I needed the source code of a game. So naturally I developed my own game: Robostrat. Robostrat is a turn-based strategy game played on a small hex-tiled board for two players. It is designed to have a rule set simple enough that it is easy for a human to learn, but complex enough that a bot would need to be more than just a tree-searching script over the available moves. The ultimate goal is to develop a workflow for AI training that is transferable to other turn-based strategy games.

One GPU and a guided approach

Since my private data center consists of a sole Radeon RX 7900 XT, a pure self-play approach with no guidance (i.e. throw compute at it until it git gud) was not a reliable strategy to train a bot. Instead I needed a more guided approach, where specific weaknesses of the bot are identified by an outside observer and training is allocated to try to fix these weaknesses. This, together with a very exploratory approach to various bot designs, has been the main focus of this project.

How the AI works

The bot is a single neural network. It reads the board as a few dozen numbers per hex (terrain, owner, unit type, health, supply and so on) and answers with two things: how likely it is to play each possible action, and how likely it thinks it is to win. The network is a small convolutional network of about 1.1 million parameters, which looks at every hex the same way and so is not tied to one map. A turn in Robostrat is a sequence of single actions (move this unit, attack with that one, buy a unit, end the turn), so the bot picks one action at a time.

It is trained with PPO, a standard reinforcement learning algorithm, by playing about a thousand games at a time against itself and a pool of its older versions, on a GPU copy of the rules. Training runs in rounds of roughly 40 minutes on a mix of generated maps and the standard one. A new version only replaces the current best after being tested against it on eight maps that neither of them has trained on, and against a fixed older version as an independent check. The bot you play on this site is one of those champions, exported to run in your browser: no search, just the network, at about 17 milliseconds per action.

Where it stands

The current status of the project is that I have managed to train a bot that is somewhat competent on a variety of maps, in the sense that it sometimes wins against me. Once I learn the quirks from the last session of training, however, the bot usually fails to score any new victories against me. Regarding game design, the training process has helped me to identify and fix some flaws, but the workflow is still some way from the autonomous game balancing loop I had hoped for.

On this site

Published on this website is a version of the game, which can be played in a browser against one of the bots I trained, or watched while the bot plays against itself. Either side can also be played by a language model, which reads the same board text as the scouts in the devlog below. I also publish a devlog here, which dives deeper into some of the experiments done and challenges encountered in this project.