Agent Uprising – Building a game with agentic code agents

Agent Uprising

Building Agent Uprising: a tower defense coded by AI agents, designed by humans

Agent Uprising is a browser tower defense game with an unusual constraint: it ships zero binary assets. No sprite sheets, no sound files, no image downloads. Every pixel of the turrets, enemies, scenery and effects, and every crack of a sniper shot, is generated by code from JSON recipes. The bulk of the code was written by AI coding agents, while the design and architecture were created by humans. That division of labor shaped almost every technical decision in the game, and this post walks through the interesting parts: the drawing pipeline, the sound engine, the entity art system, and the testing and balancing infrastructure that made it safe to let agents write the code.

How we used build agents

Agent Uprising started life as a test. We had built our own harness for long, unattended agentic coding runs, the kind you start before bed and read about over coffee, and we needed a project hard enough to tell us whether it actually worked. A tower defense game is a good torture test: it has a simulation, rendering, input, an economy, balance and a lot of content, and it fails in ways that are obvious the moment you try to play it. Every run used self-hosted models on our own hardware, so a night of work cost electricity rather than API credits. It also meant working with models that are capable but much less forgiving than the big hosted ones, which is exactly the pressure we wanted on the harness.

The harness follows an orchestrator and worker pattern. A manager agent holds the goal, but it has no shell and no file access, so it cannot quietly start coding on its own. Its only job is to break the work into task cards and dispatch them one at a time, reading each result before deciding what comes next. Each worker gets one card and nothing else: no memory of earlier tasks, no view of the manager’s conversation, no peek at another worker’s transcript. It does the work and returns a short structured report, and that summary is all the manager ever sees. The isolation is the point. A card that cannot be completed on its own comes back blocked, which forces the manager to write complete, self-contained cards and keeps every worker’s context small enough for a local model to hold. Every task, attempt, result and token lands in a ledger on disk, so a run that dies at 3am resumes where it stopped instead of starting over.

The most important idea is that “done” is not the agents’ call. Each run carries a set of verification scripts we wrote by hand, and a run only counts as finished when every one of them passes. A failing check does not end the run; it goes back to the manager as a new problem to plan around. We learned fast that what those checks measure decides what you get. Our first run produced a tidy MVP in about seventeen minutes. It passed every check, it drew a lovely board, and you could not play it. From then on the checks judged behaviour rather than code: they drove the game through a small player-facing contract (press these buttons, read this state), never named an internal, and asked whether a wave could actually be started, won and lost. That same philosophy is why the game itself leans so hard on tests, validators and the balance script described below. The agents needed the same machine-checkable definition of correct that the harness did.

We gave the briefs the same discipline. Specs describe what we want and why, not how to build it, and the design decisions are left to the model. The rigour lives in the checks instead of in step-by-step instructions. Over four overnight runs the harness took the game from an empty folder to an MVP, made that MVP genuinely playable, added the freeze, rocket and railgun turrets plus research upgrades and new maps in a run of almost eleven hours, and then spent another fourteen-hour night building themed maps, one task per map. Every run surfaced something the harness got wrong, and every harness fix shaped the next run.

The unattended runs got the game to a solid, data-driven foundation. From there we switched to interactive coding agents with a human in the loop for the campaign, the title screen, the polish and most of the remaining maps. That handoff worked because of what the overnight runs left behind: content as JSON, validators that explain their own errors, and a test suite and balance tool that say plainly whether the game is still the same game. Those turned out to be exactly the properties the rest of this post is about.

The engine and the drawing pipeline

The game loop is deliberately simple: requestAnimationFrame, a variable timestep, and a clamp of 100 ms on dt so a backgrounded tab cannot spiral the simulation. A TimeScaler multiplies dt by 0, 1, 2, 3, 4 or 8 for pause and speed modes, and the simulation, audio and animation all read the same scaled clock.

Rendering is one visible canvas with up to four offscreen canvases behind it, and the core idea is that nothing static should ever be drawn twice. Each map’s background, scenery, grid and roads are pre-rendered once into a full-resolution board canvas, cached by dimensions, and stamped with a single drawImage per frame. Roads are stroked as one merged path so junctions blend instead of overlapping. When the window resizes, the cache key changes and the board rebuilds.

The second big win is glow. Late waves put hundreds of glowing units on screen, and per-shape shadowBlur was what made those waves stutter. The fix: units are re-drawn into a half-resolution glow canvas with a flag set on the context (`ctx.glowPass = true`) so their art skips everything that is not light-emitting, that whole canvas is blurred exactly once, and the result is composited with globalCompositeOperation set to lighter. A tiny feature probe (set `ctx.filter` to `blur(1px)` and read it back) decides whether the layered path is available; old Safari and jsdom fall back to plain shadowBlur automatically. Display scaling is capped at 4K pixels so a giant monitor cannot melt the blur passes.

The sound engine

Every sound in the game is synthesized. A sound is a JSON recipe: layers of sine, square, saw, triangle or noise, each with an optional log-shaped frequency slide from A to B, an attack/hold/decay envelope, and an optional state-variable filter. Noise comes from a seeded mulberry32 generator, which means a sound is exactly its recipe: the same explosion on every machine, byte for reproducible if you want it.

The synthesis math lives in a plain-JS module with no Web Audio in it, so it can run in Node and be unit tested by checking the numbers it produces. At page load the recipes are pre-rendered into AudioBuffers; during gameplay no oscillators or filters run at all, only BufferSource plus gain nodes. Voice management keeps chaos under control: a per-sound minimum interval, a per-sound voice cap, and a global cap of 24 voices with priority eviction, so an explosion is heard over the rattle of rapid fire. Sounds pan across the board with playAt(x), buses mix sfx, ui and music, and a DynamicsCompressor limiter sits on the master. At 8x speed the sound engine thins out voicing instead of changing pitch.

Drawing turrets and enemies in code

There is no entity component system and no pretense of one: plain classes plus a data-driven art module per style. Each of the eight turrets has a hand-designed vector art file. Barrels track targets with a rate-limited turn so they swing instead of snapping, firing plays a 0.18 s recoil and a 0.07 s muzzle flash, and overlays draw the things that live above enemies: sniper sight lines, Tesla chain lightning. Specialisations add corner chevrons so you can read a build at a glance.

Enemies ship in thirteen art styles, from swarmers to juggernauts, with per-enemy lane offsets so crowds walk parallel lanes instead of single file. Status is drawn physically: health bars with color thresholds, shield arcs, a hexagonal frost crust with spikes when slowed. For big waves the same trick as the glow layer applies: the vector art is rendered once per (style, colour, size) into a small cached canvas and stamped as a sprite for the rest of the run, so the sheets are generated from code at runtime and discarded on resize. Projectiles come from a 200-object pool and draw as a glow dot with a 30 ms tracer streak.

Test-driven development: the reliability layer for agent runs

This is where the human/agent split really paid off. Agents are fast and they are tireless, but a chunk of code you did not write yourself needs a machine-checkable definition of correct. The test suite is that definition: 68 spec files and roughly 800 test cases, run with vitest, and the engine is built so the suite never needs a browser. Every engine test uses the same two-line idiom, a fake canvas object that accepts any draw call, and then steps the simulation by hand: `for (let t = 0; t < 60 * 60 && game.waveActive; t++) game.update(1 / 60)`. Fourteen spec files contain full headless playthroughs.

Some favorites: a synthetic fork map is injected purely as data, then the test asserts every enemy starts at a spawn, ends at an exit, and costs exactly one life. On the Serpentine map, 40 enemies walk the whole curve and the test asserts none of them ever triggered the pathing engine’s anti-stuck safety net, which proves the real maps are clean rather than assuming it. And the safety net for new content itself: a loop over every turret in gameData.json asserting each one unlocks, builds, upgrades to top level with stats that never decrease, and renders its panel. Add a turret to the JSON and it inherits all of that coverage for free.

Validation is the other half of the contract. Load-time validators (validate.js for game data, the GameMap constructor for maps, validateCampaign for the campaign) throw human-readable, file-named errors before a window ever draws: `map "candy" route "r" names unknown waypoint(s): b`, `waypoint "exit" (5000, 380) is outside the 1280x720 world`. For an agent, that error message is the fix instruction. Every task an agent was handed ended with “npm test must pass”, which turned the suite into a regression net against the drift and drift-adjacent mistakes that large-scale agent coding produces.

Data-driven content so many agents could build many maps

The maps are the clearest example of designing for parallel agent work. All 37 maps (27 free-play plus 10 campaign) are pure JSON: named waypoints, routes that share stretches, build rules, waves, and a theme block for colors, particles and scenery. No map touches code. Scenery placement is seeded and deterministic, so a map that looked good in the editor looks identical every run and nobody has to hand-place trees. The GameMap constructor is the schema, the validators are the linter, and the authoring guide in src/data/README.md tells a contributor exactly which two commands to run.

The practical result: we could hand map N to one agent and map N+1 to another, in separate sessions, with zero merge conflicts, because each map is one file with a built-in definition of done. The campaign works the same way: 10 stages of pure data with turret, specialisation and research whitelists, and stageGameData hands the engine a filtered copy of the world so a stage literally cannot see content it has not granted.

The balance script: winnable, but not too easy

Difficulty is checked, not hoped for. tools/balance.mjs is a headless auto-player that plays every map with three simulated skill levels (good buys freely, medium three turrets per wave, casual one) and scores each candidate build tile per dollar, factoring in damage, coverage and crowd control, with a synthetic crowd injected so splash and slow turrets get credit for the ground they hold. A map is in band when the good player wins with 5 to 14 lives left, the medium player loses around waves 6 to 8, and the casual player falls by wave 4. Out of band prints OFF BAND and exits 1. A `--without rocket,freeze` flag enforces a design invariant: no single turret may be mandatory. Runs are seeded for reproducibility, and the slow-weight constant in the scoring function is an empirical playtest finding, documented in-file. Every campaign stage was tuned through this tool first, and then verified by a human: every stage has been played through by the author, twice.

Closing

The engine is vanilla JavaScript on a single canvas, the audio is arithmetic, the art is paths, and the content is JSON. None of that is nostalgia. Every layer of indirection earns its keep by making the thing checkable, and in a project where agents wrote most of the code, checkable was the most valuable property we could design for. The humans decided what the game should be; the agents built it; the tests, validators and balance bots made sure it was the same game when they finished.

Joss Malasuk Avatar

Written by

Previous post