Skip to main content

Example: a game built by coding agents

A terminal Pac-Man in Python, standard library only, built with the method for agents on open models: a person wrote the plan and the tasks, OpenCode agents on models served by ColabHive executed them — each task in its own git worktree — and a command judged every one. It is public, including everything that went wrong.

Repository: github.com/jminich-ctrl/opencode, tag docs-2026-09-21; the example is in ejemplo-pacman/. The repository's documents are in Spanish and its README is in English. It also holds the complete OpenCode configuration of the model team described in Choosing a model team.

Estimated time: 10 minutes to read · To run the game and its tests: Python 3.11 or later


The setup​

PieceWhat it is
Plan8 tasks with dependencies, the files each one may touch, and stages that never run two tasks on the same file together. Written and approved by a person before any agent started
ExecutorThe OpenCode build agent: first Qwen3.8-27B FP8, then Qwen3-Coder-30B-A3B
Judgescripts/gate.sh: the whole test suite, no file changed outside the task's declared scope, no new dependency, no generated files, no stray TODO
Runnerscripts/correr-tarea.sh: one worktree, one branch and one OpenCode data directory per task, several tasks in parallel

The first task, the maze parser, was written by hand with its tests, as the quality reference the agent would imitate.

First round: 0 of 4 green on the first attempt​

TaskFinished byAgent attemptsResult
T01 mazea person—the reference, 15 tests
T02 Pac-Man movementthe agent3green on the third — but without the tunnel its task asked for
T03 pelletsa person2 failedthe agent never emptied the eaten cell
T04 ghostsa person1 unfinishedstill exploring when it was stopped
T05–T08a person—frightened mode, lives, rendering, game loop

What broke, and what each failure changed:

  1. Two opencode run at once died with database is locked. OpenCode keeps its state in one SQLite file. Each task now gets its own XDG_DATA_HOME — never its own XDG_CONFIG_HOME, or it loses the provider and the models.
  2. The gate passed a task that had not been done. The agent died before touching anything, and the gate found a healthy repository. Now a task with no changes is red, and an OpenCode error in the log makes the task red whatever the gate says.
  3. The agent reported a green gate that was red. It ran the gate, watched it fail three times, and wrote "GATE VERDE" in its summary. The runner searched the whole log for that text and believed it. Now the runner reads only the output of the gate it runs itself, after the agent has finished.
  4. A thinking model deliberated instead of acting. Qwen3.8-27B wrote 325 and 345 lines of reasoning in two attempts without touching a file. Qwen3-Coder-30B-A3B, an instruct model, wrote code on its first attempt. The thinking model moved to the reviewer role, where deliberation helps.
  5. A green gate did not mean a complete task. The agent that skipped the tunnel also wrote the tests, so no test could notice. A person found it in the diff review.
  6. The agent never commits. Its files stay untracked in the worktree, so merging the branch brings nothing: the integrator copies them by hand. (The runner could commit on green; it does not yet.)

Fewer than 30% green on the first attempt means the plan is the problem, not the model — and here it was: one task contradicted the map, and every task asked for a module and its tests in one step.

The game passed 54 tests and was bad to play​

With all eight tasks closed and every test green, the first minute of play found five problems: a single ghost, Pac-Man and a ghost passing through each other, a ghost stuck oscillating, lost keystrokes, and a game that felt frozen. They traced back to the plan, not to the agents. The map had one ghost and no task asked for four; the collision rule never defined two characters swapping cells; "no pathfinding" was an explicit decision; the loop read one key per 150 ms turn. And the plan's own final check — a real game played by hand — had been skipped.

What a command cannot check had dropped out of the plan, and the green gate felt like done. The method now turns those qualities into an explicit human check with concrete questions and an owner.

Second round: 3 of 4 green on the first attempt​

Four new tasks on the existing code, with the lessons applied: explicit contracts, the required tests named in each definition of done, and the human check written into the task.

TaskFileGate on the first attempt
T09 collision when characters crossjuego.pygreen
T10 ghosts with personalitiesentidades.pygreen
T11 smooth controls__main__.pygreen
T12 colours and scorerender.pyred — because of the gate

All ten tests required by name existed. T12 did its job and failed on a rule written by hand into the gate during the first round; the gate now reads the allowed files from the task itself.

Playing it — the human check done this time — found five more bugs with every test green (62 of them at that point): the game crashed at start (an argument in the wrong position), ghosts moved twice per turn, the loop slept without reading the keyboard, the end screen closed by itself, and drawing failed outside a terminal. All five were in code with no tests, excused because it "needs a real terminal". A fake screen — an object with the few curses calls the loop uses — now tests the loop, and would have caught all five.

The agent of T11 also reported green without playing the game, although its task said playing was the only way to check it. An instruction is not a check: a human check is done by a person.

What to take from it​

  • Write the plan as if it were the product. Almost every problem traced back to it.
  • Keep the verdict out of the agent's reach, and read it only from the command you ran.
  • Name the tests that must exist. Otherwise the agent that writes the code decides what is tested.
  • "Needs a terminal" usually means "needs a test double". What you cannot automate is how it feels, not whether it starts.
  • Play it, use it, open it. In both rounds, the worst problems were found by playing the game, not by the tests.

Run it​

git clone --branch docs-2026-09-21 https://github.com/jminich-ctrl/opencode.git agentes
cd agentes/ejemplo-pacman
python3 -m src.pacman # play (on Windows: pip install windows-curses)
bash scripts/gate.sh # the judge: tests, scope and hygiene

Launching agents needs OpenCode configured for ColabHive — the Quickstart for one model, or colabhive/README.md in the repository for the four-model team. The twelve tasks are done, so write a new one with the repository's task template: the two defects still open — ghosts that get stuck in some corners, and a chasing target that can fall outside the map — are good candidates.

Authors: José Luis Minich, Maximiliano Lucius.