A method for agents on open models
This is not the method you would use with the largest proprietary models. The difference comes down to one sentence:
A large model holds a goal and finds its way. A smaller model executes a bounded, verifiable task. You do the decomposition, the model does the execution, and a command is the judge.
What follows is the method ColabHive uses to drive projects with coding agents on open 20–30B models, and the failures that shaped it.
Three principles
The unit of work is the task, not the feature
A task sized for a 20–30B model:
- touches one or two files, never ten;
- has one goal, stated in a sentence;
- is checked by a command that passes or fails;
- would take a person on the team 30 to 90 minutes.
If you cannot write its definition of done in one line, the task is too big. Split it.
The judge is a command, not the model
A smaller model says "done" with the same confidence whether it works or not. No task is finished because the agent says so; it is finished when the gate passes. A task with no command to check it is badly defined.
The corollary is expensive to learn: reading the verdict has to be out of the agent's reach too. It happened: the model ran the gate, watched it fail three times, and wrote "the gate is green" in its summary. The runner searched the log for that text and believed it. Now the runner reads only the output of the gate it runs itself. Instructions asked the model not to misreport; it did anyway. What holds is the verification architecture, not rules of conduct.
What a command cannot check does not vanish from the project — it vanishes from the plan. Because every criterion must be verifiable, the plan fills with what can be tested and silently drops what is qualitative: that it feels smooth, that the result is usable. Nobody writes it down, so nobody builds it, and a green gate feels like done. When a quality matters and cannot be tested, turn it into an explicit human gate with concrete questions, an owner and a time box.
A green gate does not prove the task is complete either. It checks what the tests cover; if the same agent writes the code and the tests, it can leave out a whole requirement without anything showing it. Name in the definition of done the tests that must exist, or split "write the tests" and "make them pass" into two tasks.
Width, not depth
On hardware you run, a task costs electricity, not billed tokens; on the hosted service capacity is billed per accelerator-hour, not per token. Either way, one more task in parallel is cheap. So:
- launch several tasks in parallel, each in its own worktree — see Running agents in parallel;
- if one comes out wrong, throw it away and relaunch it with a better prompt;
- do not negotiate with the model to rescue a bad answer: it costs more of your time than a relaunch;
- what fails twice, you do with a larger model or by hand.
The cycle
| Gate | What it is | Who |
|---|---|---|
| G0 — Plan approved | A plan with numbered tasks, their dependencies and their criteria. No agent starts without it; it is the gate that saves the most time | a person approves |
| G1 — Task gate | On the task's worktree: the full test suite passes, no file outside the task's declared scope changed, no new dependencies or stray TODOs. Red in any of the three means not done | automatic |
| G2 — Diff review | Read it like the pull request of someone new: is the fix right or a patch that makes the test pass? Do the tests really fail against the old code? Did it touch what it should not? | a person, helped by a read-only reviewer agent |
| G3 — Integration | All tasks of a stage merged, the full gate on the trunk — and a human end-to-end use test, written into the plan with concrete questions and an owner | automatic, then a person |
G3's human test is not optional. In the example project — a Pac-Man game built with this method — the gate was green, all 54 tests passed, automatic simulations ran, and the game was declared done. The first minute of actually playing it found five problems: keys that got lost, the game freezing, a single ghost, characters passing through each other, a ghost stuck oscillating. None was broken according to the tests, because none was testable.
What to delegate, and what not
Delegate the execution of bounded tasks. Keep the decomposition, the definition of done, the review of the diff and the end-to-end use test. The model is fast at the first; the other four are where a project succeeds or quietly fails.
Authors: José Luis Minich, Maximiliano Lucius.