Skip to main content

Choosing a model team

OpenCode lets each agent role use its own model. The useful question for picking them is not a benchmark score but a behavior: does the model act, or does it deliberate? A builder has to call tools and move on; a reviewer earns its keep by thinking before it answers. The same model can be the right choice for one role and the wrong one for another.

Measured on 2026-09-21​

Each model got the same tool-calling request ten times — the check colabhive agents init runs.

ModelTool calls (10 tries)Median time to the tool callBehavior
gpt-oss-20b10/100.7 sacts
Qwen3 Coder 30B-A3B AWQ10/101.1 sacts
Qwen3-8B10/106.6 sdeliberates
Qwen3.8-27B-FP810/109.9 sdeliberates

All four call tools reliably; what separates them is how long they think first. Context windows are not in this table on purpose: the window is the one each replica runs and it changes when a model is loaded on different hardware — read it from colabhive agents models or /v1/models.

A team that works​

RoleModelWhy
build (writes code)Qwen3 Coder 30B-A3B AWQacts instead of deliberating, and calls tools fast
tester (writes and runs tests)Qwen3 Coder 30B-A3B AWQthe same executor
plan (splits the work)gpt-oss-20breasons briefly and answers fast
reviewer (reads a diff, finds bugs)Qwen3.8-27B-FP8here the deliberation is the point
title, summary, exploregpt-oss-20bthe fastest to answer; a model that deliberates spends seconds on a title

In OpenCode, with the model ids from the catalog:

{
"model": "colabhive/5d21e32a-3bbb-4040-9c34-3b06c4415b84",
"agent": {
"plan": { "model": "colabhive/5d21e32a-3bbb-4040-9c34-3b06c4415b84", "temperature": 0.3 },
"build": { "model": "colabhive/f5d76140-5b4d-41c0-88da-dad6d11f341a", "temperature": 0.2,
"steps": 80 },
"tester": { "mode": "subagent", "model": "colabhive/f5d76140-5b4d-41c0-88da-dad6d11f341a",
"temperature": 0.2, "description": "Writes and runs tests for a change." },
"reviewer": { "mode": "subagent", "model": "colabhive/1af07b1f-5832-4451-a61c-76d1fe43115a",
"temperature": 0.1, "description": "Reviews a diff looking for bugs. Read-only.",
"permission": { "edit": "deny", "bash": "ask" } },
"title": { "model": "colabhive/5d21e32a-3bbb-4040-9c34-3b06c4415b84" },
"summary": { "model": "colabhive/5d21e32a-3bbb-4040-9c34-3b06c4415b84" },
"explore": { "model": "colabhive/5d21e32a-3bbb-4040-9c34-3b06c4415b84", "temperature": 0.1 }
}
}

A complete configuration built around this team — provider, agents, prompts and a warm-up script — is in the repository of Example: a game built by coding agents.

Every model referenced here also needs its entry under provider.colabhive.models — see OpenCode setup. colabhive agents init writes those entries for the models it picks; add the rest by id.

  • Deny edit to the reviewer. A reviewer that can edit "fixes" what it should only report.
  • steps caps the builder's iterations. Raise it if the agent runs short; if it runs out of steps while deliberating, the model is the problem, not the cap.
  • Write description for the model, not for people. It is what the main agent reads to decide which subagent to call.
  • Check the ids. These are the public catalog models as of the date above; colabhive agents models --json lists the ones your account can use, and whether each one is warm right now.

A scorer next to the team: ranking and verifying actions​

Not every role needs a model that writes. When the agent already has the candidates — the next commands it could run, the patches the builder proposed, the plans the planner drafted — a scorer can rank them or check each against the current state, and return a probability instead of an opinion to parse. CLM-8B does that: POST /v1/rank orders free-form candidates for a context and a question, and POST /v1/systemone answers yes/no, choice and rubric questions about a state.

from colabhive import ColabHive
from colabhive.clm import Noul

client = ColabHive(api_key="hive_...", account_id="YOUR_ACCOUNT_ID")

state = "Task: make test_parser pass.\nLast output: 1 failed, KeyError 'lang' in parse_header()"
ranked = client.clm.rank(
model="clm-v0.1-8b",
context=state,
question="Which next step makes progress on the task?",
answers=["Read parse_header() in parser.py", "Rewrite the test to skip 'lang'", "Run the whole suite again"],
)
check = client.clm.system_one(
model="clm-v0.1-8b",
state=state,
questions={"safe": Noul(instructions="Is 'git push --force origin main' an appropriate next step?")},
)
print(ranked.ranked[0].candidate, check.answers["safe"].noul)
  • It is not an OpenCode provider model: it is not in /v1/models and does not chat. Call it from a tool, a hook or your orchestration script, with the candidates the agent produced.
  • The reference head was not trained on your agents' trajectories. Record states and the actions that worked, and train a head with task clm — see Fine-tune a CLM head — before you let its score gate an action.
  • It reads English only, and score questions are not reliable zero-shot: see the model page's known limitations.