Coding Agents
A coding agent — OpenCode, SuperClaw (which is built on OpenCode), or any
client that speaks the OpenAI Chat Completions API with tools — can use ColabHive as its model
provider. The agent keeps running on your machine and executes its tools there; ColabHive serves the
model behind /v1/chat/completions.
Start here
| Page | For |
|---|---|
| Quickstart | Install the CLI and run colabhive agents init: it picks a model, checks that it really returns tool calls, and writes the agent's configuration |
| OpenCode setup, by hand | The provider block field by field, when you manage the configuration yourself |
| Choosing a model team | Which model for each agent role — planner, builder, reviewer, small tasks — measured |
| Running agents in parallel | How many sessions one replica serves at once, measured, and how to size a batch of tasks |
| Known limits | Cold starts and timeouts, tool calling that varies by model, context windows, request limits |
| Example: a game built by coding agents | A whole project run with the method, public repository included: the plan, the gate, what failed and what changed |
Why run an agent on ColabHive
Where the model runs. Source code, financial data and customer records are what coding agents are
most useful on, and what many teams cannot send to a third-party model API. On the hosted service
the model is served from capacity ColabHive operates. With a
Private Cluster it can run only on nodes you
enrolled: set the account's node-eligibility policy to own_hardware_only and placement is restricted
to those nodes — see the Data Protection API. The account's policy is
resolved on the server for every request, including /v1/chat/completions, so the agent needs no
change to respect it. An endpoint's own policy applies when model is that endpoint's id or name. What is and is not a published commitment is listed in
Trust & Operations.
Batch work. The typical fit is many small, well-defined tasks — a migration applied file by file, tests written module by module, a review pass over a set of changes — each one judged by a command that passes or fails. Open models of 20–30B parameters handle a narrow task with a clear check well; they do not hold a long, open-ended goal the way the largest proprietary models do. Split the work so each task fits in one sitting of the agent. How to split it, and how to judge each piece, is in A method for agents on open models.
What the platform gives an agent today
- Tool calling through the standard
tools/tool_callsfields — see Tool Calling. - The context window the replica actually runs, published per model in
/v1/modelsasmax_model_len, so the agent compacts against a real number. - Prefix caching on vLLM-served models: a repeated prompt prefix (system prompt, files already
read) is served from the replica's cache, and
usage.prompt_tokens_details.cached_tokenssays how many tokens were. - Warm state in the model list:
max_model_len_source: "resident"means a replica is serving that model right now.
Terms used in this section
| Term | Meaning |
|---|---|
| Endpoint | A model you can call. Its id is what goes in model |
| Replica | One running copy of a model on a GPU node. An endpoint can have several |
| Warm / cold | Warm: a replica is serving now and answers in seconds. Cold: none is, and the first request waits for a load |
| Node-eligibility policy | Which nodes may run your work: any node, or only hardware your account enrolled (own_hardware_only) |