Skip to main content

Coding Agents

A coding agent — OpenCode, SuperClaw (which is built on OpenCode), or any client that speaks the OpenAI Chat Completions API with tools — can use ColabHive as its model provider. The agent keeps running on your machine and executes its tools there; ColabHive serves the model behind /v1/chat/completions.

Start here​

PageFor
QuickstartInstall the CLI and run colabhive agents init: it picks a model, checks that it really returns tool calls, and writes the agent's configuration
OpenCode setup, by handThe provider block field by field, when you manage the configuration yourself
Choosing a model teamWhich model for each agent role — planner, builder, reviewer, small tasks — measured
Running agents in parallelHow many sessions one replica serves at once, measured, and how to size a batch of tasks
Known limitsCold starts and timeouts, tool calling that varies by model, context windows, request limits
Example: a game built by coding agentsA whole project run with the method, public repository included: the plan, the gate, what failed and what changed

Why run an agent on ColabHive​

Where the model runs. Source code, financial data and customer records are what coding agents are most useful on, and what many teams cannot send to a third-party model API. On the hosted service the model is served from capacity ColabHive operates. With a Private Cluster it can run only on nodes you enrolled: set the account's node-eligibility policy to own_hardware_only and placement is restricted to those nodes — see the Data Protection API. The account's policy is resolved on the server for every request, including /v1/chat/completions, so the agent needs no change to respect it. An endpoint's own policy applies when model is that endpoint's id or name. What is and is not a published commitment is listed in Trust & Operations.

Batch work. The typical fit is many small, well-defined tasks — a migration applied file by file, tests written module by module, a review pass over a set of changes — each one judged by a command that passes or fails. Open models of 20–30B parameters handle a narrow task with a clear check well; they do not hold a long, open-ended goal the way the largest proprietary models do. Split the work so each task fits in one sitting of the agent. How to split it, and how to judge each piece, is in A method for agents on open models.

What the platform gives an agent today​

  • Tool calling through the standard tools / tool_calls fields — see Tool Calling.
  • The context window the replica actually runs, published per model in /v1/models as max_model_len, so the agent compacts against a real number.
  • Prefix caching on vLLM-served models: a repeated prompt prefix (system prompt, files already read) is served from the replica's cache, and usage.prompt_tokens_details.cached_tokens says how many tokens were.
  • Warm state in the model list: max_model_len_source: "resident" means a replica is serving that model right now.

Terms used in this section​

TermMeaning
EndpointA model you can call. Its id is what goes in model
ReplicaOne running copy of a model on a GPU node. An endpoint can have several
Warm / coldWarm: a replica is serving now and answers in seconds. Cold: none is, and the first request waits for a load
Node-eligibility policyWhich nodes may run your work: any node, or only hardware your account enrolled (own_hardware_only)