Skip to main content

Known limits for coding agents

These are the places where an agent session on ColabHive behaves differently from a hosted proprietary API. Each one has a workaround.

A cold model makes the first request wait​

A model with no replica serving is loaded on the first request. For a 20–30B model that takes minutes, and an agent usually shows nothing while it waits.

  • Pick a warm model. colabhive agents init ranks warm models first, and colabhive agents models shows the state of each one.
  • Streamed requests are kept alive while the model loads, but not forever. The gateway waits up to 20 minutes, load and generation included, when model is an endpoint id or its name; when it is a catalog model name, it waits 1.25× that model's worst measured load time, and at least 2 minutes.
  • OpenCode aborts a request after 5 minutes by default. colabhive agents init sets "timeout": 1200000 on the provider so the first answer of a loading model is not cut off. If you write the configuration by hand, set it too.
  • A non-streamed request gets 503 with Retry-After just before the 300-second limit of the API edge while the model is still loading, or 504 if the model was ready but the answer takes longer — see OpenAI-compatible API.
  • Warm it before a long session: send one short request, wait for the answer, then start the agent.

Shared catalog models cannot be kept warm by you​

The platform loads and unloads the models of the public catalog as demand changes; you cannot pin one. Only the account that owns an endpoint can change its scaling policy: PATCH /endpoints/{endpoint_id}/scaling on a public catalog endpoint returns 404.

  • To keep a model warm, serve it from your account. Import a HuggingFace repository that is not already in the catalog — a repository already listed cannot be imported again under the same name. Imported models are published to the catalog and approved automatically. Then, with a key whose role is operator or higher, set the new endpoint to {"scaling_mode": "minimum", "min_replicas": 1, "max_replicas": 1} with PATCH /endpoints/{endpoint_id}/scaling.

Tool calling depends on the model​

/v1/models reports tool_calling_configured: true when a tool-call parser is declared for a model. That says the server can parse tool calls; it does not say the model uses tools well. The field that would say so, supports_tool_calls, is always null: the platform does not claim a capability from configuration.

  • Check before you rely on it. colabhive agents init and colabhive agents doctor send a real request with a tool and require a well-formed tool_call. The same check by hand is in OpenCode setup.
  • A model without a tool parser rejects the request. Sending tools with tool_choice: "auto" to it returns 400 with code: "engine_rejected_request"; retrying does not help, pick another model.
  • One call is not a long task. A model that passes the check can still lose track over dozens of steps; smaller models do so sooner. Keep tasks narrow, and judge each one with a command that passes or fails — see the method.
  • A model sometimes answers in prose instead of calling a tool. It is a sampled decision, so it varies between identical requests. The CLI's check asks twice before it gives up on a model; in an agent session, rephrase the step or ask again.

Prefer the model id to its name​

Both work as model in a chat request, but the name in /v1/models is a display label that can be renamed, and the id never changes. Use the id in configuration files. See OpenCode setup.

The context window is the one the replica runs​

max_model_len in /v1/models is the window the serving replica actually runs, which can be smaller than the model's native window when the board it runs on has less memory, and it changes when the model is loaded somewhere else. max_model_len_source is resident when that number comes from a serving replica, and configured when no serving window is known — normally because no replica is serving — in which case it is what the catalog declares and the replica that loads may run a smaller one. Configure the agent to compact before the published window, and re-read it after a cold start.

Prefix caching lives in each replica​

On models served by vLLM — most chat models — a prompt prefix the replica has already processed (the system prompt, files the agent has read) is served from its cache instead of being recomputed, and usage.prompt_tokens_details.cached_tokens says how many tokens were. A few models have it turned off. The cache belongs to one replica: a request routed to another replica of the same model starts without it.

Measured on 2026-09-21, the same prompt of 9,000–12,000 tokens sent twice:

ModelFirst requestSecond requestcached_tokens on the second
Qwen3.8-27B-FP810.0 s1.5 s11,760 of 12,381
gpt-oss-20b8.9 s0.7 s9,136 of 9,148

Keep the stable part of the prompt — system prompt, instructions, files — at the start, and put what changes at the end, so each request reuses the longest possible prefix.

Requests per minute​

Each API key has a request limit per minute set by the account's plan: 60 on Starter, 300 on Pro and 1,000 on Team. Every response carries X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset; past the limit the API answers 429 with code: "rate_limited" and a Retry-After header. Several agent sessions sharing one key share its limit.

Request size​

A request body can be up to 11.9 MiB of JSON. An agent that pastes whole directories into one request hits that before it hits most context windows. See OpenAI-compatible API.