Known limits for coding agents
These are the places where an agent session on ColabHive behaves differently from a hosted proprietary API. Each one has a workaround.
A cold model makes the first request wait
A model with no replica serving is loaded on the first request. For a 20–30B model that takes minutes, and an agent usually shows nothing while it waits.
- Pick a warm model.
colabhive agents initranks warm models first, andcolabhive agents modelsshows the state of each one. - Streamed requests are kept alive while the model loads, but not forever. The gateway waits up to
20 minutes, load and generation included, when
modelis an endpointidor its name; when it is a catalog model name, it waits 1.25× that model's worst measured load time, and at least 2 minutes. - OpenCode aborts a request after 5 minutes by default.
colabhive agents initsets"timeout": 1200000on the provider so the first answer of a loading model is not cut off. If you write the configuration by hand, set it too. - A non-streamed request gets
503withRetry-Afterjust before the 300-second limit of the API edge while the model is still loading, or504if the model was ready but the answer takes longer — see OpenAI-compatible API. - Warm it before a long session: send one short request, wait for the answer, then start the agent.
Shared catalog models cannot be kept warm by you
The platform loads and unloads the models of the public catalog as demand changes; you cannot pin
one. Only the account that owns an endpoint can change its scaling policy:
PATCH /endpoints/{endpoint_id}/scaling on a public catalog endpoint returns 404.
- To keep a model warm, serve it from your account. Import a
HuggingFace repository that is not already in the catalog — a repository already listed cannot be
imported again under the same name. Imported models are published to the catalog and approved
automatically. Then, with a key whose role is operator or higher, set the new endpoint to
{"scaling_mode": "minimum", "min_replicas": 1, "max_replicas": 1}withPATCH /endpoints/{endpoint_id}/scaling.
Tool calling depends on the model
/v1/models reports tool_calling_configured: true when a tool-call parser is declared for a model.
That says the server can parse tool calls; it does not say the model uses tools well. The field that
would say so, supports_tool_calls, is always null: the platform does not claim a capability from
configuration.
- Check before you rely on it.
colabhive agents initandcolabhive agents doctorsend a real request with a tool and require a well-formedtool_call. The same check by hand is in OpenCode setup. - A model without a tool parser rejects the request. Sending
toolswithtool_choice: "auto"to it returns400withcode: "engine_rejected_request"; retrying does not help, pick another model. - One call is not a long task. A model that passes the check can still lose track over dozens of steps; smaller models do so sooner. Keep tasks narrow, and judge each one with a command that passes or fails — see the method.
- A model sometimes answers in prose instead of calling a tool. It is a sampled decision, so it varies between identical requests. The CLI's check asks twice before it gives up on a model; in an agent session, rephrase the step or ask again.
Prefer the model id to its name
Both work as model in a chat request, but the name in /v1/models is a display label that can be
renamed, and the id never changes. Use the id in configuration files. See
OpenCode setup.
The context window is the one the replica runs
max_model_len in /v1/models is the window the serving replica actually runs, which can be smaller
than the model's native window when the board it runs on has less memory, and it changes when the
model is loaded somewhere else. max_model_len_source is resident when that number comes from a
serving replica, and configured when no serving window is known — normally because no replica is
serving — in which case it is what the catalog declares and the replica that loads may run a smaller
one. Configure the agent to compact before the published window, and re-read it after a cold start.
Prefix caching lives in each replica
On models served by vLLM — most chat models — a prompt prefix the replica has already processed (the
system prompt, files the agent has read) is served from its cache instead of being recomputed, and
usage.prompt_tokens_details.cached_tokens says how many tokens were. A few models have it turned
off. The cache belongs to one replica: a request routed to another replica of the same model starts
without it.
Measured on 2026-09-21, the same prompt of 9,000–12,000 tokens sent twice:
| Model | First request | Second request | cached_tokens on the second |
|---|---|---|---|
| Qwen3.8-27B-FP8 | 10.0 s | 1.5 s | 11,760 of 12,381 |
| gpt-oss-20b | 8.9 s | 0.7 s | 9,136 of 9,148 |
Keep the stable part of the prompt — system prompt, instructions, files — at the start, and put what changes at the end, so each request reuses the longest possible prefix.
Requests per minute
Each API key has a request limit per minute set by the account's plan: 60 on Starter, 300 on Pro and
1,000 on Team. Every response carries X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset; past the
limit the API answers 429 with code: "rate_limited" and a Retry-After header. Several agent
sessions sharing one key share its limit.
Request size
A request body can be up to 11.9 MiB of JSON. An agent that pastes whole directories into one request hits that before it hits most context windows. See OpenAI-compatible API.