Skip to main content

Fine-tune a CLM head

CLM-8B scores candidates with a frozen Qwen3-8B encoder and two small heads. The reference head is general; a head trained on your own labeled data is what makes the scores useful in your domain. Training changes only the heads — the encoder is never modified — so a trained head is a file of about 75 MB, and serving it loads nothing new: it rides with each request to the same clm-v0.1-8b replicas.

Template: clm-head-infonce (framework clm). One job, on one GPU, in two phases:

  1. Embed. The encoder reads every text in the dataset once and stores the vectors on the node. This phase is most of the time and is resumable: a job re-dispatched to the same node shortly after an interruption continues from the last completed shard of 10,000 rows.
  2. Train the heads. The encoder is released and the two heads are trained on the stored vectors, with early stopping on a validation split.

The job writes artifacts/clm_head.pt and artifacts/clm_head_manifest.json. The manifest binds the head to the exact encoder it was trained on (repository, revision, dtype, pooling, tokenizer and chat template); a head is refused by any other encoder, so it can never score silently on the wrong one.

Choose the task​

TaskYour dataWhat the head learnsLongest text
clm (default)trajectories: a state and the action taken in itto score the action that fits a state above the others (contrastive, InfoNCE)8,192 tokens
choicemultiple-choice questions in the /v1/systemone format, with the right answerto answer noul, choice and score questions like yours2,048 tokens

Use choice when you will call the head with fixed questions (routing, triage, rubric scoring). Use clm when the candidates are open-ended actions, such as the next step of an agent.

Dataset formats​

Upload one JSONL file (one JSON object per line) as a dataset — see Preparing Datasets and the Datasets API. A dataset with more than one .jsonl file is refused.

Task clm​

{"state": "Ticket: VPN drops every 10 minutes since the update.", "action": "Ask for the client version and the OS", "task_id": "t-1841", "step_idx": 0}
{"state": "Ticket: VPN drops every 10 minutes since the update.\nUser: 5.2.1 on macOS 15", "action": "Link the known-issue article for 5.2.1 and offer the 5.2.2 build", "task_id": "t-1841", "step_idx": 1}
{"state": [{"role": "user", "content": "Refund my last order"}], "action": "Look up the order and check the refund window", "task_id": "t-2207", "step_idx": 0}
FieldRequiredMeaning
stateyesThe context: text, or a list of chat messages (role + content). A head trained on chat messages expects chat messages at serving time.
actionyesThe action taken in that state: a non-empty string.
task_idyesGroups the steps of one task. The validation split is by task_id, so at least two distinct values are needed.
step_idxyesThe step within the task. Rows with the same task_id and step_idx are never used as negatives of each other.
trajectory_id, rewardnoKept with the row; not used by the loss.

Task choice​

Each row is one /v1/systemone request plus its answers. gold maps each question id to its right label, and optionally to a full probabilities distribution (soft targets).

{"id": "q-001", "state": "My card was charged twice for the same order.", "questions": {"intent": {"type": "choice", "instructions": "What does the customer need?", "criteria": {"refund": "Money back for a charge", "card_lost": "Block a lost or stolen card", "transfer": "Send money to someone"}}}, "gold": {"intent": {"label": "refund"}}}
{"id": "q-002", "state": "I can't find my card anywhere since yesterday.", "questions": {"intent": {"type": "choice", "instructions": "What does the customer need?", "criteria": {"refund": "Money back for a charge", "card_lost": "Block a lost or stolen card", "transfer": "Send money to someone"}}}, "gold": {"intent": {"label": "card_lost", "probabilities": {"card_lost": 0.9, "refund": 0.1}}}}
FieldRequiredMeaning
idyesRow id. Rows are split into train, validation and test by id, so a row never appears in two splits; at least three distinct ids are needed.
stateyesText, an object or an array (as in /v1/systemone). Chat messages are refused: choice heads read prose.
questionsyes{question_id: {type, instructions, criteria}}, exactly as in /v1/systemone.
goldyes{question_id: {label, probabilities?}}. label is an option key for choice, true/false for noul, and the level index ("0", "1", …) for score. A question without a gold label is not trained on.

state, questions and gold may also be JSON-encoded strings.

Hyperparameters​

Model ID: clm-head-infonce

This is the template's strict contract: unknown keys and out-of-range values are refused when the run is created. The same table is on the model page.

ParameterDefaultRange / optionsNotes
task"clm""clm", "choice"clm: state/action pairs (contrastive). choice: multiple-choice questions in the System One format.
loss"infonce""infonce", "softce"softce requires task choice.
targets"soft""soft", "hard"Task choice: soft keeps the gold distribution, hard its most probable label.
epochs201 to 200Maximum epochs of the head training; early stopping may end sooner.
patience51 to 200Epochs without a better validation score before training stops.
seed12340 to 2147483647Seed of the data split and of the head training.
batchnull8 to 8192null = task default: 2048 for clm, 256 for choice.
lrnullgreater than 0, at most 1.0null = task default: 2e-3 · √(1024/width) · √(batch/1024) for clm, 5e-4 for choice.
max_tokensnull16 to 8192Encoder tokens kept per text. null = task limit: 8192 for clm, 2048 for choice. A value above the task's limit fails before the GPU is used.
batch_tokensnull16 to 8192Padded tokens per encoder pass while the dataset is embedded. null = max_tokens.
max_total_tokens500000001 to 200000000Upper bound on the encoder tokens read from the dataset; a larger dataset fails before encoding starts.
holdout_frac0.10.01 to 0.49Share held out for validation (and, for choice, for test).

null means "the task's default": leave the key out, or send null, and the job picks the value for the task you chose.

Run it​

from colabhive import ColabHive

client = ColabHive(api_key="hive_...", account_id="YOUR_ACCOUNT_ID")

dataset = client.datasets.upload(name="support-intents", file="./intents.jsonl")
job = client.training.create(
model="clm-head-infonce",
dataset_id=dataset.id,
job_name="support-intents-head",
hyperparameters={"task": "choice", "loss": "softce", "targets": "soft"},
)
job.wait()
print(job.status, job.get_metrics())

Metrics reported while the job runs:

  • task clm: loss during training; val_loss, val_top1 and within_task_top1 (the selection metric) on the validation tasks;
  • task choice: acc and soft_ce on validation, and test_acc and test_soft_ce on the held-out test ids at the end.

Before the GPU is used, the job validates the whole dataset and the configuration and fails with a named error if something is wrong: dataset_invalid (with the line number), dataset_too_small, max_total_tokens_exceeded, choice_max_tokens_exceeded, choice_chat_state_unsupported. A non-finite loss fails the job instead of producing a head.

Serve the head​

Register the finished run as an endpoint, then pass that endpoint as model:

endpoint = client.training.register_for_inference(
run_id=job.id,
name="support-intents-head",
description="CLM head for support ticket intents",
visibility="account",
)

r = client.clm.system_one(
model=endpoint.endpoint_id, # or its name, "support-intents-head"
state="I lost my card on the train this morning.",
questions={"intent": {"type": "choice", "instructions": "What does the customer need?",
"criteria": {"refund": "Money back for a charge",
"card_lost": "Block a lost or stolen card",
"transfer": "Send money to someone"}}},
)
print(r.answers["intent"].choice, r.colabhive.head_sha256)

The request runs on a clm-v0.1-8b replica with your head; colabhive.head_sha256 in the answer names the head that scored it. The endpoint follows the account rules of any trained model: other accounts cannot call it unless you publish it. Never pass a runtime key containing :trained: as model: it is refused, because it is not bound to your account.

Start from a head you already trained​

Pass the earlier run (or its model version) as base. The new job starts from that head instead of the reference head:

from colabhive import ArtifactRef

job = client.training.create(
model="clm-head-infonce",
dataset_id=new_dataset.id,
base=ArtifactRef.job(previous_job.id), # or ArtifactRef.model_version(version_id)
hyperparameters={"task": "choice"},
)

base must be a trained CLM head over the same encoder: anything else is refused with 422 (clm_base_not_a_head, clm_base_encoder_mismatch) before a job exists. Without base, training starts from the reference head.

Hardware​

The job needs one GPU that holds the encoder in bf16 with room for its activations. It never splits the encoder across GPUs, and nothing below bf16 is used.

HardwareHead trainingReference run: choice on a fixed Banking77 subsetEncoder throughput while embeddingMeasured on
NVIDIA RTX 3090 24 GB{{MEASURED:clm.status.training.nvidia_rtx3090}}{{MEASURED:clm.training_time.banking77_choice.nvidia_rtx3090}}{{MEASURED:clm.encoder_tokens_per_s.nvidia_rtx3090}} tokens/s{{MEASURED:clm.evidence_date.training.nvidia_rtx3090}}
Intel Arc Pro B70 32 GB{{MEASURED:clm.status.training.intel_b70}}{{MEASURED:clm.training_time.banking77_choice.intel_b70}}{{MEASURED:clm.encoder_tokens_per_s.intel_b70}} tokens/s{{MEASURED:clm.evidence_date.training.intel_b70}}
AMD Instinct {{MEASURED:clm.amd.board}}{{MEASURED:clm.status.training.amd_instinct}}{{MEASURED:clm.training_time.banking77_choice.amd_instinct}}{{MEASURED:clm.encoder_tokens_per_s.amd_instinct}} tokens/s{{MEASURED:clm.evidence_date.training.amd_instinct}}
Intel, two 16 GB GPUsnot offered———
CPU{{MEASURED:clm.status.training.cpu}}{{MEASURED:clm.training_time.banking77_choice.cpu}}{{MEASURED:clm.encoder_tokens_per_s.cpu}} tokens/s{{MEASURED:clm.evidence_date.training.cpu}}

The reference run has {{MEASURED:clm.training.banking77_rows}} rows. The embedding phase scales with the tokens in your dataset: divide them by the throughput above for a first estimate. The job also reserves {{MEASURED:clm.training.cpu_cores}} CPU cores and {{MEASURED:clm.training.ram_gb}} GB of RAM on the node. These are measurements of the reference run, not guarantees.

Training on your own nodes is not charged; on capacity ColabHive operates it is billed per accelerator-hour — see Pricing.


Authors: José Luis Minich, Maximiliano Lucius.