Train and distill
Use pairjudge[train] with Python 3.10+. Only one CPU or one CUDA device is
qualified. Start with a small local dataset and a fresh output directory.
python -m pairjudge.training --cfg examples/configs/quickstart.yaml
CSV uses Arena JSON-encoded lists; JSONL/parquet use actual lists. All three
winner columns are required. Hard labels are strict one-hot; soft labels must
be finite, nonnegative and sum to one. Default loaders do not relabel human
preferences or remove rows. Nulls/misaligned rounds are rejected. Public
experiment exclusions are separately audited in its frozen manifest.
Provide an exact 40-character model_revision for Hub backbones. Configure
max_length, packer (all other PackerConfig fields), and label_order.
device: cpu requires dtype: float32 or auto; bf16: true is an explicit
compatibility alias for a BF16 request and is rejected on CPU. bf16: false
no longer means FP16. dtype: auto selects CPU FP32 or supported CUDA
BF16/FP16. gradient_checkpointing can bound CUDA activations.
Use a separate validation_path. Training rejects shared normalized prompts
or unordered response pairs. Without it, connected-group holdout uses
eval_holdout and seed; this is validation, not an independent final
test. Keep test and pseudo-label pools separate before training.
The explicit PyTorch loop sums CE or KL over each microbatch and divides by
the actual example count in the whole accumulation group. Its last smaller
group is normalized by its own count. AdamW, cosine schedule, warmup and
gradient norm clipping are explicit; FP16 overflows/nonfinite losses fail
without exporting success. No implicit distributed Trainer contract is used.
Validation runs at each completed/partial epoch; its lowest log-loss selects
the export. LoRA saves the trained classification head alongside adapters;
validation happens before merge changes the model object.
Merged exports move to CPU FP32 before applying adapter deltas. Qualification
must compare the reloaded adapter and merged model: BF16 merging changed the
public pilot's probabilities materially, whereas FP32 merging passed the
declared tolerance. This increases public weight storage to about 2 GB.
Two phases
First train a hard-label teacher. Then label a separate, authorized pool:
python -m pairjudge.pseudo_label --model output/judge-quickstart/merged --data pool.parquet --out pool_pl.parquet --swap-debias
The output keeps full canonical probabilities, with teacher revision/artifact
hash, packer, precision and inference mode in parquet metadata and a JSON
sidecar. These are teacher labels, not independent human annotation.
For the student set label_mode: soft, retain human one-hot rows as valid soft
distributions, and add pseudo-labeled training rows only. Keep validation/test
human-only when reporting human preference quality. Record pool and teacher
provenance in the training config. tests/test_user_paths.py actually executes
teacher → pseudo-label → soft student using tiny offline models; it establishes
the pipeline, not a distillation quality improvement.
The public model experiment is reproducible via examples/public_judge.py:
freeze once, inspect manifest, reconstruct private text files from exact
dataset revision and recorded IDs, train using the published seed configs,
select on validation, then evaluate final test once. The frozen report contains
no conversation text. See evaluation.