Compatibility and limitations

Core: Python 3.9+; no Torch. Locally verified Python 3.9.25 with NumPy 2.0.2 / Pandas 2.3.3. CI also tests Python 3.14 core.

Model paths: Python 3.10+, single CPU FP32 or single CUDA. Two actual offline training/loading combinations are tested:

Combination Torch Transformers PEFT Accelerate
Minimum 2.6.0 4.57.6 0.18.1 1.12.0
Current qualification 2.14.1 CPU 5.18.0 0.21.2 1.15.0

Local CPU tests used Python 3.12.13 on Windows; CI tests minimum on Linux Python 3.10 and current on Linux Python 3.12. Real public training/inference uses Windows RTX 4080, Torch 2.6.0+cu124 and the minimum combination. Training is BF16; the corrected merged export and recommended inference are FP32. Version ranges express API compatibility intent, not qualification of every version/device/model Cartesian product. Qwen2.5-0.5B and an offline tiny GPT2 are exercised. Other compatible classifier families need their own measured qualification; remote code, 8-bit loading, MPS and multiple devices are outside the current scope.

The small example is uncalibrated, has no external benchmark guarantee, and cannot reliably verify truth, safety or human values. Long text loses context under its 512-token packer. Some rounds are omitted; even retained content may be insufficient. Tie labels in Arena pool both preference ties and “both bad” cases. The model can favor verbosity, style, or familiar domains. Use its probabilities for experiments and diagnostics alongside human review.

The public Arena dataset may overlap previous experiments or pretraining. The newly frozen test is held out from this delivery's training/selection; it is not claimed to be an uncontaminated external benchmark. Two seeds expose limited training variability, while grouped bootstrap describes uncertainty from this particular test sample. Neither establishes cross-domain robustness.

The original competition ensemble, larger models, pseudo-label recipe and award remain historical provenance; the released example does not reproduce their scale or quality. No new distillation, packing-quality or calibration gain is claimed without an ablation.