Use a trained judge

Install python -m pip install 'pairjudge[judge]==0.3.0' in a Python 3.10+ environment. For CPU-only Torch, install from the official CPU wheel index first. The first run downloads about 2 GB; it does not upload your input.

pairjudge compare --prompt "Why does ice float?" --a "Its crystal structure lowers its density." --b "Cold objects always rise." --swap --device cpu
pairjudge batch --input pairs.jsonl --output results.jsonl --swap
pairjudge demo --port 7860

The built-in model is a small example with measured limitations. An omitted --revision pins its published commit automatically. Other Hub models are resolved to a single snapshot before tokenizer/weights load. Local directories must be trained bundles. A plain backbone is intentionally rejected.

JSONL keys are id, prompt, response_a, response_b. Pair columns may be strings for a single turn, or equal-length lists of strings for multiple turns. List entries cannot be null. Each response is the candidate's answer to the corresponding shared prompt; histories must already be aligned. Flat strings in the low-level packer/DataFrame API are invalid; the CLI explicitly converts single-turn strings to lists. Duplicate IDs/indexes are allowed and retain positional order. Empty files produce no predictions.

Results contain canonical a_wins, b_wins, tie probabilities; verdict; model commit/device/dtype; and per-round original/kept content-token counts. ambiguous means two numerical maxima within absolute tolerance 1e-8, not a learned preference tie. No confidence threshold or calibration is implied.

--swap computes both orders and averages after exchanging A/B columns. It provides swap-equivariant distributions, with additional measured cost. To inspect each direction separately, use the local demo or the Python API's single-pass prediction on each input order.

from pairjudge import PairwiseJudge, from_pairs
from pairjudge.catalog import DEFAULT_MODEL, DEFAULT_REVISION
judge = PairwiseJudge.from_pretrained(DEFAULT_MODEL, revision=DEFAULT_REVISION,
                                     device="cpu", local_files_only=True)
data = from_pairs(["Explain gravity."], ["Mass attracts mass."], ["It is magic."])
print(judge.predict(data, swap_debias=True))

After the initial download, --offline --cache-dir PATH works from that same cache. An empty cache in offline mode fails clearly. CPU supports FP32; CUDA supports explicit FP32/FP16/BF16 with hardware checks. Auto uses one GPU if available and respects the bundle's saved default precision. Quantization, remote code and multi-device maps are unqualified and rejected. Adapter inference additionally needs pairjudge[train] for PEFT; the public merged example needs only [judge].

Both-blank answers and a budget unable to retain response content fail. One-blank, identical answers and missing prompt context are diagnosed; no fixed probability override is applied. CLI/UI cap pairs at 60,000 characters and 64 rounds, HTTP bodies at 180,000 bytes, and CLI batches at 10,000 rows. The local demo accepts one inference request at a time (concurrent requests receive HTTP 429). It is loopback-only and has no remote fonts, scripts, telemetry or conversation logging. It is a development demo, not a hosted production service. The documentation site runs no model inference.