Troubleshooting¶
Missing dependencies or CUDA¶
Core estimate/recipe/demo needs no torch. Actual training needs the qualified
quickstart stack. Use a fresh project environment if unrelated
packages cause import failures. CPU uses explicit --device cpu --base-dtype fp32
--optimizer adamw_torch and full/LoRA. QLoRA and 8-bit optimizers require CUDA.
nvidia-smi working does not establish that the installed torch wheel has CUDA;
check canifinetune doctor and torch.cuda.is_available(). Do not replace global
CUDA or other project environments to fix this package.
The verified combination is torch 2.6.0/cu124 and bitsandbytes 0.49.2. Use the
provided constraints instead of an unbounded upgrade. Windows console logs with
non-ASCII text should use $env:PYTHONUTF8="1" in PowerShell. Linux uses
export PYTHONUTF8=1. CI sets UTF-8 explicitly.
Dataset or configuration rejected¶
Read the line number and generated dataset_format.md. Preserve system/multi-turn
records; unsupported roles or templates fail explicitly. Assistant-only chat needs
native generation masks. Empty replies and no shifted supervised labels fail.
Increase sequence length, shorten input deliberately, or explicitly opt into
truncation: right; right truncation still cannot remove all response labels.
Unknown config keys fail. Regenerate old recipes rather than copying legacy
save_steps/TRL settings into the strict schema. Output must be empty; choose a
fresh directory. --force overwrites generated recipe files only, never a run.
OOM or a busy display GPU¶
Reduce micro-batch/sequence length, use attention-only targets, QLoRA or a smaller
model, and re-estimate the changed configuration. Accumulation preserves effective
batch after reducing micro-batch; it does not shrink a single micro-batch. A load
OOM will not be fixed by checkpointing. Inspect run.json/bench stage locally.
An OOM receipt has no exact peak and cannot be evaluated as successful.
Keep actual total capacity in --gpu-vram-gb and pass currently free capacity
separately with --available-vram-gb. Background allocations change; don't compare
device-level usage directly with process allocator peaks. Safety is an additional
planning allowance, not part of observed reserved memory. A short probe cannot
guarantee a long run will fit. Do not terminate unrelated Python processes.
BF16, attention, Liger¶
Unsupported BF16 is an explicit error. Choose fp16 for a supported adapter path or fp32, then re-estimate. Full fp16 weight training is rejected. Use SDPA/eager explicitly if Flash Attention is unavailable; no loader retry removes attention. Flash Attention and Linux/Triton Liger have not received 0.4.0 CUDA qualification. Liger's estimate is a low-confidence stock upper planning proxy. Bench rejects it.
Offline and gated models¶
--offline uses curated metadata, local directories or cached config; unknown
uncached metadata fails clearly. Training additionally needs cached weights and
tokenizers. Curated metadata does not establish access to gated weights. Obtain
access under the model's terms yourself and authenticate using the normal Hub
workflow. Never put credentials in datasets, YAML, issues or logs for sharing.
Remote model code is disabled; models requiring it are unsupported.
Reload or save failure¶
Check the run is successful and contains its model/adapter artifact. Failed saves
remove .saving; they do not become successful checkpoints. Adapter evaluation
reloads the same pinned base and applies the saved adapter; it never silently
ignores adapter errors. These are inference artifacts, not a full resume snapshot.
For a single unambiguously matched GPU, discovery retains both CUDA-driver and nvidia-smi free-memory readings and uses the lower one for a planning budget. These readings disagreed on the qualification host during concurrent GPU pressure; that observation does not prove its underlying cause. Stage snapshots label their CUDA-driver source. A conservative reading still cannot reserve future memory.