Skip to content

0.4.0 implementation and qualification record

Baseline: 053ce9e115b213e0ded3634e00d564ef97561f9b (clean main). At the start, GitHub and the live PyPI JSON API both reported 0.3.0; its release workflow had succeeded. Main had no branch protection or repository rulesets; the maintainer has ADMIN permission. Work uses a branch and CI before integration.

Confirmed findings

Review lead Finding at the baseline
Two tokenization paths Confirmed: prepared labels are discarded when TRL imports; SFTTrainer reprocesses raw rows.
First user/assistant Confirmed in both formatters; system and later turns disappear in the legacy formatter, while modern TRL can prioritize messages. Actual behavior depends on TRL version.
BF16 fallback Confirmed: load falls back to fp16 but TrainingArguments does not enable fp16 for a bf16 request.
Broad TypeError retry Confirmed in bench, training and evaluation; drops attention and eventually dtype.
Training CI Confirmed: the 80 passing baseline tests do not execute a training recipe.
Calibration evidence Historical regression data informed the estimator; it is not independent validation.
Configuration divergence Confirmed: QLoRA bench overrides compute dtype; fused optimizer name does not actually enable fused AdamW.
Trusted publishing Already present and working; retain release.yml and its current no-environment OIDC binding.
Pre-quantized input Transformers 5.8.1's actual config merge retains checkpoint 4-bit settings despite an int8 request. Reject declared pre-quantized bases before weights and in metadata.
PyPI documentation links The published 0.3.0 relative changelog URL resolves within PyPI and returns 404. Main README now uses versioned absolute repository links.

Design decisions

Use one package-backed Transformers Trainer runtime, one validated configuration and one dataset processor. Default remains all-token loss. Assistant-only chat requires native generation masks; unsupported templates fail explicitly. No string searches or prompt-prefix guesses. Reject truncation by default; opt-in right truncation must retain shifted causal supervision and response tokens. Keep historical measurements unchanged and avoid fitting the new validation data.

Predeclared local validation cohort

Before observing new results, freeze a bounded cohort on the detected RTX 4080 (16 GiB): local tiny Qwen2 smoke (CPU full/LoRA, CUDA LoRA/QLoRA), then Qwen2.5-0.5B-Instruct at sequence 256 and 512, LoRA and NF4 double-quant QLoRA, batch 1, rank 8, attention targets, AdamW, checkpointing on, bf16, SDPA, stock loss, 3 updates. Maximum new weight download 1.5 GB; maximum one larger process at a time; stop on an unexpected OOM. Compare published 0.3.0 and candidate on the same observations. This is one GPU and one additional real model, not a cross-hardware or long-training guarantee. Recheck free memory before each run.

Baseline verification: python -m pytest -q: 80 passed. Windows Python 3.12, torch 2.6.0+cu124, Transformers 5.8.1, PEFT 0.19.1, Accelerate 1.13.0, bitsandbytes 0.49.2. nvidia-smi and torch report changing free memory; no other processes will be stopped. Disk free space exceeds 500 GB.

Implemented and verified locally

One Transformers runtime replaces the data-path split; no silent dtype/attention retry. Native masks and offset boundaries preserve system/multi-turn/all-token semantics. Strict configuration, staged saving and real reload are shared by installed recipes. A loopback stdlib demo and explicit redacted export are shipped. Calibration schema 3 separates compatible fits from prospective/community evidence. The original 70% core coverage gate is retained; optional torch runtime has actual CPU integration and candidate CUDA qualification rather than fabricated core coverage.

No-torch Python 3.12 core: 102 tests passed (71.44% coverage), ruff/format/mypy/generated Python checks passed. Minimum and recommended CPU stacks each passed 14 real integration tests. Four bounded prospective GPU observations succeeded; predictions are unchanged from 0.3.0 and conservatively high (MAPE 46.2%). No claim of accuracy improvement or cross-hardware validation. See validation-0.4.0.md and its raw files. A resource-pressure repeat was interrupted and excluded, without stopping unrelated processes. Historical raw files are intact.

The actual 0.3.0 tiny recipe completed one update after a GBK UnicodeDecodeError prevented TRL import and selected its legacy fallback. This is version/environment specific; it does not establish that modern TRL always fails. The statically confirmed divergent processing paths were removed regardless of that import result.

The pre-quantization config merge probe loaded no weights and allocated no GPU model. It is evidence of configuration precedence, not int8/4-bit accuracy or physical memory qualification. A local config-only checkpoint regression proves the new guard runs before weight lookup. Doctor installation guidance now matches the qualified torch/bitsandbytes constraints.

Delivery status

Completed: implementation, installed candidate qualification, annotated v0.4.0, GitHub Release, official PyPI 0.4.0 and public-install verification. Release commit is 2edf707ac2040a2dc7a55dfe26dfa9f427b8fc45. The release workflow passed all three jobs; public wheel/sdist hashes exactly match the tested files. Fresh official-index installs outside the checkout passed torch-free core and real CPU full/LoRA plus CUDA LoRA/QLoRA update/save/reload qualification. The pinned 0.5B quickstart also passed two QLoRA updates and adapter reload/generation. See release verification for commands, identities, receipts, compatibility, attribution and remaining unqualified scope. No external publication blocker remains. The shipped demo is local; no hosted service is claimed.

GPU-pressure diagnosis retained disagreeing CUDA-driver/nvidia-smi readings; single-device budgets now use the lower reading. The interrupted extra repeat revealed this safeguard gap; it was not counted as an OOM or successful sample.