Troubleshooting#
Find the symptom, run the check, apply the fix.
Installation#
error: Failed to spawn: pimmuv runwas called outside the checkout. Run from the repository root or useuv run --project /path/to/particle-imaging-models.ModuleNotFoundErrorforspconv,pointopsor another native packageRun
uname -srm(training needs Linux x86-64),uv lock --check, thenuv sync --locked.torch.cuda.is_available()isFalseCheck
nvidia-smiand that the driver supports CUDA 12.6. Containers need--nv(Apptainer) or--gpus all(Docker).
Data#
PILArNet data root not foundSet
data_rootin the config, exportPILARNET_DATA_ROOT_V3, or download into~/.cache/pimm/pilarnet/v3. On a cluster, put the variable in the site profile’senv:so compute nodes see it.PILArNet revision='v1' is not supportedThe recipe names revision
v1. Override each split tov3, for exampledata.train.revision=v3 data.val.revision=v3.Key ... not foundin a transformThe reader didn’t emit the key, or an earlier transform removed or renamed it. Apply the transforms one at a time and print the keys after each.
- Point fields have different lengths
A per-point key is missing from
index_valid_keys, so subsampling skipped it. See Transforms.
Launcher and Slurm#
- A launcher flag is rejected
Launcher flags are dotted, such as
--resources.nproc-per-node. Training values go after--askey=value.uv run pimm launch --helplists the flags.- The job asked for the wrong account, partition or GPUs
Run the same command with
--dry-runand read the script. Check the site profile; without--site,pimm submitusess3df.- A multi-node job hangs at NCCL initialization
Run one GPU on one node with the same image and data first, then check the site’s network-interface settings and the rendezvous values in the dry run.
Training#
batch_sizeassertion failsBatch sizes are global and must divide by the number of GPUs.
- CUDA runs out of memory after many steps
Events differ in size, so one large event can overflow a batch. Log the total points per batch and the GPU memory to find it.
- Loss is NaN
On a few events, log the loss components and gradient norm (
GradientNormLogger), then rerun withenable_amp=False.- No model parameters loaded
The checkpoint’s keys don’t match the model. Check the weight path, prefix and key mapping in the startup report.
- The wrong parameters are trained
Add
ParameterCounterand print the names withrequires_grad=Truebefore the first step.
Checkpoints and evaluation#
Incomplete checkpoint directoryThe checkpoint has no
.complete. Useuv run python -m pimm.utils.path latest-checkpoint <run>/modelto find the newest complete one.- The saved epoch restarted on resume
The number of GPUs or workers changed; see Checkpoints and resume.
model_best.pthis never writtenevaluateisFalse, there is no validation split, orCheckpointSavercomes before the evaluator inhooks.
Reporting a problem#
Open an issue with the pimm commit, the exact command, the resolved config, the first traceback, and your OS, GPU, driver and CUDA versions.