Scale your training job across multiple GPU nodes

Converts a single-machine training script into a multi-node Slurm job, handling distributed PyTorch setup, NCCL config, and the specific OOM and timeout patterns that break at scale.

Best for: ML engineers moving from a single GPU to a cluster without rewriting the whole pipeline.

Engineering / pipelines-dataatomicfor-engineerslight-setupfrom-file

Source

Creator's repository · nvidia/skills

View on GitHub

License: Apache-2.0

Security

Verified — safe to install
Passed all 3 independent security checks
Checked by 3 independent security firms
Does it try to trick the AI?NoSAFE · Gen Agent Trust Hub
Does it sneak in hidden code?NoNo alerts · Socket
Does it have known bugs?NoLow risk · Snyk