nvidia / skills / nemo-mbridge-multi-node-slurm Scale your training job across multiple GPU nodes Converts a single-machine training script into a multi-node Slurm job, handling distributed PyTorch setup, NCCL config, and the specific OOM and timeout patterns that break at scale.
Best for: ML engineers moving from a single GPU to a cluster without rewriting the whole pipeline.
Engineering / pipelines-data atomic for-engineers light-setup from-file
Source Creator's repository · nvidia/skills
License: Apache-2.0
Security Verified — safe to install
Passed all 3 independent security checks
Checked by 3 independent security firms
Does it try to trick the AI? No SAFE · Gen Agent Trust Hub
Does it sneak in hidden code? No No alerts · Socket
Does it have known bugs? No Low risk · Snyk
Converts a single-machine training script into a multi-node Slurm job, handling distributed PyTorch setup, NCCL config, and the specific OOM and timeout patterns that break at scale.
↓ Download .skillDrag into Cowork or Claude Desktop — no terminal needed.
Open in Claude Code →
Copy install command→
1,467 installs
repository nvidia/skills
installs 1.5k
added Jun 8, 2026
last indexed Jun 8, 2026